What AI confidence scores actually tell your plugin

Confidence Gap: an isometric gauge splits into two chutes, one flows free, one waits at a gate, over faint dial line art.

Here is a gate I see in a lot of AI plugin builds, and it gets one thing backwards:

// $response is the decoded API payload
$probs = array_map( fn( $lp ) => exp( $lp->logprob ), $response->logprobs );
$avg   = array_sum( $probs ) / max( count( $probs ), 1 );

if ( $avg >= 0.9 && bbioon_grounding_check( $answer ) ) {
    bbioon_auto_apply( $answer );
} else {
    bbioon_queue_for_review( $answer, $avg );
}

This gates an AI answer before it writes to your catalog or your support queue. The problem is what $avg measures. It is the mean token probability the model reported while generating. Developers tend to read AI confidence scores as a probability of truth, when they sit closer to a measure of how fluently the model could produce those tokens. Sara A. Metwalli’s article on the confidence trap at Towards Data Science explains the mechanism, and it opens with a Nobel Prize question that got a confident answer before the prize existed.

What the number measures

Classifiers convert raw logits into percentages with softmax, which forces the outputs to sum to one and stretches small gaps into large ones. Cat: 0.97 means cat won among the options the model was offered, and nothing in training taught it to say none of the above. Show a toaster to a cat-and-dog classifier and you still get dog: 0.98. Hosted LLMs have the same problem one level down: logprobs describe how sure the sampler felt about each next token, and a fluent hallucination samples as smoothly as a fact. The number never saw your data, so it cannot know your answer is wrong.

Getting AI confidence scores you can act on

Option one is logprobs. Add logprobs: true and top_logprobs to the request, average the token probabilities that come back, and use the result for triage. As a truth estimate it is weak. As a wobble detector it is decent, since low averages cluster where the model gets unsure mid-answer. The OpenAI cookbook has a working example of the response shape. Not every provider exposes it.

Option two is asking the model to grade itself: return JSON with an answer field and a confidence field. It works with every provider and ships in most plugins. It also has the same flaw, because the model reports its own certainty as fluently as it reports facts. When it invents a Nobel laureate, it attaches 99 percent to the invention. I only use it for sorting a review queue.

Option three is calibration, and it only exists for models you train yourself. If you run your own classifier for product images or spam, fit temperature scaling on a validation split. The Guo et al. paper from 2017 shows that a single scalar parameter recovers most of the calibration that training burns off. You cannot recalibrate a hosted LLM, but you can log whether each auto-applied answer was later corrected and re-tune the threshold against that log.

For an internal tool where a wrong answer costs someone a reread, logprobs with a forgiving threshold is fine. For anything customer-facing, or anything that writes to products, orders and refunds, let the number sort the queue and make the gate a grounding check in PHP. Does the SKU exist? Does the order ID belong to this customer? Is the price inside the range the database knows?

Where this bites a WooCommerce build

Support triage that auto-denies refunds is the obvious case. Bulk description generation is sneakier, because a model writing fluent copy about a product it half-invented carries high token confidence through every sentence. Metwalli’s doctor example, a 99 percent cancer probability, maps to an owner reading a 98 percent fraud score on a refund claim and clicking approve. Confidence rises on the inputs where you want caution: odd products, mixed-language tickets, edge cases outside training. Most of what I check before running an AI product advisor comes down to boring database checks, and the transparency patterns for showing uncertainty in the interface are worth borrowing.

If you are adding an AI feature to a WordPress or WooCommerce site and have not decided what happens when the answer is wrong, I can help with that part, meaning the thresholds and grounding checks and the review queue that sits behind them.

What I still do not know is whether a logprob threshold keeps its meaning when a provider quietly updates the model behind the same API name. Nothing in the response tells you the snapshot changed. It matters because a threshold tuned on one model becomes a different filter overnight, while the plugin keeps reporting the old number with the old certainty.

author avatar
Ahmad Wael
I'm a WordPress and WooCommerce developer with 15+ years of experience building custom e-commerce solutions and plugins. I specialize in PHP development, following WordPress coding standards to deliver clean, maintainable code. Currently, I'm exploring AI and e-commerce by building multi-agent systems and SaaS products that integrate technologies like Google Gemini API with WordPress platforms, approaching every project with a commitment to performance, security, and exceptional user experience.

Leave a Comment