Fast uncertainty estimates for black-box language models / P(correct) in a single forward pass
An external calibrator that estimates the correctness of responses from black-box API models. It needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states.
University of Maryland · Ritual AI · Fudan University · Columbia University
Accepted to COLM 2026
In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions.
Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, many industrial LLM products use closed-source API models, and many such API models like GPT do not return log-probabilities and may not allow fine-tuning.
Rather than extracting uncertainty from the target model itself, we train a small calibrator that takes the question, the response, and the identity of the model that produced it, plus any associated images for vision-language tasks, and outputs a probability of correctness.
We train Pinocchio jointly on responses from seven LLMs that span a wide capability range, across 20 benchmarks: 9 text-only and 11 vision-language.
It reaches 0.862 AUROC predicting the correctness of held-out responses from those same models.
Since strong models achieve very high scores on popular benchmarks, the central challenge in training a calibrator is assembling enough examples where strong models are wrong.
We address this deficit by selecting benchmarks with both correct and incorrect responses, including vision-language data where models tend to be weaker, and pooling responses from several target models.
We evaluate zero-shot transfer using the same calibrator. It scores thirteen target models absent from training, spanning eight organizations; several were released after the calibrator was trained, including GPT-6 Astra. Mean AUROC is 0.814 across the thirteen targets, a modest drop from the 0.862 it achieves on models whose responses we used during training.
Lower-capability training sources improve broad transferability. Trained on the four high-capability models alone, mean transfer AUROC is 0.649. Adding GPT-5-mini raises it to 0.778, and adding GPT-5.2 and Qwen3.5 raises it to 0.814.
Transfer, especially to lower-capability target models, improves as lower-capability source models are added. This effect is especially pronounced on individual targets such as Mistral-7B, where each successive training mixture raises transfer.
Approximately 100 labeled samples can be used to re-calibrate after zero-shot transfer. On four open models absent from training, Platt scaling fit on 100 labeled examples reduces pooled expected calibration error (ECE) from 0.259 to 0.076.
Platt scaling preserves AUROC because it is monotone, so each model's ranking is unchanged. Isotonic regression lowers pooled ECE further, to 0.058, but introduces ties.
The calibrator reads the question and response and predicts a single token, i (incorrect) or ii (correct). We extract the correctness probability via softmax over the two logits:
These are the calibrator's logits, not the target's. A single forward pass produces a correctness probability without sampling, logit access, or modification of the target model.
On the held-out test set, Pinocchio reaches 0.863 AUROC, compared with 0.649 for the strongest single-pass black-box baseline. A lightweight text-only 0.8B checkpoint matches our largest model's AUROC.
Training spans 20 benchmarks, 9 text-only and 11 vision-language, including GPQA Diamond, SimpleQA, BBEH, HLE, MMMU and MathVista.
# pip install pinocchio-uq
from pinocchio import Pinocchio
judge = Pinocchio() # load once
response = client.chat.completions.create(model="gpt-5", messages=messages)
p_correct = judge.score(response, messages=messages)
if p_correct < 0.5:
escalate(response)
We release code for adding uncertainty estimation to existing repos in only two additional lines of code. The package ships a lightweight text-only 0.8B checkpoint.
@inproceedings{hayes2026pinocchio,
title = {Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models},
author = {Hayes, Kevin David and Pal, Arka and Zhang, Haosong and
Goldstein, Tom and Goldblum, Micah},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026}
}