PINOCCHIO // EXTERNAL CALIBRATOR FOR BLACK-BOX LLMS

Fast uncertainty estimates for black-box language models  /  P(correct) in a single forward pass

Pinocchio

An external calibrator that estimates the correctness of responses from black-box API models. It needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states.

Kevin David Hayes, Arka Pal, Haosong Zhang, Tom Goldstein, Micah Goldblum

University of Maryland · Ritual AI · Fudan University · Columbia University

Accepted to COLM 2026

0.863
AUROC predicting correctness on the held-out test set, compared with 0.649 for the strongest single-pass black-box baseline
0.814
mean AUROC in zero-shot transfer to thirteen unseen models across eight organizations
0.076 ECE
after Platt scaling on 100 labeled examples, down from 0.259, on four unseen open models
1 pass
single forward pass, without sampling, logit access, or modification of the target model
01

The access problem

In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions.

Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, many industrial LLM products use closed-source API models, and many such API models like GPT do not return log-probabilities and may not allow fine-tuning.

Rather than extracting uncertainty from the target model itself, we train a small calibrator that takes the question, the response, and the identity of the model that produced it, plus any associated images for vision-language tasks, and outputs a probability of correctness.

Fig. 01-Athe sealed endpoint, one channel out
prompt CLOSED API logits weights internal states fine-tuning sealedsealedsealedsealed question + response no logits or weights PINOCCHIO P(correct)
Pinocchio requires no access to the target model's logits, weights, or internal states. It reads the question and the response, plus any image, and outputs a probability of correctness.
Black-box alternatives pay a steep cost: verbalized confidence is poorly calibrated. On the held-out test set it reaches 0.610 AUROC, against 0.863 for Pinocchio.
02

One calibrator, seven sources

We train Pinocchio jointly on responses from seven LLMs that span a wide capability range, across 20 benchmarks: 9 text-only and 11 vision-language.

It reaches 0.862 AUROC predicting the correctness of held-out responses from those same models.

Since strong models achieve very high scores on popular benchmarks, the central challenge in training a calibrator is assembling enough examples where strong models are wrong.

We address this deficit by selecting benchmarks with both correct and incorrect responses, including vision-language data where models tend to be weaker, and pooling responses from several target models.

High-capability 4 models

Claude Fable 5 Claude Opus 5 GPT-5.6 Kimi 3

Lower-capability 3 models

GPT-5-mini GPT-5.2 Qwen3.5-397B
All seven train one calibrator jointly.
03

Transfer to models it never saw

We evaluate zero-shot transfer using the same calibrator. It scores thirteen target models absent from training, spanning eight organizations; several were released after the calibrator was trained, including GPT-6 Astra. Mean AUROC is 0.814 across the thirteen targets, a modest drop from the 0.862 it achieves on models whose responses we used during training.

Trace 03-Azero-shot transfer, no target labels
mean of 13 · 0.814 Gemini 3.1 Pro0.863 Gemini 3 Flash0.832 Mistral-7B0.828 Devstral-2-24B†0.826 LLaMA-3.1-8B0.825 OLMo-2-7B0.825 Claude Sonnet 4.60.807 GPT-6 Astra†0.804 DeepSeek-R1-Distill-32B0.802 gemma-2-9b0.801 Olmo-3.1-32B†0.793 granite-4.1-30b†0.793 Claude Opus 4.60.783 0.50.70.9 AUROC, zero-shot transfer to all 13 target models
Zero-shot transfer AUROC on target models absent from training, spanning eight organizations. † Released after the calibrator was trained.
04

Why it transfers: capability diversity

Lower-capability training sources improve broad transferability. Trained on the four high-capability models alone, mean transfer AUROC is 0.649. Adding GPT-5-mini raises it to 0.778, and adding GPT-5.2 and Qwen3.5 raises it to 0.814.

Transfer, especially to lower-capability target models, improves as lower-capability source models are added. This effect is especially pronounced on individual targets such as Mistral-7B, where each successive training mixture raises transfer.

Trace 04-Amean transfer AUROC by training sources
0.50.70.9 0.649 0.778 0.814 4 high-capability+ GPT-5-mini+ GPT-5.2, Qwen3.5 mean transfer AUROC over eleven unseen models
Lower-capability training sources improve broad transferability. Mean AUROC is measured over eleven unseen models shared by all three training mixtures.
05

Correct it with about 100 labels

Approximately 100 labeled samples can be used to re-calibrate after zero-shot transfer. On four open models absent from training, Platt scaling fit on 100 labeled examples reduces pooled expected calibration error (ECE) from 0.259 to 0.076.

Platt scaling preserves AUROC because it is monotone, so each model's ranking is unchanged. Isotonic regression lowers pooled ECE further, to 0.058, but introduces ties.

Trace 05-Areliability, before and after 100 labels (schematic)
perfect 01 confidence 01 accuracy before ECE 0.259 after Platt, 100 labels ECE 0.076 ranking preserved
Curves are schematic. ECE values are pooled over Mistral-7B, gemma-2-9b, OLMo-2-7B and granite-4.1-30b, with recalibration fit on 100 labeled examples and evaluated on the rest over 25 random splits.
06

How it reads

The calibrator reads the question and response and predicts a single token, i (incorrect) or ii (correct). We extract the correctness probability via softmax over the two logits:

$$P(\text{correct}) = \frac{e^{z_{\text{ii}}}}{e^{z_{\text{i}}} + e^{z_{\text{ii}}}}$$

These are the calibrator's logits, not the target's. A single forward pass produces a correctness probability without sampling, logit access, or modification of the target model.

On the held-out test set, Pinocchio reaches 0.863 AUROC, compared with 0.649 for the strongest single-pass black-box baseline. A lightweight text-only 0.8B checkpoint matches our largest model's AUROC.

Question: {Q} Answer: {A} Is the answer correct? (i) No (ii) Yes
Trace 06-AAUROC against baselines
Pinocchio0.863 verb. + length0.649 verbalized0.610 LLM-as-judge0.532 0.50.751.0 AUROC, held-out test set (1,953)

Training spans 20 benchmarks, 9 text-only and 11 vision-language, including GPQA Diamond, SimpleQA, BBEH, HLE, MMMU and MathVista.

07

Usage

quickstart.pypython
# pip install pinocchio-uq
from pinocchio import Pinocchio

judge = Pinocchio()  # load once

response = client.chat.completions.create(model="gpt-5", messages=messages)
p_correct = judge.score(response, messages=messages)

if p_correct < 0.5:
    escalate(response)

We release code for adding uncertainty estimation to existing repos in only two additional lines of code. The package ships a lightweight text-only 0.8B checkpoint.

08

Cite

pinocchio.bibbibtex
@inproceedings{hayes2026pinocchio,
  title     = {Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models},
  author    = {Hayes, Kevin David and Pal, Arka and Zhang, Haosong and
               Goldstein, Tom and Goldblum, Micah},
  booktitle = {Conference on Language Modeling (COLM)},
  year      = {2026}
}