ConfTuner: Tokenized Brier Score for Calibrated LLM Confidence Distributions

The article compares TypeSafe AI's Jev model with the NeurIPS 2025 ConfTuner paper, which introduces Tokenized Brier Score to train LLMs to output calibrated probability distributions over confidence tokens, achieving significant ECE reduction across multiple benchmarks with minimal training time, and extends the approach to decision tokens via JevTuner.

Machine Heart
Machine Heart
Machine Heart
ConfTuner: Tokenized Brier Score for Calibrated LLM Confidence Distributions

Recent attention on TypeSafe AI's Jev model highlights a paradigm where models output decisions alongside calibrated probabilities. Though Jev remains closed-source, its API reveals four core traits: direct decision-plus-probability output, parallel probability prediction, probability as the core prediction target, and pursuit of probability calibration so that reported confidence matches actual accuracy.

These traits align closely with the NeurIPS 2025 paper ConfTuner from the National University of Singapore. ConfTuner introduces Tokenized Brier Score , a method that directly trains the probability distribution over candidate confidence tokens (0% to 100%) rather than only the final scalar confidence value.

Turning "I'm 80% sure" into a trainable token probability distribution

After a model answers a question, it generates a confidence token (e.g., "Confidence: 80%"). At that step, the model has already computed logits for all vocabulary tokens. ConfTuner extracts the logits corresponding to confidence tokens representing 0% through 100%, applies softmax over them, and obtains a full probability distribution over confidence values. Crucially, these candidate probabilities are produced in a single forward pass — no separate generation per value.

This mirrors Jev's interface: both expose the entire probability distribution over candidates and focus on how that distribution is produced and calibrated, not just the final picked value.

Learning "how confident" from only correctness labels

ConfTuner's core is the Tokenized Brier Score loss. Let p_i be the probability assigned to confidence token c_i (where c_i is the numeric confidence, e.g., 0.8), and y ∈ {0,1} indicate whether the answer is correct. The loss is: L = Σ_i p_i (c_i - y)^2 Intuitively, if the model is wrong ( y=0) but places high probability on high-confidence tokens (e.g., 99%), the loss penalizes heavily. If correct ( y=1), the loss encourages higher confidence. This requires only binary correctness labels — no human-annotated confidence values.

The paper proves this loss is a Proper Scoring Rule : minimizing expected loss forces the probability mass to concentrate on the token whose numeric confidence matches the true accuracy, providing theoretical grounding for calibration.

Experimental validation: significant calibration error reduction

Experiments fine-tuned LLaMA, Qwen, and Ministral on 2,000 HotpotQA samples (4×A40 GPUs, ~4 minutes) and evaluated on GSM8K, TriviaQA, StrategyQA, and TruthfulQA. The primary metric is Expected Calibration Error (ECE) — lower is better.

Average ECE across five datasets dropped from 0.2768, 0.3781, 0.4393 to 0.1082, 0.2872, 0.1884 for the three base models respectively. Reliability diagrams show markedly reduced area between predicted confidence and actual accuracy, outperforming baselines.

Training efficiency is notable: ConfTuner takes ~4 minutes versus 26 minutes for LACIE and 120 minutes for SaySelf under the same hardware.

Calibrated probabilities enable downstream gains: routing low-confidence answers to self-correction or stronger models. Under equal compute budget, ConfTuner-based cascading improved accuracy by up to 9.3% on HotpotQA and 5.5% on TruthfulQA.

From confidence tokens to decision tokens

The ConfTuner team open-sourced JevTuner (https://github.com/liushiliushi/JevTuner) to explore extending the same principle to business decisions. Instead of confidence tokens, candidate decision tokens (e.g., "billing", "technical", "sales") occupy a single token slot; their logits yield a full decision probability distribution. Training uses the correct decision label with multi-class Brier Loss.

Example inference output:

{
  "decision": "billing",
  "probabilities": {
    "billing": 0.84,
    "technical": 0.09,
    "sales": 0.03,
    "other": 0.04
  }
}

JevTuner is not a replication of Jev's undisclosed internals but a feasibility demonstration of the shared direction: treat candidate probabilities as first-class model outputs and train the model to be accountable for those probabilities themselves.

As AI decisions enter automated workflows, mere correctness is insufficient; the reliability of the reported probabilities becomes the foundation for industrial-scale trust.

Paper: https://arxiv.org/abs/2508.18847

Code: https://github.com/liushiliushi/ConfTuner

Extension: JevTuner https://github.com/liushiliushi/JevTuner

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

NeurIPS 2025Jevprobability calibrationConfTunerECEJevTunerLLM confidenceTokenized Brier Score
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.