Jun 24, 2026
CALIBER: Calibrating confidence before and after reasoning in language models
CALIBER improves reasoning model calibration by eliciting state-dependent confidence estimates before and after reasoning with matched supervision targets, achieving up to 52.5% ECE reduction and best scores across multiple benchmarks.

Authors
Conor Finlay*, Joshua Kurien*, Saurabh Dash, Marzieh Fadaee, Beyza Ermis * First authors
Abstract
We introduce CALIBER (Calibration Before and After Reasoning), which elicits both estimates and supervises each with the target matched to its information state. Under this unified protocol, CALIBER reduces Expected Calibration Error (ECE) by 52.5% over the strongest single-confidence baseline on BigMathDigits for the 7B model, while achieving the best Brier score and AUROC, and remains within 2.1 points of the best accuracy. Further, on a larger 30B model, CALIBER achieves the best ECE on BigMathDigits while remaining competitive in Brier score and AUROC. Out of distribution, it achieves the best ECE and Brier score on GPQA and TriviaQA, and remains competitive on SimpleQA. Ablations further show that this position–target alignment is most beneficial under distribution shift where it consistently reduces calibration error across all out-of-distribution benchmarks.
Related works









