Work on semantic uncertainty has mostly asked whether a model can tell right answers from wrong ones, and skipped whether its confidence numbers mean anything. This AISTATS 2026 paper studies both. The finding is unusually practical: tuning a single token-level temperature parameter consistently improves calibration, discrimination, and semantic entropy across models, datasets, and confidence measures, beating fixed-temperature baselines and more elaborate calibration methods.
.jpg)
Testimonials
