Checking a long model response claim by claim is expensive and breaks down on open-ended prompts. This ICML 2026 paper proposes a cheaper first-pass signal. Generate several responses, embed them, and measure how widely those embeddings spread across the unit sphere. Greater dispersion reliably signals lower factual consistency. The method needs no labeled data, no fine-tuning, and no hyperparameter selection, and works with open or closed weight embedding models.
.jpg)
Testimonials
