Skip to main content
Aggregate arXiv cs.AI 人工智能 17 Aug 2026 - 12:30

Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference.…

  • We study a different possibility: stable miscalibration, where a confi…
  • We combine two diagnostics: a label-aware output-level audit score tha…
  • On a multi-domain binary factual audit set, this audit score tracks wh…

摘要引擎:抽取

正文提要

arXiv:2608.13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement. On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines rank the same gain more strongly. Internally, self-critical prompting consistently reduces hidden-state sensitivity across layers in three open-weight models. This supports prompt-induced local stabilization rather than a purely output-level abstention pattern, but it does not imply calibration: audit-defined overconfident errors are not clearly more locally sensitive than confidently correct answers, so some high-confidence errors may be stable and miscalibrated rather than simply fragile.

来源:https://arxiv.org/abs/2608.13591

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表