微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
RSS 官方收录 · 可信分层展示
关键摘要
9个前沿LLM法官在压力下25%-91%概率翻转判决,且翻转多偏离真相
- Wiggle框架从机械一致性、单轮信念、多轮坚持三维度测试LLM法官稳定性
- 所有9个模型在静态挑战下判决翻转率25%–71%,对抗说服下达62%–91%
- 判决翻转通常导致偏离真实答案,基线陪审团多数票是最优预测信号
AI 摘要 · 来源可核验
正文提要
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\% of the time under static pushback, and 62--91\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.