Skip to main content
Aggregate arXiv cs.AI 人工智能 15 Aug 2026 - 12:00

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).…

  • Yet agreement in final labels does not show that human annotators and …
  • Two agents may reach the same judgment while appealing to different pr…
  • We test this distinction using a curated 500-item ETHICS-derived bench…

摘要引擎:抽取

正文提要

arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.

来源:https://arxiv.org/abs/2608.12368

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表