Skip to main content
Aggregate arXiv cs.AI 人工智能 2 Sep 2026 - 13:30

Asymmetries in Spontaneous and Instructed Deception

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to.…

  • However, much of the study on deception in models involves instructed …
  • We investigated the relationship between instructed and spontaneous (u…
  • 1-70B-Instruct.

摘要引擎:抽取

正文提要

arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.

来源:https://arxiv.org/abs/2609.00180

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表