Skip to main content
Aggregate arXiv cs.AI 人工智能 19 Aug 2026 - 13:00

KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 17150v1 Announce Type: new Abstract: To effectively collaborate with u…
  • Yet user simulators used to evaluate and train LLMs do not explicitly …
  • To close this gap, we introduce KNOWSIM, an evaluation framework built…

摘要引擎:抽取

正文提要

arXiv:2608.17150v1 Announce Type: new Abstract: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.

来源:https://arxiv.org/abs/2608.17150

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表