Skip to main content
Aggregate arXiv cs.AI 人工智能 25 Aug 2026 - 12:00

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 21374v1 Announce Type: new Abstract: Literature reviews are essential …
  • We introduce LitReview Arena, a battle-style evaluation platform with …
  • From this protocol, we collect approximately 3k expert judgments, each…

摘要引擎:抽取

正文提要

arXiv:2608.21374v1 Announce Type: new Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.

来源:https://arxiv.org/abs/2608.21374

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表