Skip to main content
Aggregate arXiv cs.AI 人工智能 15 Aug 2026 - 15:00

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

RSS 官方收录 · 可信分层展示

关键摘要

arXiv:2608.…

  • 12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Res…
  • We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a R…
  • The framework operates through two synergistic components: the Academi…

摘要引擎:抽取

正文提要

arXiv:2608.12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

来源:https://arxiv.org/abs/2608.12788

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表