微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
当前信源:arXiv cs.AI
· 清除信源筛选
1 / 24
Designing AI Pipelines for Decision-Ready ITSM Intelligence
arXiv:2608.…
12670v1 Announce Type: new Abstract: IT service management (ITSM) syst…
This paper presents a sociotechnical AI pipeline, designed and evaluat…
The pipeline combines LLM-based schema normalization, HDBSCAN sub-topi…
RSS 官方收录 · 可信分层展示
General Probabilities of Causation with Causal Knowledge
arXiv:2608.…
12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs)…
Tian and Pearl first derived theoretically sharp bounds for binary PoC…
Mueller et al.
RSS 官方收录 · 可信分层展示
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment.…
The steering decision is the pre-commit choice at that boundary: proce…
We introduce SteerBench-Work, an incident-anchored, bidirectional benc…
Release v2026-05 contains 106 scenarios anchored in public incidents, …
RSS 官方收录 · 可信分层展示
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately.…
The dominant delivery model is installation: once installed, a skill's…
This leaves the long tail with no practical path to use and forces tea…
We observe that installation bundles three separable functions: conten…
RSS 官方收录 · 可信分层展示
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
Judges are typically validated by accuracy on golden data, but accurac…
We introduce the \emph{Wiggle Framework}, a unified stress test for ep…
The framework decomposes judge robustness along three dimensions: Mech…
RSS 官方收录 · 可信分层展示
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
arXiv:2608.…
12599v1 Announce Type: new Abstract: Multi-turn dialogues let users re…
No existing instrument measures this influence per clause, predicts it…
\sysname{} closes the three gaps through the model API alone: a contra…
RSS 官方收录 · 可信分层展示
DiG-bench: Discovery in Games
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process.…
Despite its importance, there is a gap in the current AI benchmark lan…
To address this gap, we release a new benchmark: DiG-bench (Discovery …
DiG-bench consists of a set of 70 independent games.
RSS 官方收录 · 可信分层展示
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
arXiv:2608.…
12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requ…
We present ThyroidXAgent, a clinician-interactive agentic AI system th…
The system was developed using OpenThyroidDB, a multicentre, multitask…
RSS 官方收录 · 可信分层展示
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
arXiv:2608.…
12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires…
Additionally, surfacing reasoning mistakes that the model makes would …
Due to the difficulty of this complex task on long reasoning traces, s…
RSS 官方收录 · 可信分层展示
Trie Automata for Constrained Decoding over Large Finite Sets
arXiv:2608.…
12574v1 Announce Type: new Abstract: Large language models increasingl…
Current constrained decoding systems handle this through general-purpo…
We introduce the trie automaton, a specialized mechanism that exploits…
RSS 官方收录 · 可信分层展示
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
arXiv:2608.…
12555v1 Announce Type: new Abstract: Predictive explanation methods at…
We introduce the Causal Attribution Score (CAS), a compact score archi…
CAS starts from an identified interventional coalition game, allocates…
RSS 官方收录 · 可信分层展示
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
arXiv:2608.…
12522v1 Announce Type: new Abstract: LLM-based program evolution syste…
We introduce $\varepsilon$-MemEvo, a framework for cross-task knowledg…
$\varepsilon$-MemEvo stores prior experience as task-agnostic tactic m…
RSS 官方收录 · 可信分层展示
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
arXiv:2608.…
12476v1 Announce Type: new Abstract: Long-term agent memory is usually…
We introduce Governed Persistent Memory (GPM), an auditable bitemporal…
Five executable clauses cover ledger integrity, source binding, confli…
RSS 官方收录 · 可信分层展示
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
arXiv:2608.…
12428v1 Announce Type: new Abstract: Memory is a core component of AI …
However, existing memory systems often remain fixed after development,…
We present MindMemOS, a portable and self-evolving memory operating la…
RSS 官方收录 · 可信分层展示
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
arXiv:2608.…
12426v1 Announce Type: new Abstract: Large language models are increas…
Individual constraints are handled proficiently, but the compositional…
We introduce Constraint Saturation Evaluation (CSE), a procedurally ge…
RSS 官方收录 · 可信分层展示
Research Assistant: AstraZeneca's Agentic System for R&D
arXiv:2608.…
12395v1 Announce Type: new Abstract: We describe Research Assistant, a…
The system provides a chat-style interface that brings together eviden…
It supports both a fast mode for direct question answering and a multi…
RSS 官方收录 · 可信分层展示
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
arXiv:2608.…
12385v1 Announce Type: new Abstract: As large language models serve mo…
The two inference phases stress hardware differently: prompt prefill i…
Conventional width or depth scaling increases both costs together beca…
RSS 官方收录 · 可信分层展示
Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
arXiv:2608.…
12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot pe…
Existing adaptation methods struggle to calibrate update magnitude und…
To calibrate adaptation to evidence quality, we propose PAC-Bayes-regu…
RSS 官方收录 · 可信分层展示
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
arXiv:2608.…
12373v1 Announce Type: new Abstract: Large language models are increas…
We test nine models from six providers and ask whether the language of…
We use single-turn game-theoretic vignettes in which a model advises a…
RSS 官方收录 · 可信分层展示
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.…
This position paper argues that in many settings, particularly high-st…
We review evidence that cognitive alignment improves understandability…
We outline the gaps between existing alignment methods and what is nee…
RSS 官方收录 · 可信分层展示
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).…
Yet agreement in final labels does not show that human annotators and …
Two agents may reach the same judgment while appealing to different pr…
We test this distinction using a curated 500-item ETHICS-derived bench…
RSS 官方收录 · 可信分层展示
Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
arXiv:2608.…
12371v1 Announce Type: new Abstract: Stream-processing systems increas…
This paper proposes \emph{MAS-DecStream}, whose main contribution is \…
Edge-cluster agents refine natural-language offloading proposals from …
RSS 官方收录 · 可信分层展示
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
arXiv:2608.…
12346v1 Announce Type: new Abstract: This position paper argues that m…
By mapping current alignment techniques to the possibility and actual …
We need to discuss this dual-use potential now, as its risk is exacerb…
RSS 官方收录 · 可信分层展示
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
arXiv:2608.…
12345v1 Announce Type: new Abstract: Language models are increasingly …
We introduce IntegrityBench, a benchmark evaluating misconduct classif…
Evaluating 18 frontier model variants, we find that under peak pressur…
RSS 官方收录 · 可信分层展示
上滑下一条
上滑 · j/k · m/u · h 隐藏 · a 稍后 · o 原文 · e 详情 · t 今日 · i 模式 · f 搜索