Skip to main content

XMT

短闻

信源:arXiv cs.AI · 短平快可信阅读。

抖音式连刷节奏 · 视频号式正能量克制 · 第三种:可信信流——短平快,可核验。

今日 稍后 搜索 RSS

当前信源:arXiv cs.AI · 清除信源筛选

Aggregate arXiv cs.AI 人工智能 45″

FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence

arXiv:2608.…

  • 26310v1 Announce Type: new Abstract: Large language models can now gen…
  • Existing evaluation approaches largely depend on model-based natural-l…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems

arXiv:2608.…

  • 26306v1 Announce Type: new Abstract: A large language model (LLM) guar…
  • This creates an Execute-stage time-of-check to time-of-use (TOCTOU) ha…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

SKILL.state: Scalable Long-Horizon Agent Skills

arXiv:2608.26263v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills.…

  • Existing agent runtimes maintain execution by continually appending ob…
  • We present SKILL.

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Assessing mentalization in humans and large language models

arXiv:2608.…

  • 26291v1 Announce Type: new Abstract: Mentalization - the ability to in…
  • Large language models (LLMs) demonstrate behaviour consistent with hum…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation

arXiv:2608.…

  • 26236v1 Announce Type: new Abstract: We present a six-stage framework …
  • Instantiating the framework on the NSAI subdomain produced a multi-yea…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

LLM Agents for Time-Series: A Survey

arXiv:2608.…

  • 26226v1 Announce Type: new Abstract: LLM-based agents are increasingly…
  • This survey adopts a problem-driven taxonomy that organizes these syst…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

arXiv:2608.…

  • 26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of rea…
  • We introduce the Token Economy Score (TES), a marginal benchmarking me…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy

arXiv:2608.26225v1 Announce Type: new Abstract: Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them.…

  • The machinery such orchestrators reach for is the service mesh's: retr…
  • We report a failure study of a production agentic software-delivery pl…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Same Model, Different Harness: Different Coding-Agent Results

arXiv:2608.…

  • 26218v1 Announce Type: new Abstract: A coding agent combines a model w…
  • We ask whether changing the harness changes the result when the model …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

GameWAM: A World Action Model for Video Games

arXiv:2608.26200v1 Announce Type: new Abstract: Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls.…

  • Existing game agents map visual and task context directly to actions b…
  • World-Action Models (WAMs) unify these objectives, but remain largely …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling

arXiv:2608.…

  • 26199v1 Announce Type: new Abstract: We ask whether AI agents powered …
  • In these environments, engineers issue repetitive, dependency-ordered …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

arXiv:2608.…

  • 26193v1 Announce Type: new Abstract: Multimodal large language models …
  • Models may predict correct answers while neglecting people-centric cue…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Agentic AI for operating scientific instruments for nanoscale characterization

arXiv:2608.26198v1 Announce Type: new Abstract: Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert decision-making.…

  • A trained user defines the experimental intent, translates it into ins…
  • Existing automation usually addresses only parts of this workflow thro…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs

arXiv:2608.…

  • 26191v1 Announce Type: new Abstract: Incident risk prediction from lon…
  • We propose structured evidence routing, a router-predictor-reviewer wo…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

arXiv:2608.…

  • 26190v1 Announce Type: new Abstract: World models enable agents to rea…
  • In this work, we propose a compatibility prediction Latent World Model…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Invocation-Level Reliability of Tool-Using Agents

arXiv:2608.…

  • 26189v1 Announce Type: new Abstract: Tool-using agents fail two ways: …
  • We measure a correct-invocation rate that separates the two, under bot…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments

arXiv:2608.26188v1 Announce Type: new Abstract: Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel.…

  • We ask whether such judgments track measured risk or the patterns atta…
  • We probe seven instruct-tuned models under three conditions that disso…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion

arXiv:2608.…

  • 26185v1 Announce Type: new Abstract: Equal participation in co-located…
  • We present SecondVoice, a mixed-reality system that enables people to …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education

arXiv:2608.…

  • 26184v1 Announce Type: new Abstract: AI programming tutors provide sca…
  • We present TutorTrace, a dataset and behavioral abstraction pipeline t…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI

arXiv:2608.…

  • 26182v1 Announce Type: new Abstract: Large language models (LLMs) are …
  • As a result, robots may present hallucinated capabilities, unclear beh…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Knowledge Cards: Structured Knowledge for AI Systems

arXiv:2608.…

  • 26176v1 Announce Type: new Abstract: AI systems whose outputs inform r…
  • Established documentation artefacts already capture important aspects …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

AI Revealed Preferences

arXiv:2608.26178v1 Announce Type: new Abstract: There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons.…

  • We test 20 language models and find a range of preferences---stable di…
  • We run three forced-choice experiments on revealed rather than stated …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript

arXiv:2608.…

  • 26167v1 Announce Type: new Abstract: Hallucination and abstention benc…
  • Seven large language models were evaluated on the TAME Pain speech cor…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving

arXiv:2608.…

  • 26164v1 Announce Type: new Abstract: Large language models can interpr…
  • This paper presents ChemOntoRule, a proof-of-concept symbolic core for…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian

arXiv:2608.…

  • 26162v1 Announce Type: new Abstract: Safety-critical mental-health sup…
  • This paper presents Anian, a safety-gated multimodal AI backend for pe…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum

arXiv:2608.…

  • 26160v1 Announce Type: new Abstract: Nowadays semantic models provide …
  • Therefore, AI processes and resources are often described using incomp…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions

arXiv:2608.…

  • 26157v1 Announce Type: new Abstract: Natural-language analytics over e…
  • Existing text-to-SQL systems often ground generation in database schem…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Selection Bias Correction in Retail Intelligence

arXiv:2608.…

  • 26156v1 Announce Type: new Abstract: Retail intelligence often relies …
  • This simulation study investigates selection bias in inflation estimat…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG

arXiv:2608.…

  • 26153v1 Announce Type: new Abstract: Clinical electroencephalography (…
  • Most toolboxes focus on visualization or preprocessing, providing limi…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

arXiv:2608.26151v1 Announce Type: new Abstract: Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.…

  • 9% in mature markets eroding billions in revenue annually.
  • Predictive models can flag at-risk customers accurately, yet they are …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图