Skip to main content

XMT

短闻

信源:arXiv cs.AI · 短平快可信阅读。

抖音式连刷节奏 · 视频号式正能量克制 · 第三种:可信信流——短平快,可核验。

今日 稍后 搜索 RSS

当前信源:arXiv cs.AI · 清除信源筛选

Aggregate arXiv cs.AI 人工智能 45″

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

arXiv:2608.18580v1 Announce Type: new Abstract: Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging.…

  • Each task couples an instruction, an initialized environment, a refere…
  • Meanwhile, multi-stage synthesis can discard the goals, dependencies, …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement

arXiv:2608.…

  • 18543v1 Announce Type: new Abstract: Modern e-commerce platforms often…
  • This is particularly challenging for exploratory intents such as best …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

arXiv:2608.…

  • 18534v1 Announce Type: new Abstract: Large language models are increas…
  • In financial reconciliation, the evidence needed for diagnosis is dist…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

arXiv:2608.…

  • 18531v1 Announce Type: new Abstract: Industrial explainable-recommenda…
  • We separate generation from selection: explanations are produced ahead…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval

arXiv:2608.…

  • 18521v1 Announce Type: new Abstract: Dense-caption retrieval has recen…
  • However, these methods largely inherit the same InfoNCE objective, who…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

arXiv:2608.…

  • 18504v1 Announce Type: new Abstract: Universal multimodal retrieval ai…
  • Recent MLLM-based embedding methods typically derive representations f…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

arXiv:2608.18423v1 Announce Type: new Abstract: Language model agents now execute bounded tasks reliably.…

  • Whether they can sustain effective decision-making over long horizons,…
  • FM-Bench (Football Management Benchmark) measures this.

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions

arXiv:2608.…

  • 18409v1 Announce Type: new Abstract: Combinatorial scheduling poses a …
  • This challenge is especially pronounced in resource-constrained settin…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification

arXiv:2608.18397v1 Announce Type: new Abstract: Wearable stress classifiers can achieve strong average performance while failing completely for a particular individual.…

  • On WESAD, a Random Forest reaches 93.
  • 0% mean accuracy yet yields F1 = 0 for Subject 14, whose cross-signal …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations

arXiv:2608.…

  • 18389v1 Announce Type: new Abstract: AI code agents are increasingly d…
  • We evaluate whether coding agents that repair repository-level issues …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

arXiv:2608.…

  • 18336v1 Announce Type: new Abstract: When evaluating language models o…
  • This approach assumes that partial knowledge is worth proportional cre…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

arXiv:2608.…

  • 18324v1 Announce Type: new Abstract: Machine-verifiable workflows prod…
  • We test whether these records can supervise bounded models, consolidat…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition

arXiv:2608.…

  • 18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces r…
  • We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a traini…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

arXiv:2608.18307v1 Announce Type: new Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests.…

  • This leaves an under-instrumented middle layer: realistic component-ce…
  • , toggle a button set) that are short enough to diagnose and rich enou…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

arXiv:2608.…

  • 18300v1 Announce Type: new Abstract: LLM-as-a-Judge, which leverages a…
  • However, most work treats a judge as a static artifact, evaluating it …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

arXiv:2608.…

  • 18289v1 Announce Type: new Abstract: The extraction of structured info…
  • While proprietary solutions dominate commercial applications, a rapidl…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study

arXiv:2608.…

  • 18261v1 Announce Type: new Abstract: Serving a 235B-parameter Mixture-…
  • We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measur…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Redakto - The Incognito Tab for LLMs

arXiv:2608.18260v1 Announce Type: new Abstract: Large Language Models (LLMs) are being increasingly used in everyday applications.…

  • A major challenge in the context of LLMs or Artificial Intelligence (A…
  • These challenges have become more urgent with novel EU legislation.

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks

arXiv:2608.…

  • 18238v1 Announce Type: new Abstract: Genomic analysis on viruses such …
  • These approaches use pairwise codon or nucleotide distance matrices to…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices

arXiv:2608.18194v1 Announce Type: new Abstract: This paper presents new theoretical results on generalizing the Jaccard distance for lattices and real valuations.…

  • We demonstrate that when the valuation is strictly positive, monotone,…
  • Moving to relatively complemented distributive lattices (which safely …

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Looped Language Models Improve Compositional Tool Calling

arXiv:2608.…

  • 18171v1 Announce Type: new Abstract: Looped language models have shown…
  • We study this question in compositional tool-calling settings, where m…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Adversarial Review: Structured Disagreement for Grounded Agentic Code Review

arXiv:2608.…

  • 18167v1 Announce Type: new Abstract: Early multi-agent LLM systems oft…
  • Recent alternatives treat agents as passive tools (subagents), yet thi…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

RDFdL: Integrating RDF with Differential Dynamic Logic

arXiv:2608.…

  • 18165v1 Announce Type: new Abstract: Knowledge graphs modeled in RDF a…
  • , systems described by differential equations, which is a critical gap…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu

arXiv:2608.…

  • 18142v1 Announce Type: new Abstract: It is challenging to detect hate …
  • A good example of such a challenge is Roman Urdu which is broadly used…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

arXiv:2608.…

  • 18136v1 Announce Type: new Abstract: Conversational agents now act for…
  • Banking is the clearest case: the same agent that answers a question c…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Improving Rural Medication Safety with AI: A Scoping Review

arXiv:2608.18135v1 Announce Type: new Abstract: Introduction: Medication errors (MEs) represent a significant threat to global healthcare systems, contributing to patient harm.…

  • Introducing artificial intelligence (AI) in rural healthcare enhances …
  • The aim is to explore the applications and effectiveness of AI technol…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis

arXiv:2608.18133v1 Announce Type: new Abstract: Reliable transformer fault diagnosis is essential for maintaining power system stability.…

  • The IEEE Key Gas Method (KGM), a widely utilized approach in Dissolved…
  • This study presents An enhanced model combining Fuzzy Logic with the I…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

arXiv:2608.…

  • 18117v1 Announce Type: new Abstract: This position paper argues that A…
  • The barrier is not missing data; high-quality regional benchmarks alre…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

arXiv:2608.18131v1 Announce Type: new Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric.…

  • When such safety filters fail for non-English languages, the consequen…
  • For spoken language technologies deployed across India's linguisticall…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图
Aggregate arXiv cs.AI 人工智能 45″

Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry

arXiv:2608.…

  • 18111v1 Announce Type: new Abstract: Foundation models such as GPT and…
  • Yet solving a geometry problem and drawing the figure it depends on ar…

RSS 官方收录 · 可信分层展示

信流 详情 原文 分享图