微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
当前信源:arXiv cs.AI
· 清除信源筛选
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
arXiv:2608.18580v1 Announce Type: new Abstract: Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging.…
Each task couples an instruction, an initialized environment, a refere…
Meanwhile, multi-stage synthesis can discard the goals, dependencies, …
RSS 官方收录 · 可信分层展示
Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement
arXiv:2608.…
18543v1 Announce Type: new Abstract: Modern e-commerce platforms often…
This is particularly challenging for exploratory intents such as best …
RSS 官方收录 · 可信分层展示
FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems
arXiv:2608.…
18534v1 Announce Type: new Abstract: Large language models are increas…
In financial reconciliation, the evidence needed for diagnosis is dist…
RSS 官方收录 · 可信分层展示
Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
arXiv:2608.…
18531v1 Announce Type: new Abstract: Industrial explainable-recommenda…
We separate generation from selection: explanations are produced ahead…
RSS 官方收录 · 可信分层展示
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
arXiv:2608.…
18521v1 Announce Type: new Abstract: Dense-caption retrieval has recen…
However, these methods largely inherit the same InfoNCE objective, who…
RSS 官方收录 · 可信分层展示
UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
arXiv:2608.…
18504v1 Announce Type: new Abstract: Universal multimodal retrieval ai…
Recent MLLM-based embedding methods typically derive representations f…
RSS 官方收录 · 可信分层展示
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
arXiv:2608.18423v1 Announce Type: new Abstract: Language model agents now execute bounded tasks reliably.…
Whether they can sustain effective decision-making over long horizons,…
FM-Bench (Football Management Benchmark) measures this.
RSS 官方收录 · 可信分层展示
Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions
arXiv:2608.…
18409v1 Announce Type: new Abstract: Combinatorial scheduling poses a …
This challenge is especially pronounced in resource-constrained settin…
RSS 官方收录 · 可信分层展示
When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification
arXiv:2608.18397v1 Announce Type: new Abstract: Wearable stress classifiers can achieve strong average performance while failing completely for a particular individual.…
On WESAD, a Random Forest reaches 93.
0% mean accuracy yet yields F1 = 0 for Subject 14, whose cross-signal …
RSS 官方收录 · 可信分层展示
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
arXiv:2608.…
18389v1 Announce Type: new Abstract: AI code agents are increasingly d…
We evaluate whether coding agents that repair repository-level issues …
RSS 官方收录 · 可信分层展示
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
arXiv:2608.…
18336v1 Announce Type: new Abstract: When evaluating language models o…
This approach assumes that partial knowledge is worth proportional cre…
RSS 官方收录 · 可信分层展示
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
arXiv:2608.…
18324v1 Announce Type: new Abstract: Machine-verifiable workflows prod…
We test whether these records can supervise bounded models, consolidat…
RSS 官方收录 · 可信分层展示
SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
arXiv:2608.…
18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces r…
We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a traini…
RSS 官方收录 · 可信分层展示
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
arXiv:2608.18307v1 Announce Type: new Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests.…
This leaves an under-instrumented middle layer: realistic component-ce…
, toggle a button set) that are short enough to diagnose and rich enou…
RSS 官方收录 · 可信分层展示
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
arXiv:2608.…
18300v1 Announce Type: new Abstract: LLM-as-a-Judge, which leverages a…
However, most work treats a judge as a static artifact, evaluating it …
RSS 官方收录 · 可信分层展示
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
arXiv:2608.…
18289v1 Announce Type: new Abstract: The extraction of structured info…
While proprietary solutions dominate commercial applications, a rapidl…
RSS 官方收录 · 可信分层展示
Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
arXiv:2608.…
18261v1 Announce Type: new Abstract: Serving a 235B-parameter Mixture-…
We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measur…
RSS 官方收录 · 可信分层展示
Redakto - The Incognito Tab for LLMs
arXiv:2608.18260v1 Announce Type: new Abstract: Large Language Models (LLMs) are being increasingly used in everyday applications.…
A major challenge in the context of LLMs or Artificial Intelligence (A…
These challenges have become more urgent with novel EU legislation.
RSS 官方收录 · 可信分层展示
GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks
arXiv:2608.…
18238v1 Announce Type: new Abstract: Genomic analysis on viruses such …
These approaches use pairwise codon or nucleotide distance matrices to…
RSS 官方收录 · 可信分层展示
On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices
arXiv:2608.18194v1 Announce Type: new Abstract: This paper presents new theoretical results on generalizing the Jaccard distance for lattices and real valuations.…
We demonstrate that when the valuation is strictly positive, monotone,…
Moving to relatively complemented distributive lattices (which safely …
RSS 官方收录 · 可信分层展示
Looped Language Models Improve Compositional Tool Calling
arXiv:2608.…
18171v1 Announce Type: new Abstract: Looped language models have shown…
We study this question in compositional tool-calling settings, where m…
RSS 官方收录 · 可信分层展示
Adversarial Review: Structured Disagreement for Grounded Agentic Code Review
arXiv:2608.…
18167v1 Announce Type: new Abstract: Early multi-agent LLM systems oft…
Recent alternatives treat agents as passive tools (subagents), yet thi…
RSS 官方收录 · 可信分层展示
RDFdL: Integrating RDF with Differential Dynamic Logic
arXiv:2608.…
18165v1 Announce Type: new Abstract: Knowledge graphs modeled in RDF a…
, systems described by differential equations, which is a critical gap…
RSS 官方收录 · 可信分层展示
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu
arXiv:2608.…
18142v1 Announce Type: new Abstract: It is challenging to detect hate …
A good example of such a challenge is Roman Urdu which is broadly used…
RSS 官方收录 · 可信分层展示
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
arXiv:2608.…
18136v1 Announce Type: new Abstract: Conversational agents now act for…
Banking is the clearest case: the same agent that answers a question c…
RSS 官方收录 · 可信分层展示
Improving Rural Medication Safety with AI: A Scoping Review
arXiv:2608.18135v1 Announce Type: new Abstract: Introduction: Medication errors (MEs) represent a significant threat to global healthcare systems, contributing to patient harm.…
Introducing artificial intelligence (AI) in rural healthcare enhances …
The aim is to explore the applications and effectiveness of AI technol…
RSS 官方收录 · 可信分层展示
Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis
arXiv:2608.18133v1 Announce Type: new Abstract: Reliable transformer fault diagnosis is essential for maintaining power system stability.…
The IEEE Key Gas Method (KGM), a widely utilized approach in Dissolved…
This study presents An enhanced model combining Fuzzy Logic with the I…
RSS 官方收录 · 可信分层展示
Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
arXiv:2608.…
18117v1 Announce Type: new Abstract: This position paper argues that A…
The barrier is not missing data; high-quality regional benchmarks alre…
RSS 官方收录 · 可信分层展示
Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs
arXiv:2608.18131v1 Announce Type: new Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric.…
When such safety filters fail for non-English languages, the consequen…
For spoken language technologies deployed across India's linguisticall…
RSS 官方收录 · 可信分层展示
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry
arXiv:2608.…
18111v1 Announce Type: new Abstract: Foundation models such as GPT and…
Yet solving a geometry problem and drawing the figure it depends on ar…
RSS 官方收录 · 可信分层展示
下滑 · j/k · m/u · i 信流 · a 稍后 · o 原文 · e 详情 · t 今日 · f 搜索