微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开 ,或复制链接后用系统浏览器访问。
综合
官方
企业
汇聚
1 / 24
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).…
Yet agreement in final labels does not show that human annotators and …
Two agents may reach the same judgment while appealing to different pr…
We test this distinction using a curated 500-item ETHICS-derived bench…
RSS 官方收录 · 可信分层展示
Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
arXiv:2608.…
12371v1 Announce Type: new Abstract: Stream-processing systems increas…
This paper proposes \emph{MAS-DecStream}, whose main contribution is \…
Edge-cluster agents refine natural-language offloading proposals from …
RSS 官方收录 · 可信分层展示
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
arXiv:2608.…
12346v1 Announce Type: new Abstract: This position paper argues that m…
By mapping current alignment techniques to the possibility and actual …
We need to discuss this dual-use potential now, as its risk is exacerb…
RSS 官方收录 · 可信分层展示
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
arXiv:2608.…
12345v1 Announce Type: new Abstract: Language models are increasingly …
We introduce IntegrityBench, a benchmark evaluating misconduct classif…
Evaluating 18 frontier model variants, we find that under peak pressur…
RSS 官方收录 · 可信分层展示
Position: Reasoning is a Learnable Rule-Based Process
arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today.…
Historically the purview of symbolic AI, recent advances have mainly e…
Despite immense interest and rapid progress, the generative AI communi…
This position contends that definitional ambiguity leaves the construc…
RSS 官方收录 · 可信分层展示
限量手办 + 实景体验,浙江人行NAVIAI2026WRC 福利提前曝光
8 月 19 日至 23 日,2026 世界机器人大会(WRC2026)将在北京亦庄亦创国际会展中心举办。浙江人形机器人创新中心将携旗下 NAVIAI 人形机器人亮相 C-105 展位,集中展示人形机器人在多元场景的落地应用成果。…
本次展会上,NAVIAI 将呈现工业制造、智慧零售、家庭服务、遥操数采多场景的实操能力,现场演示拆垛分拣搬运、货品递送、烹饪清洁、远程数据采…
另有文娱演绎类场景在 8 月 19 日开展当天作为特别环节限时展示。
大会期间,浙江人形机器人创新中心有限公司首席科学家熊蓉教授将于 8 月 21 日站上主论坛,分享 NAVIAI 的技术演进与未来落地规划,多…
RSS 官方收录 · 可信分层展示
智谱发布 GLM-5.3,编程能力更强;传苹果训练国内专用 AI 模型;微信:朋友圈现在、过去、未来都不会有二次编辑功能 | 极客早知道
智谱正式发布 GLM-5.3,拥有更强编程能力 8 月 14 日,据介绍,与 GLM-5.2 相比,GLM-5.3 基座模型未变,但通过极致的后训练 Scaling 大大提高了模型的智能上界。…
3 拥有更强的编程能力,在内部自建体感评测中较 GLM-5.
2 提升 50%,在包括 TerminalBench3.
0、Agents'LastExam(CLI)在内的公开基准测试中取得开源第一。
RSS 官方收录 · 可信分层展示
早报|曝苹果与阿里合作训练AI模型/微信:永不推出朋友圈二次编辑/售价20万,追觅首台手机交付
曝苹果与阿里合作,为中国市场训练自研 AI 模型 微信确认朋友圈永不推出二次编辑功能 售价 20 万元,追觅首台 AURORA 手机交付:24K 足金镶宝石 Google DeepMind 或裁员三分之一以上,资源转向 Flash WorkBuddy 接入 GLM-5.…
3 广州推出「Token 贷」:按算力合同和 Token 消耗额度授信 曝 DeepSeek 正研发情感 AI 模型 调查:美国年轻人普遍不…
3,同一基座靠后训练提升编程与网络安全能力 Ling-3.
0-tiny 与 ASystem AReno 打通单机 Agentic RL 训练闭环 Suno Studio 2.
RSS 官方收录 · 可信分层展示
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
arXiv:2608.…
11403v1 Announce Type: new Abstract: Self-consistency (SC) via majorit…
On the full GPQA Diamond benchmark (198 graduate-level science questio…
6% of problems for Qwen2.
RSS 官方收录 · 可信分层展示
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
arXiv:2608.…
11381v1 Announce Type: new Abstract: We study whether the localized nu…
Larix maps a 16-lens European listed-real-estate analysis framework to…
Across 19 firms spanning seven regulatory wrappers, decomposition impr…
RSS 官方收录 · 可信分层展示
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
arXiv:2608.…
11343v1 Announce Type: new Abstract: Multimodal retrieval and classifi…
The March 2026 release of Gemini Embedding 2, Google's first natively …
Simultaneously, frontier Large language models (LLMs) have also demons…
RSS 官方收录 · 可信分层展示
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
arXiv:2608.…
11354v1 Announce Type: new Abstract: Modern recommender systems treat …
As interfaces evolve from static layouts toward generative UIs and imm…
We propose an Inverse Theory of Mind (IToM) pipeline that reasons back…
RSS 官方收录 · 可信分层展示
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations.…
It succeeded by turning a distant ambition into a mission architecture…
AI now faces a similar transition: frontier models can solve difficult…
We introduce Apodex Discovery, a framework for building and evaluating…
RSS 官方收录 · 可信分层展示
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability.…
We show, across three open agent-trace benchmarks (TheAgentCompany, $\…
Leaderboards rank specialization, not capability.
We arrive at this through a four-facet Generalizability Theory varianc…
RSS 官方收录 · 可信分层展示
Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning
arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals.…
Existing approaches exhibit a "when-what" dissociation: traditional DN…
We attribute this to the absence of a unified reasoning paradigm.
Inspired by how humans inspect surveillance videos - glancing globally…
RSS 官方收录 · 可信分层展示
Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
arXiv:2608.…
11258v1 Announce Type: new Abstract: Gradient injection helps Particle…
We propose Adaptive Hybrid PSO (AHPSO), which uses a sigmoid function …
Under budget-normalized comparison (PSO given equivalent total functio…
RSS 官方收录 · 可信分层展示
Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures
arXiv:2608.…
11255v1 Announce Type: new Abstract: Accurate prediction of vapor--liq…
While deep learning models can provide accurate predictions, they ofte…
In this work, we propose a symbolic machine learning approach to disco…
RSS 官方收录 · 可信分层展示
Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning
arXiv:2608.…
11252v1 Announce Type: new Abstract: Agentic AI systems routinely tran…
We prove this class of safeguard is structurally incomplete.
Modelling a covering of context space by its nerve and evidence by a r…
RSS 官方收录 · 可信分层展示
AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search
arXiv:2608.…
11250v1 Announce Type: new Abstract: Language models can propose many …
We present AgonAlpha, an architecture that searches over frozen resear…
To our knowledge, AgonAlpha is the first alpha-mining system to combin…
RSS 官方收录 · 可信分层展示
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
arXiv:2608.11248v1 Announce Type: new Abstract: Long-term memory is essential for language agents operating across extended interactions and evolving tasks.…
Existing memory-augmented agents mainly focus on storing and retrievin…
In particular, previously distilled insights can become outdated, over…
To address this issue, we study insight-level memory maintenance for l…
RSS 官方收录 · 可信分层展示
Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
arXiv:2608.…
11247v1 Announce Type: new Abstract: Recent advances in language model…
Each agent sees what the others assert before it answers, so peer opin…
We measure that displacement in 23 open-weight models, 19 conditions, …
RSS 官方收录 · 可信分层展示
Towards the Harness of Embodied Agents
arXiv:2608.…
11246v1 Announce Type: new Abstract: The success of coding agents has …
We ask whether the same paradigm extends to embodied agents in the phy…
We present Thea, a harness in which an agentic loop orchestrates robot…
RSS 官方收录 · 可信分层展示
Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
arXiv:2608.…
11245v1 Announce Type: new Abstract: Online education offers unprecede…
To address these challenges, we introduce AI Tutor, a reinforcement le…
In the short term, AI-Tutor draws on cognitive theory to guide learner…
RSS 官方收录 · 可信分层展示
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model
arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability.…
Existing standard application workflows rely on keyword-based document…
To address these limitations, this study develops a multimodal knowled…
The framework introduces 1) a multimodal knowledge graph (MKG) for uni…
RSS 官方收录 · 可信分层展示
上滑下一条
上滑 · j/k · m/u · h 隐藏 · a 稍后 · o 原文 · e 详情 · t 今日 · i 模式 · f 搜索