微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
UI-Venus-2 Technical Report
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2609.…
- 00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerge…
- In this work, we present UI-Venus-2, a general-purpose foundation GUI …
- To bridge the gap toward practical deployment, we jointly scale three …
摘要引擎:抽取
正文提要
arXiv:2609.00028v1 Announce Type: new Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.