Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use Paper • 2607.01084 • Published Jul 1
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published Jun 28 • 24
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published Jun 28 • 24
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Paper • 2605.27284 • Published May 26 • 9
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Paper • 2510.10396 • Published Oct 12, 2025