Apple's MintAct unifies GUI agents across mobile, desktop, and web, hitting SOTA 48.9 on OSWorld-Verified
apple · hf · 2026-09-21
Apple introduced MintAct, a family of vision-language models (2B, 4B, 8B) that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use — matching per-domain specialists across all capabilities.
Infrastructure: Hundreds of concurrent environment instances on heterogeneous backends serve both trajectory collection and online RL; an asynchronous RL framework gives explicit control over cross-domain training distribution and stays stable under noisy feedback and off-policy drift.
Results: State-of-the-art 48.9 on OSWorld-Verified among comparable model sizes.
More from Embodied
- DeformSmith generates physics-grounded deformable assets for robot manipulation — Can Li · 2026-09-21
- RewardAI launches OM-1 robot foundation model trained on human data alone, zero-shot across arms and humanoids — zipengfu · 2026-09-21
- Tsinghua's EMERGE-Policy dispatches role-based sub-agents to combine VLA, world models and verifiers — jiqizhixin · 2026-09-21
- First human vs humanoid robot T800 fight opens a new entertainment genre — CyberRobooo · 2026-09-21
- Meta to launch camera-less smart glasses at Connect as privacy backlash grows — Scobleizer · 2026-09-21
- Movement Trend Guidance lifts 3D diffusion policies without explicit trajectories — 72% vs 49% on real robots — dalian-university-of-technology · 2026-09-21