Xiaohongshu Open-Sources 280B Agent Model, Outperforming Claude & GPT in Long-Horizon Tasks

机器之心 · wechat · 2026-08-14

Xiaohongshu's dots lab has open-sourced the dots3-note-preview model (280B total parameters, 16B active, 512K context). Following its sibling model's perfect score at IMO, this release targets long-horizon, open-ended tasks like travel planning and wedding prep.

Equipped with multimodal perception, the model continuously learns and corrects itself via Reinforcement Learning and self-critiquing mechanisms. In hands-on tests, it independently beat the game Slay the Spire II, solved ARC-AGI3 puzzles, and executed end-to-end visionOS app development.

Alongside the model, the lab released two new benchmarks: VibeSearchBench and VibeLifeBench. Results show that top-tier models like Claude Opus 5 and GPT-5.5 fail to reach passing marks, highlighting the industry's struggle with stable, multi-hour agentic performance.

Related event: Xiaohongshu Open-Sources dots3-note: 280B Model for Long-Horizon Agents(8 posts)→

Original post →

More from Models

Models channel →