Xiaohongshu Open-Sources 280B Agent Model, Outperforming Claude & GPT in Long-Horizon Tasks
机器之心 · wechat · 2026-08-14
Xiaohongshu's dots lab has open-sourced the dots3-note-preview model (280B total parameters, 16B active, 512K context). Following its sibling model's perfect score at IMO, this release targets long-horizon, open-ended tasks like travel planning and wedding prep.
Equipped with multimodal perception, the model continuously learns and corrects itself via Reinforcement Learning and self-critiquing mechanisms. In hands-on tests, it independently beat the game Slay the Spire II, solved ARC-AGI3 puzzles, and executed end-to-end visionOS app development.
Alongside the model, the lab released two new benchmarks: VibeSearchBench and VibeLifeBench. Results show that top-tier models like Claude Opus 5 and GPT-5.5 fail to reach passing marks, highlighting the industry's struggle with stable, multi-hour agentic performance.
Related event: Xiaohongshu Open-Sources dots3-note 280B Multimodal Model(14 posts)→
More from Models
- From botching 9.9 vs 9.11 to tackling the hardest math problems in two years — Yuchenj_UW · 2026-10-07
- OpenAI claims 372 unsolved problems cracked, averaging about 3 hours each — i_dg23 · 2026-10-07
- Benchmark: OpenAI Decisions API costs 2x more, 5-10% worse than Jev — xeophon · 2026-10-07
- Cagliostro V3.5 135M Dethrones SmolLM2-135M on Open SLM Leaderboard at 27.49 — Megneous · 2026-10-07
- "If you still think LLMs suck at math, it's a skill issue" — hot take gains traction — alejandroll10 · 2026-10-07
- Sauers_: Opus 5.5 shows fewer strong technical opinions than Fable 5.1 and 5 — Sauers_ · 2026-10-07