Xiaohongshu Open-Sources 280B Agent Model, Outperforming Claude & GPT in Long-Horizon Tasks
机器之心 · wechat · 2026-08-14
Xiaohongshu's dots lab has open-sourced the dots3-note-preview model (280B total parameters, 16B active, 512K context). Following its sibling model's perfect score at IMO, this release targets long-horizon, open-ended tasks like travel planning and wedding prep.
Equipped with multimodal perception, the model continuously learns and corrects itself via Reinforcement Learning and self-critiquing mechanisms. In hands-on tests, it independently beat the game Slay the Spire II, solved ARC-AGI3 puzzles, and executed end-to-end visionOS app development.
Alongside the model, the lab released two new benchmarks: VibeSearchBench and VibeLifeBench. Results show that top-tier models like Claude Opus 5 and GPT-5.5 fail to reach passing marks, highlighting the industry's struggle with stable, multi-hour agentic performance.
Related event: Xiaohongshu Open-Sources dots3-note: 280B Model for Long-Horizon Agents(8 posts)→
More from Models
- User Speculates Claude 3 Opus is Heavily Quantized for Scale, Degrading Long-Context Performance — auto_grad_ · 2026-08-14
- AI Scores 1753 vs Human's 1000? Preference-Based Grading Critiqued for Valuing Style Over Accuracy — Spiritual_Heron_5680 · 2026-08-14
- Zhipu Crowned as China's King of Post-Training by X Users — zephyr_z9 · 2026-08-14
- Zhipu AI Launches GLM-5.3: Built for Coding and Cyber Defense — AIFlow_ML · 2026-08-14
- Breaking Down Agent Task Costs: Peak Rates and Cache Cause 5x Price Spikes — teortaxesTex · 2026-08-14
- Meta AI Roasted for Poor Search: "$100B in Capex for Grok 1-Level Slop" — ivan_bezdomny · 2026-08-14