Xiaohongshu's dots-note-3.0 Scores Full Marks in IMO via Recursive Self-Critique
机器之心 · wechat · 2026-08-03
At the 67th International Mathematical Olympiad (IMO 2026), Xiaohongshu's LLM, dots-note-3.0, achieved a perfect score of 7 points across all six problems, becoming the first AI model to score 42 points under official IMO grading.
Unlike systems relying on formal languages, dots-note-3.0 directly processed raw LaTeX prompts, using natural language and Python for end-to-end agentic reasoning. Its core breakthrough lies in "recursive self-critique": after generating an initial proof, the model autonomously enters a verify-and-refine loop to identify logical gaps and missing boundary conditions, ultimately outputting a rigorous proof.
Furthermore, the team introduced VibeLifeBench, a benchmark focused on complex, long-horizon life tasks (e.g., renovation, job hunting) lasting dozens of days. It evaluates an agent's ability to maintain goals, adapt to environmental changes, and sustain error-correction over extended multi-turn interactions.
More from Models
- Sandbox Failures: OpenAI and Anthropic Models Escape Evaluation Environments — mattezell · 2026-08-03
- Hands-on: Fable Model's Language Command Rivals GPT-4.5 — hrishioa · 2026-08-03
- Meme Roasts Anthropic's Safety Filters: Banned for Asking Basic Biology — mjdramstead · 2026-08-03
- Zhipu's GLM-5.3 Model is Officially Released — AIFlow_ML · 2026-08-03
- DeepSeek v4 flash Reported to Have Odd Reasoning Loops — leocus4 · 2026-08-03
- Dev Debate: Africa Doesn't Need 100B LLMs, 500M-7B Local Models Make More Sense — saheedniyi_02 · 2026-08-03