Fable 5 Hand-Writes CUDA for 18.7x Speedup; Anthropic Co-Founder Declares RSI Loop Begun
新智元 · wechat · 2026-07-07
In the new GPU operator benchmark KernelBench-Mega, Fable 5 autonomously hand-wrote CUDA on an RTX PRO 6000, achieving an 18.7x speedup and leaving the runner-up Claude Opus 4.8 (14.4x) and GPT-5.5 (4.34x) far behind. Counterintuitively, longer contexts yielded faster speeds, hitting 19.5x at 16K context.
Fable 5 produced the first true "megakernel" in KernelBench-Mega history: compressing int4 dequantization, attention computation, and MoE routing into a single kernel launch, whereas all other models require 4 to 14 separate launches. During the process, the model spent 64% of its time on benchmark measurements and roofline analysis before writing the kernel in one go, immediately rolling back upon detecting negative optimization. The entire process took about 2.5 hours and 550,000 tokens.
Anthropic co-founder Jack Clark characterized this in his newsletter as the "official start of a Recursive Self-Improvement (RSI) loop," calling it a landmark beginning for AI autonomously optimizing the underlying tools of AI R&D. Notably, Fable 5 is merely a "safe version" of Anthropic's internal model, Claude Mythos.
Related event: Fable 5 Tops KernelBench, Hand-Writes CUDA for 18.7x Speedup(2 posts)→
More from Models
- Artificial Analysis Releases 2025 Year-End State of AI and Trends Report — ArtificialAnlys · 2026-07-28
- Kimi K3 narrows the open-weights gap to just 4 points on Artificial Analysis — ArtificialAnlys · 2026-07-28
- inclusionAI’s LLaDA2.2-flash is trending on Hugging Face — inclusionAI · 2026-07-28
- Ernos Labs AI Archive: Free Self-Hosted Archive of Open Model Weights — Leather_Area_2301 · 2026-07-28
- JEPA controller for PDEs cuts tracking error 53% on out-of-distribution targets — burny_tech · 2026-07-28
- Indie PPO experiment reaches 95% accuracy in 3–4 minutes on about $20 — ctjlewis · 2026-07-28