Fable 5 Hand-Writes CUDA for 18.7x Speedup; Anthropic Co-Founder Declares RSI Loop Begun
新智元 · wechat · 2026-07-07
In the new GPU operator benchmark KernelBench-Mega, Fable 5 autonomously hand-wrote CUDA on an RTX PRO 6000, achieving an 18.7x speedup and leaving the runner-up Claude Opus 4.8 (14.4x) and GPT-5.5 (4.34x) far behind. Counterintuitively, longer contexts yielded faster speeds, hitting 19.5x at 16K context.
Fable 5 produced the first true "megakernel" in KernelBench-Mega history: compressing int4 dequantization, attention computation, and MoE routing into a single kernel launch, whereas all other models require 4 to 14 separate launches. During the process, the model spent 64% of its time on benchmark measurements and roofline analysis before writing the kernel in one go, immediately rolling back upon detecting negative optimization. The entire process took about 2.5 hours and 550,000 tokens.
Anthropic co-founder Jack Clark characterized this in his newsletter as the "official start of a Recursive Self-Improvement (RSI) loop," calling it a landmark beginning for AI autonomously optimizing the underlying tools of AI R&D. Notably, Fable 5 is merely a "safe version" of Anthropic's internal model, Claude Mythos.
Related event: Fable 5 Tops KernelBench, Hand-Writes CUDA for 18.7x Speedup(2 posts)→
More from Models
- Unreleased 'GPT 6 Sol' model spotted in OpenAI's API — SteveEricJordan · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11