Fable 5 Hand-Writes CUDA for 18.7x Speedup; Anthropic Co-Founder Declares RSI Loop Begun

新智元 · wechat · 2026-07-07

In the new GPU operator benchmark KernelBench-Mega, Fable 5 autonomously hand-wrote CUDA on an RTX PRO 6000, achieving an 18.7x speedup and leaving the runner-up Claude Opus 4.8 (14.4x) and GPT-5.5 (4.34x) far behind. Counterintuitively, longer contexts yielded faster speeds, hitting 19.5x at 16K context.

Fable 5 produced the first true "megakernel" in KernelBench-Mega history: compressing int4 dequantization, attention computation, and MoE routing into a single kernel launch, whereas all other models require 4 to 14 separate launches. During the process, the model spent 64% of its time on benchmark measurements and roofline analysis before writing the kernel in one go, immediately rolling back upon detecting negative optimization. The entire process took about 2.5 hours and 550,000 tokens.

Anthropic co-founder Jack Clark characterized this in his newsletter as the "official start of a Recursive Self-Improvement (RSI) loop," calling it a landmark beginning for AI autonomously optimizing the underlying tools of AI R&D. Notably, Fable 5 is merely a "safe version" of Anthropic's internal model, Claude Mythos.

Related event: Fable 5 Tops KernelBench, Hand-Writes CUDA for 18.7x Speedup(2 posts)→

Original post →

More from Models

Models channel →