20B Looping paper says it matches Qwen3 Coder 30B with 10% of pretraining tokens
Dany0 · reddit · 2026-07-21
A paper on a 20B Looping model claims it matches or beats Qwen3 Coder 30B while using only 10% of the pre-training tokens.
- The post says the run may have cost only hundreds of thousands of dollars.
- The author notes it did not beat GPT-OSS 20B on every dimension, but the efficiency jump is the real headline.
- The bigger implication: training a true LLM from scratch on 3.5T tokens instead of 35T could eventually make home-scale pretraining more realistic.
Related event: Loopie Cyclic Transformer Matches 30B Baselines with Fractional Tokens(7 posts)→
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22