Banbury Road's Kardashev-0.7 trains 32 distinct models together with reinforcement learning
MarceloDeAviz · reddit · 2026-10-06
Banbury Road announced Kardashev-0.7, a population of 32 distinct models trained jointly with reinforcement learning for "Population Scaling" — letting models learn complementary specializations so that model count and collaboration become a new scaling axis. An open evaluation question: how much does joint population training improve over an ensemble of independently trained models at the same total compute?
Related event: Kardashev-0.7 Trains a Swarm of 32 Models with RL(2 posts)→
More from Research
- Cambridge team's new work on pullback geometry for data on mixtures of manifolds — skoularidou · 2026-10-06
- CMU L3 Lab brings GradAlign RL data selection and sim2real papers to COLM 2026 — wellecks · 2026-10-06
- A statistical framework for LLM watermarks: optimal detection rules via hypothesis testing — weijie444 · 2026-10-06
- 4DCodeBench: new benchmark tests coding agents on reconstructing dynamic 3D scenes from video — CSProfKGD · 2026-10-06
- EMNLP paper: LLMs learn novel tasks more reliably from rules than from in-context examples — najoungkim · 2026-10-06
- New COLM study: Looped Transformers do implicit reasoning over parametric knowledge, boosting generalization — hhsun1 · 2026-10-06