Kardashev-0.7: a 32-model RL-trained swarm claims frontier performance at 0.7-2% inference cost
Scobleizer · x · 2026-10-06
Kardashev-0.7 is billed as the world's first trained swarm composed of 32 distinct models, trained with a method called RLPS (Reinforcement Learning for Population Scaling).
Key points:
- The 32 models organically develop specialization and complementary capabilities during training, mirroring how civilizations advance through specialized minds working together
- Claims frontier-level performance at 0.007x-0.02x of the inference cost and 0.03x of the required memory
- The team's stated direction is scaling intelligence by model count — building 'civilizations of models' that learn to build on one another
Related event: Kardashev-0.7 Trains a Swarm of 32 Models with RL(2 posts)→
More from Models
- SelfBench turns real GitHub PRs into evals: open-weight models cost more and do worse — ycombinator · 2026-10-06
- Embedded Grok on X reportedly lacks per-user context isolation, called out as a major flaw — altryne · 2026-10-06
- Liquid AI's d1 decision model adds vision, beats GPT-6.1 on 4 of 6 tasks at up to 200x lower cost — JosephJacks_ · 2026-10-06
- Anthropic reviewers alerted police to a Claude chat threatening a sheriff's office, leading to an arrest — rohanpaul_ai · 2026-10-06
- flow-1: RL-trained model matches GPT-6-sol at trace debugging while 23x cheaper — kalyan_kpl · 2026-10-06
- First large-scale 3B/8B continuous diffusion LMs match pass@1 and beat pass@k vs masked dLMs — ArashVahdat · 2026-10-06