Open-weights Kolibri-1 plays Breakout with no fine-tuning at ~25ms per move
Nils_Reimers · x · 2026-10-06
A developer got Aleph Alpha's open-weights Kolibri-1 to play Atari Breakout autonomously, with no fine-tuning and no generated text.
- The model reads ball/paddle coordinates and scores four candidate actions; the highest-probability label becomes the move.
- Inference uses vLLM's pooling/classification path: classifierfromtoken with method=nopostprocessing copies LM output weights into a classification head, yielding 25ms per decision.
- Input is emulator coordinates, not pixels; paddle-following strategy lives in the prompt.
- Tesseracted Labs hosts a free shared arcade page where anyone can watch live decisions, with full prompts and outputs visible.
Related event: Open-source Kolibri-1 Plays Breakout Zero-shot in 25ms per Decision(2 posts)→
More from coding & agent
- Cloudflare launches cf, an agent-first CLI to query observability data via the API — dinasaur_404 · 2026-10-06
- celld: Self-Hosted Cloudflare Workers-Style Runtime With a Cell per AI Agent — letandrewcook · 2026-10-06
- 9 Prompt Rules Cut Agent Thinking Up to 29% With Zero Task Loss, Across 664 Runs — PilgrimofHaqq2 · 2026-10-06
- Low Effort Can Burn More Tokens: Qwen3.8-27B-pi Fine-Tunes Coding Agent Effort Ordering — lmoroney · 2026-10-06
- Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options — DrainBramage · 2026-10-06
- Yacine's 90-Minute Chat With a1zhang: Why Harnesses Boost LLM Generalization — yacinelearning · 2026-10-06