One model plays many games in parallel from screenshots alone — text input adds nothing

maximelabonne · x · 2026-10-06

Maxime Labonne showed a single model playing multiple games (like Wordle) in parallel, with a key finding: feeding only screenshots works because the model's built-in OCR reliably extracts the screen state and makes correct decisions — adding text input doesn't help. He argues this goes far beyond text generation with LLMs, with much bigger implications for the gaming industry.

Related event: Labonne demos one model playing multiple games in parallel from screenshots alone(5 posts)→

Original post →

More from Multimodal

Multimodal channel →