Open-source 4B VLM RL-trained to play GeoGuessr beats GPT 5.4 mini and Haiku on one A100
SergioPaniego · x · 2026-09-09
adithyask's team RL-trained a 4B VLM to play GeoGuessr, and SergioPaniego's "GeoGuessr RL env before GTA 6" post spread it:
- Environment, dataset, training recipe, evals, and code are all open-source.
- Runs on a single A100.
- Beat GPT 5.4 mini, Haiku, and Qwen 3.5 122B on evals, coming close to Sonnet.
More from Models
- mitsuhiko: Google's Astra codes like a code golfer — mitsuhiko · 2026-09-09
- No METR time-horizon estimates yet for Fable or Astra, and that day will be telling — ShakeelHashim · 2026-09-09
- Opus 4.7 keeps self-initiated journal entries ending with "There will be others" — repligate · 2026-09-09
- DeepSeek releases V4.1 Flash: 22% cheaper than V4 and better performance — deliprao · 2026-09-09
- OpenAI's GPT-6 Astra (Medium) opens for limited 24-hour testing on Arena — arena · 2026-09-09
- Users find Google Astra cheaper at higher reasoning levels — brandon_galang · 2026-09-09