Six decision models play Pac-Man: jev 1.13 tops leaderboard at 2,750 avg score
facethef · reddit · 2026-10-08
A Reddit user benchmarked six popular decision models by having them play Pac-Man in real time (all respond within ms): jev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya.
Leaderboard (100 runs per model, mean score with ±2 standard errors):
| Model | Avg score | High score | Avg latency |
|---|---|---|---|
| jev 1.13 | 2,750 | 6,380 | 290 ms |
| GPT-6 Luna | 2,568 | 5,920 | 179 ms |
| Clef Flash | 2,538 | 4,260 | 256 ms |
| Clef | 2,476 | 4,820 | 398 ms |
| Kev 4B | 1,506 | 5,280 | 231 ms |
| Laya | 639 | 1,200 | 104 ms |
jev 1.13 leads on average score, though the top models statistically tie; Laya is fastest but lowest-scoring. The repo is open-source so anyone can plug in their own fine-tuned local or hosted decision model and join the leaderboard. You can also play Pac-Man yourself with the models controlling the ghosts, mixed or uniform.
More from Models
- Leak: Grok Voice Mode coming to X — talk to Grok out loud in the app — nima_owji · 2026-10-09
- Subscription Claude models deliver far fewer thinking tokens, measured five ways — _AustinCalvert_ · 2026-10-09
- GPT-6 in ChatGPT is the best AI writer yet — barely needs editing, says user — VraserX · 2026-10-09
- Reddit User Finds GPT-6 Appears to Lack Direct Access to Saved Memories — ItsAGarbageAccount · 2026-10-09
- TypeSafe's Text-Free Model Jev Matches GPT-6 Astra Accuracy at ~1/500th the Cost — JenniferHli · 2026-10-09
- Over 7% of Claude's US consumer subs pay $100+/month vs ~1% for ChatGPT — omooretweets · 2026-10-09