Six decision models play Pac-Man: jev 1.13 tops leaderboard at 2,750 avg score

facethef · reddit · 2026-10-08

A Reddit user benchmarked six popular decision models by having them play Pac-Man in real time (all respond within ms): jev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya.

Leaderboard (100 runs per model, mean score with ±2 standard errors):

| Model | Avg score | High score | Avg latency |

|---|---|---|---|

| jev 1.13 | 2,750 | 6,380 | 290 ms |

| GPT-6 Luna | 2,568 | 5,920 | 179 ms |

| Clef Flash | 2,538 | 4,260 | 256 ms |

| Clef | 2,476 | 4,820 | 398 ms |

| Kev 4B | 1,506 | 5,280 | 231 ms |

| Laya | 639 | 1,200 | 104 ms |

jev 1.13 leads on average score, though the top models statistically tie; Laya is fastest but lowest-scoring. The repo is open-source so anyone can plug in their own fine-tuned local or hosted decision model and join the leaderboard. You can also play Pac-Man yourself with the models controlling the ghosts, mixed or uniform.

Original post →

More from Models

Models channel →