Four LLMs Play Doom: Jev Averages 5.63 Kills, 4.5x a Finetuned Qwen3.5-4B
shniydder · reddit · 2026-09-20
The author wired Jev, Laya, a finetuned ModernCE-base-nli, and a LoRA-tuned Qwen3.5-4B to play the same-seeded ViZDoom scenarios, comparing decision quality and latency.
- Input: a deterministic Python adapter converts ViZDoom's object labels, bounding boxes, and HUD values into short text; output is one button action at 5 decisions/second, with the game clock not pausing.
- Local models ran on a single DGX Spark (GB10, 128GB unified memory); Jev used TypeSafe's hosted API. No model had Doom-specific training.
- Averages over 8 seeds: Jev scored 5.63 mean kills in Defend the Center, well ahead of Qwen3.5-4B (3.63), while Laya and ModernCE each managed 1.25. Survival times in Health Gathering were close (13.0s vs 11.3s).
- Latency: Laya p50 16ms and ModernCE 8ms locally, versus 117ms for Jev (API round trip) and 147ms for Qwen.
Takeaway: Jev clearly makes better decisions but pays a latency cost; tiny local models are fast yet mediocre at gameplay. Full methodology and a longer write-up are linked in the post.
Related event: Four LLMs play Doom: Jev leads with 5.63 kills, 15x latency gap(2 posts)→
More from Models
- Stepfun's Step-5-Preview-BF16 weights quietly appear on Hugging Face, hinting at next-gen release — adefa · 2026-09-20
- User burns 100M tokens on Jev without exhausting the $5 free credit — DeryaTR_ · 2026-09-20
- Users slam Codex 'reset culture': paid subscribers hoard usage waiting for free resets — TheMoonMidas · 2026-09-20
- Andrew Chen: 400x cheaper inference could unlock ad-supported free AI-native apps — andrewchen · 2026-09-20
- Bilibili's AI Infinite Arena: 190 videos stress-testing 100+ models on real tasks — vista8 · 2026-09-20
- Users report suspiciously generous usage: full morning of building leaves 78% quota left — TheMoonMidas · 2026-09-20