Four LLMs Play Doom: Jev Averages 5.63 Kills, 4.5x a Finetuned Qwen3.5-4B

shniydder · reddit · 2026-09-20

The author wired Jev, Laya, a finetuned ModernCE-base-nli, and a LoRA-tuned Qwen3.5-4B to play the same-seeded ViZDoom scenarios, comparing decision quality and latency.

Takeaway: Jev clearly makes better decisions but pays a latency cost; tiny local models are fast yet mediocre at gameplay. Full methodology and a longer write-up are linked in the post.

Related event: Four LLMs play Doom: Jev leads with 5.63 kills, 15x latency gap(2 posts)→

Original post →

More from Models

Models channel →