Local LLM Eval: Snake Game Generation on NVIDIA GB10

niacolhealth · reddit · 2026-09-02

The author discusses the evaluation of local models beyond simple throughput benchmarks. Focusing on a recorded run on an NVIDIA GB10 device, the analysis shows a llama-server workflow generating 2,429 tokens in 70.57 seconds (34.42 tok/s) to create a functional Snake game in a single HTML file. While not a comprehensive benchmark for long-context stability, this approach provides an inspectable end-to-end chain from prompt to a running artifact. The post asks the community for examples of small tasks that revealed issues missed by standard throughput benchmarks.

Original post →

More from Infra

Infra channel →