Local models face a single-shot HTML flight simulator test across six runs

JLeonsarmiento · reddit · 2026-07-22

A Reddit user benchmarks several local models on a single-shot coding task: generate a beautiful, relaxing flight simulator in one HTML file with mountains, clouds, and endless procedural terrain.

The post compares Qwen3.6-27B, Qwen3.6-MoE, Ornith-35B, Gemma-4-26B, HuiHui-Qwen3.6-MoE, and Agents-A1 under controlled inference settings. The key point is not just raw output, but whether the HTML works at all: if it fails, the run is reset and retried up to three times. The setup uses Pi as the harness and oMLX for serving, with the author linking to their local SOTA quants collection for 48GB Macs.

Original post →

More from coding & agent

coding & agent channel →