Local models face a single-shot HTML flight simulator test across six runs
JLeonsarmiento · reddit · 2026-07-22
A Reddit user benchmarks several local models on a single-shot coding task: generate a beautiful, relaxing flight simulator in one HTML file with mountains, clouds, and endless procedural terrain.
The post compares Qwen3.6-27B, Qwen3.6-MoE, Ornith-35B, Gemma-4-26B, HuiHui-Qwen3.6-MoE, and Agents-A1 under controlled inference settings. The key point is not just raw output, but whether the HTML works at all: if it fails, the run is reset and retried up to three times. The setup uses Pi as the harness and oMLX for serving, with the author linking to their local SOTA quants collection for 48GB Macs.
More from coding & agent
- Two papers use LLMs to improve retrieval indexing and grounded answers — _reachsumit · 2026-07-22
- Grok Build turns one prompt into a full ARPG with AI-generated game assets — tetsuoai · 2026-07-22
- A dad built a controller-ready game in two hours with Grok 4.5 — minchoi · 2026-07-22
- Atomic-Chat pitches a fully offline open-source ChatGPT alternative — rohanpaul_ai · 2026-07-22
- ComfyUI Wan dance test renders a 30-second clip in 70 minutes — tostane · 2026-07-22
- Claude Code adds iOS Simulator control for side-by-side mobile testing — xiaohu · 2026-07-22