Test: Same Local Model Beats GPT 5.6 After Switching to Codex CLI — Harness Matters

L0ren_B · reddit · 2026-09-27

A user running both a local LLM and an OpenAI subscription shares a hands-on test: he runs Qwen 3.8 Flash Next at high speed on 2x3090 + RAM, but local models previously only worked for demos like 'build a 3D Mario game' and never matched GPT 5.2 or 5.6 Luna on serious work.

That changed when he had GPT 5.6 Luna configure Codex CLI for the local model (he had been using pi.dev and opencode). In Codex CLI the 3D Mario test was the best yet; on a real project Luna had worked on for days, Qwen 3.8 Flash Next ran circles around Luna in a parallel comparison — whereas Qwen and DeepSeek had previously failed to deliver on that same project.

He also ported pi-smart-web-search into Codex as a skill with great results, and wonders whether he was simply using pi.dev wrong or if an extension can bring it to Codex CLI quality.

Original post →

More from coding & agent

coding & agent channel →