Crowdsourced harness x model benchmark: 3090 beats 5090 with qwen3.8-flash-next config
dh7net · reddit · 2026-10-08
The author has been testing harness x model combinations for local coding setups and built airbench.ai so anyone can benchmark their config and share results.
- The current leader is a stranger's setup: qwen3.8-flash-next-iq3s via pi and Strata, on a 3090 24GB with 80GB system RAM and context raised to 256 — beating the author's own 5090 baseline
- The pooled data reveals the best harness per model (chart in post)
- A new "contributor" section was added; login required to submit, but all results remain publicly viewable
A useful reference for local-deploy enthusiasts: the same model performs very differently across harnesses, and crowdsourced numbers beat solo benchmarks.
More from coding & agent
- One Prompt Chain Makes Claude Run a Full Product Launch: Positioning to 7-Day Content — thisguyknowsai · 2026-10-08
- Kolibri-1 in a Real Agent Test: 78B MoE With 3.46B Active Params and 1M Context Falls Short? — WolframRvnwlf · 2026-10-08
- Agent Efficiency Is Token Efficiency: Good Test Suites Cut Costly Rework, Say XState Author — DavidKPiano · 2026-10-08
- 10 Legal Traps in Vibe-Coded Apps: 170 Lovable Builds Found Leaking Data — alex_verem · 2026-10-08
- Microsoft extends GitHub Spec Kit: presets, extensions and bundles for enterprise SDD — WirelessLife · 2026-10-08
- After an agent leaked salary data, Reddit distilled a 5-step pre-launch access checklist — Wild-Lawyer8511 · 2026-10-08