Desktop iGPU beats dual 3090s due to a Jinja template
TrifleHopeful5418 · reddit · 2026-08-25
A comparison on LiveCodeBench v6 shows Ornith-1.5-35B-A3B running on a Strix Halo iGPU nearly matches Qwen3.8-27B on dual 3090s. The key finding is that adding a LoRA and switching to a Sharp Chat Template improved scores by 15 problems, a gain larger than adding a second GPU. The author also warns that the default thinking-mode profile in GGUF failed all tests, recommending the Instruct profile instead.
More from Models
- GLM-5.3 weights will be released tomorrow — serige · 2026-08-27
- TokenSpeed adds Day-0 support for Qwen 3.8 Flash Next architecture — Alibaba_Qwen · 2026-08-27
- AI models show more creativity when talking to each other than in assistant persona — nabeelqu · 2026-08-27
- Claude Code shrinks by 45% in next version — jarredsumner · 2026-08-27
- OpenRouter leaderboard: Real token consumption data outweighs media hype — sujingshen · 2026-08-27
- Qwen 3.8-Next Released with Detailed Technical Report on Architecture — nrehiew_ · 2026-08-27