Ornith-1.5-35B-A3B runs agentic coding at 32 tok/s on an 8GB RTX 3070 laptop
Elemental_Particle · reddit · 2026-08-28
After previously favoring Qwen3.6-35B-A3B for 8GB VRAM, the author tested Ornith-1.5-35B-A3B Q4KM following community suggestions and found a new winner.
- Setup: i7-11800H + RTX 3070 laptop (8GB) + 64GB RAM, openSUSE, Unsloth Studio, Pi.dev
- On a fairly complex code-analysis project it averaged 32 tok/s and delivered a near-flawless result
- Keeps a 128K context window; the biggest gain over Qwen3.6 is how well it handles the whole agentic workflow
- Having tested many dense/MoE models, quantizations and coding fine-tunes, the author calls Ornith-1.5 the clear winner for this hardware today — while joking a new release tomorrow could change that
More from Infra
- Australia Minister: No Fossil Fuel Carve-out for Datacenters — nordicinst · 2026-08-28
- ZED Camera priced at $500? DIY alternative costs just $150 — _William_F_ · 2026-08-28
- KOTOR Remaster Path Tracer Integrates DLSS 4.5 RR — Michael_Moroz_ · 2026-08-28
- NVIDIA's NVHBM Breaks the Die-Size Limit, 3-5x VRAM per GPU — Charuru · 2026-08-28
- Tutorial: Train a Raspberry Pi to Read Gas Meter Automatically with Neural Network — JeremyCMorgan · 2026-08-28
- Ninfer Benchmark: 5090 Doubles Throughput for Qwen3 27B — Rollingsound514 · 2026-08-28