Ornith 1.5 35B-A3B hits 180 tok/s on dual 5070 Tis, 3x faster than Qwen 27B at same agent scores
Excellent-Issue-5956 · reddit · 2026-10-05
On a dual RTX 5070 Ti 16G setup (second card on an OCuLink dock, 64GB RAM, Ollama on Windows/WSL), the author swapped their local agent's daily driver from Qwen3.8 27B UD-Q4KXL to Ornith 1.5 35B-A3B (Q4KM):
- Self-built agent tests: a 9-step long session (tool calls, file reads, recall after 3 context compactions) and 10 coding tasks — both models score 9/9 and 10/10; Laguna XS 2.1 also passes both, North Mini Code 1.0 gets 4/9
- Speed: 180 tok/s at 128K context (24.4GB, fully on both cards), 3x the 27B's 55-70 tok/s. Only 3B params active per token (8 of 256 experts), just 10 of 41 layers full attention with 2 KV heads, rest linear attention — KV cache barely grows, 128K→256K adds only 2GB
- 256K works but spills 1.2GB to system RAM, dropping to 139 tok/s, so 128K stays default
- Vendor claims SWE-bench Verified 79 and Terminal-Bench 2.1 68.5, unverified by the author; Artificial Analysis hasn't scored it
- Next: trying YaRN factor 4 via llama-server for 1M context
Caveat: tests are maxed out, so this only shows it's not worse than the 27B on this workload.
More from coding & agent
- Dev builds AI skill cloning Matt Levine's writing style, won't release it over consent concerns — morqon · 2026-10-05
- DHH: Every developer needs an 'AI shed' — an always-on agent machine on Tailscale — rachittshah · 2026-10-05
- Critic Concedes OpenAI's GPT Computer-Use Now Drives Safari Like a Human, Barely Errs — kimmonismus · 2026-10-05
- Polyphonic update: one agent orchestrates Claude, Codex, Grok end-to-end — RileyRalmuto · 2026-10-05
- ESR: AI-powered decompilation of AAA games means the end of closed source is near — josephdviviano · 2026-10-05
- Evals in the agentic era should run with and without a harness, researcher says — prajdabre · 2026-10-05