Ornith 1.5 35B Hits 81.8 on GPQA with Thinking Mode, Decodes at 303 tok/s
MikePFrank · x · 2026-08-23
Enabling 'thinking mode' boosts Ornith 1.5 35B's GPQA-diamond score from 52.0 to 81.8, leveling it with top dense 27B models. The model decodes at 303 tok/s on a single RTX 5090 (Q4KM, 20GB VRAM), maintaining 258 tok/s at 32k context. Standard benchmarks include MMLU 82.2 and GSM8K 91.7. NVFP4 builds on vLLM show similar performance.
More from Infra
- Tension between data center opposition and AI industry expansion — NathanpmYoung · 2026-08-23
- Upgrading RTX A6000 thermal paste and fan makes it usable for workloads — cephaloform · 2026-08-23
- Optimized llama.cpp fork for AMD GFX906 (Mi50, Mi60, Radeon VII) — milpster · 2026-08-23
- Nvidia AI Server Prices to Rise 15%+, GB300s Around $600k — zephyr_z9 · 2026-08-23
- Is ROCm worth it on Windows for generation speed? — Low-Location5266 · 2026-08-23
- AI compute differs from gold: GPU depreciation and physical limits reshape hedging — AccBalanced · 2026-08-23