37 benchmarks, 130K decisions per model: local LLMs tested on one RTX 6000 PRO
multimodalart · x · 2026-09-22
multimodalart ran a large local-model benchmark: 37 benchmarks across 5 categories with 130K decisions per model, all on a single RTX 6000 PRO. Top results include one-pass inference techniques for Qwen3.8-27B and diffusion gemma without fine-tuning, and third-place Decider 35B-A3B, a strong fine-tune of Qwen3.5 35B base.
More from Models
- Claude counts tokens, not messages: 9 tricks to avoid hitting usage limits — HeyAmit_ · 2026-09-22
- Leaked screenshots surface of rumored OpenAI "Aeon" persistent agent — PrisonOfH0pe · 2026-09-22
- New Decision Index benchmark runs 132,422 decisions; Jev still tops at 59.5 — victormustar · 2026-09-22
- Pelican SVG test puts unreleased GPT-6 Astra head-to-head with Anthropic's Mythos — PrisonOfH0pe · 2026-09-22
- Gemini Pro users report web access silently disabled, even on paid subscription — zsolt67 · 2026-09-22
- M5 Ultra Hits 3740 tok/s Prefill on Qwen, Nearly Double Overnight — EAccelerate_42 · 2026-09-22