37 benchmarks, 130K decisions per model: local LLMs tested on one RTX 6000 PRO

multimodalart · x · 2026-09-22

multimodalart ran a large local-model benchmark: 37 benchmarks across 5 categories with 130K decisions per model, all on a single RTX 6000 PRO. Top results include one-pass inference techniques for Qwen3.8-27B and diffusion gemma without fine-tuning, and third-place Decider 35B-A3B, a strong fine-tune of Qwen3.5 35B base.

Related event: Decision Index 0.1 Launches: 132k Decisions Benchmark Jev Against 30+ Open Models(7 posts)→

Original post →

More from Models

Models channel →