MoE hits ~2000 tok/s prefill on 32GB VRAM for bulk text extraction
MkGod · reddit · 2026-09-14
On 2x RTX 5060 Ti (32GB VRAM), the author benchmarked local models for bulk structured extraction: Qwen 27B Q4KM managed 50 t/s, while the MoE Ornith-1.5-35B Q4 reached 2000 t/s prefill and 100-110 t/s generation. For read-heavy, short-output extraction workloads, the prefill jump cuts total time dramatically.
The post also asks whether modern 7B-14B dense models suffice for schema extraction, or whether MoE is the obvious pick, and whether any leaderboard tracks real local prefill/decode speeds.
More from Infra
- Cloudflare ships granular authz for Workers, granting agents access to a single Worker — dinasaur_404 · 2026-09-15
- Poll: 61% of Americans oppose AI data center construction, young adults most opposed — justin_hart · 2026-09-15
- 24,600 generations show quantization costs aren't uniform: Q2 keeps JSON perfect but tanks arithmetic 66% — Tensor_Ghost_03 · 2026-09-15
- Neoclouds: How Failed Companies Became AI's Biggest Winners — economics of the GPU cloud boom — bycloud · 2026-09-15
- Subnormal floats are expensive — but only on Intel, benchmarks show — lemire · 2026-09-15
- LLM Inference Engineer dubbed the most AI-proof job by tech commentator — ashishllm · 2026-09-15