MoE hits ~2000 tok/s prefill on 32GB VRAM for bulk text extraction

MkGod · reddit · 2026-09-14

On 2x RTX 5060 Ti (32GB VRAM), the author benchmarked local models for bulk structured extraction: Qwen 27B Q4KM managed 50 t/s, while the MoE Ornith-1.5-35B Q4 reached 2000 t/s prefill and 100-110 t/s generation. For read-heavy, short-output extraction workloads, the prefill jump cuts total time dramatically.

The post also asks whether modern 7B-14B dense models suffice for schema extraction, or whether MoE is the obvious pick, and whether any leaderboard tracks real local prefill/decode speeds.

Original post →

More from Infra

Infra channel →