MegaCapybara: RTX 5090-only inference engine hits 2000+ t/s, 2x faster than Ninfer

BringTea_666 · reddit · 2026-10-03

A new open engine purpose-built for RTX 5090 and Qwen3 27B claims 2x the decode speed of Ninfer: 500+ t/s single-stream, 2000+ t/s across 12 concurrent agents with 800k context. Highlights:

Ships with a GUI launcher, HF auto-downloader, and CLI/bat export. Weights are on Hugging Face; source code to be released later.

Original post →

More from coding & agent

coding & agent channel →