MegaCapybara: RTX 5090-only inference engine hits 2000+ t/s, 2x faster than Ninfer
BringTea_666 · reddit · 2026-10-03
A new open engine purpose-built for RTX 5090 and Qwen3 27B claims 2x the decode speed of Ninfer: 500+ t/s single-stream, 2000+ t/s across 12 concurrent agents with 800k context. Highlights:
- Dynamic kernels per model/task/context-length, plus Dflash2 speculative decoding with DSpark confidence scheduling and draft trees.
- Benchmarked weights: quantized weights ship with metadata showing KL divergence and top-1% loss vs BF16 at every setting, comparable against unsloth quants in the launcher.
- Agent serving: scheduler prioritizes multi-job throughput; fan out 30+ agents.
- Cache management: prefill once per job; context-swapped tasks resume in 0.03s without re-prefill.
- Loop Guard: escalates repetition penalty on agent loops, then fires a stop signal for frontend recovery.
Ships with a GUI launcher, HF auto-downloader, and CLI/bat export. Weights are on Hugging Face; source code to be released later.
More from coding & agent
- AI is turning programming from a writing skill into a reading skill — _jaydeepkarale · 2026-10-03
- OpenWiki: open-source AI agents build evidence-backed wikis for your codebase — abhishek__AI · 2026-10-03
- 'Cursor is cooked' takes it back: the hot take that aged like milk — gabriberton · 2026-10-03
- Obol: free self-hosted MCP gateway with argument-level Cedar policies — ToBeContinuedHermit · 2026-10-03
- Anthropic releases free workshop on building collaborative multi-agent workflows — s_mohinii · 2026-10-03
- Jevbox: open-source self-organizing document drive with hierarchical search, no vector DB — rickasaurus · 2026-10-03