Stripping antirez's ds4 to 45k lines makes Qwen3.8 Flash Next ~10% faster, bit-exact
Chida82 · reddit · 2026-10-05
Developer Chida82 forked antirez's ds4 (DwarfStar) inference engine, keeping only Qwen3.8 Flash Next and the Metal backend, cutting ds4.c from 85k to 45k lines so the code tree fits in one context window.
Key results:
- Rationale: in the coding-agent era, a patch's real cost includes how many tokens the agent must read; removing noise makes optimization cheaper and safer
- Q2: decode +9–13%, MTP from 75.8 to 86.7 tok/s; Q4: prefill +2–11%, MTP 77.8 → 85.9 tok/s
- Output is bit-exact vs stock ds4 via GGUF/greedy parity checks and interleaved A/B benchmarks; no KV cache quantization or approximate kernels
- Triaged upstream PRs: adopted 20 commits, dropped 30
- Added SSD streaming for expert weights: 27 tok/s (Q2) on a simulated 48GB machine, 35 tok/s with MTP
- Keeps syncing with upstream via git merge + rerere; StarForge repo is model/backend-agnostic
A showcase of the new engineering paradigm: slimming codebases for agent maintainability.
More from coding & agent
- Dev ships AI-built game by looping relentless AI critique, using Three.js, ElevenLabs and Suno — AIandDesign · 2026-10-05
- Karpathy shares tips for reading LLM outputs: ASD-STE100 writing and diagrams — dair_ai · 2026-10-05
- Dev rebuilt his coding agents' UI as iMessage and says life is better now — davidfromkansas · 2026-10-05
- Redditor builds offline AI agent network across Legion laptop and iPhone with Tailscale — Due_Recording_5802 · 2026-10-05
- Every MCP tool passed in isolation, yet the workflow still failed — test full chains — Stock-Pumpkin-8859 · 2026-10-05
- Instead of booking flights, these are the agent demo tasks people actually want — vivekhaldar · 2026-10-05