stuntd 0.1.2: local heads answer confident LLM calls, weakest field caps the hit rate
Inevitable-Log5414 · reddit · 2026-10-01
The author released stuntd 0.1.2, an open-source proxy that learns your LLM's typed decisions and answers confident ones with small local heads (20ms GPU, 60ms CPU), only calling the big model otherwise.
- Multi-field design: 0.1.2 supports multi-field outputs (e.g., category + urgency + needshuman). The per-field local answering idea was dropped — you pay for the full model call anyway, and half-local answers are hard to debug. It's now all-or-nothing: local only when every field is confident; otherwise the model answers and every field becomes training data.
- Counterintuitive finding: on the support demo, the three heads are individually confident on 99.9%, 92%, and 76% of tickets, but all three at once only 72.7% — the weakest field decides your local hit rate.
- New: autoretrain retrains after N captures, runs the new head in shadow mode until it consistently agrees, rolls back if it degrades; Anthropic Messages support and serve --lazy.
Code on GitHub (bladedevoff/stuntd); try it on Hugging Face Spaces.
More from Infra
- RAG and RL tool calling are bringing CPUs back into ML training, possibly needing dedicated CPU nodes — StasBekman · 2026-10-01
- China's CXMT to spend $5.2b on DRAM expansion, favoring domestic equipment suppliers — pstAsiatech · 2026-10-01
- Cloudflare launches agent-themed batch: pay-per-use gateway, AutoRouter, 6x faster containers — threepointone · 2026-10-01
- Distributed compute market adds 10 B300 nodes, rentable from a single node — markjeffrey · 2026-10-01
- 27B at Q5 with full 131k context on one 24GB RTX 3090, 13-17% faster — bjivanovich · 2026-10-01
- netkit paper: container network namespaces cost up to 31% throughput on Linux — tianyin_xu · 2026-10-01