stuntd 0.1.2: local heads answer confident LLM calls, weakest field caps the hit rate

Inevitable-Log5414 · reddit · 2026-10-01

The author released stuntd 0.1.2, an open-source proxy that learns your LLM's typed decisions and answers confident ones with small local heads (20ms GPU, 60ms CPU), only calling the big model otherwise.

Code on GitHub (bladedevoff/stuntd); try it on Hugging Face Spaces.

Original post →

More from Infra

Infra channel →