Spite: a modular Rust inference engine where every model, GPU and op is pluggable
giveen · reddit · 2026-10-05
- A developer open-sourced Spite, a Rust inference engine built on one rule: every layer is replaceable without touching any other.
- Models as modules: kernels organized under kernels/<family>/<model>/ (llama4, deepseek/v4, qwen35, etc.); adding a model is just adding a folder, auto-discovered by the dispatcher.
- GPUs as modules: same-family kernels split per GPU (sm89, rdna3, Apple M4), each exploiting FP8 tensor cores, 96MB Infinity Cache, or the Neural Engine.
- Per-op tuning: kernels can optimize only attention while FFN/rmsnorm fall back; improvements stack automatically as the dispatcher picks the best kernel per op per GPU.
- Samplers, tokenizers, KV cache and offload policies are plugin registries; each component is usable standalone.
More from Infra
- Muse ships reliability fixes after SEVs left cron jobs and scheduled tasks unrecovered — alexandr_wang · 2026-10-05
- AWS Mistakenly Suspends Account, Wabi Down for 3+ Hours With No Recourse — soleio · 2026-10-05
- Is Strix Halo the closest thing to a dream local LLM box? Unified memory vs GPUs for 20B-32B models — Robert__Sinclair · 2026-10-05
- 539 tok/s DeepSeek on 4x RTX 6000 — and a call-out that community benchmarks inflate 20-30% — HankYeomans · 2026-10-05
- GLM 5.3 flash on dual DGX Sparks gets 50-90% decode boost with new open recipe — swiebertjee · 2026-10-05
- Qualcomm's Snapdragon to power next-gen AI assistants for Meta and OpenAI — ryanshrout · 2026-10-05