Underdog's Husky Inference Engine Claims 4.5x Speedup Over MLX, 730 tok/s on MacBook
jimmykoppel · x · 2026-09-22
Underdog launched Husky, a Model-Specific Inference (MSI) engine claiming up to 4.5x speedups over Apple's MLX, with its Pareto-frontier local model hitting up to 730 tokens/sec on a MacBook.
Husky powers Underdog's invite-only, 100% local and private AI assistant (under 4GB) that uses the browser to book flights, reserve tables, order groceries and cancel subscriptions with user approval, connects directly to Gmail/Outlook for on-device email, and does fully local meeting transcription and notes — no audio or data ever leaves the machine.
More from Infra
- COLIBRI: pure-C zero-dep engine streams 2.8T-param MoE models from disk on consumer hardware — bibryam · 2026-09-22
- Sentdex benchmarks openjev: 169ms on Dell GB10 vs 137ms on RTX 3090 — Sentdex · 2026-09-22
- PyroDash cuts inference cost 96% by having a 4B model call the big one only when needed — jiqizhixin · 2026-09-22
- Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs — simonguozirui · 2026-09-22
- Cloudflare Python Workers go generally available after two-year preview — Simon Willison · 2026-09-22
- Fighting AI crawler traffic: beyond Turnstile, Cloudflare's AI Labyrinth as an option — fforres · 2026-09-22