Local AI registry: tested model recipes for 40 GPUs with real tok/s numbers
haydendevs · x · 2026-09-27
A new "Local AI" registry (localsybilsolutions.ai) helps you pick the best local model for your GPU: pick a card and get the top three models it can run, measured speeds (e.g., Qwen3.8-27B EXL3 3bpw at 174 tok/s on an RTX 4090), and a copy-paste launch command.
- Currently lists 40 GPUs and 156 recipes, with 76 tested on real hardware against six checks (loads, answers, thinks, calls tools, holds context window, keeps pace); the rest are marked "reported" until verified.
- Coverage spans DGX Spark, RTX PRO 6000 Blackwell, 5090, 4090, and budget cards, and includes third-party published recipes.
A practical shortcut for anyone doing local deployment without trial-and-error on quantization and inference stacks.
More from Infra
- Confidential computing: the answer that wins LLM providers million-dollar enterprise deals — abhijithneil · 2026-09-27
- Serve models from KitOps ModelKit on HAMi: a registry-native path to SGLang inference — HowDevelop · 2026-09-27
- How do teams actually control LLM inference costs in production? A Reddit thread asks — Ok_Philosophy_4031 · 2026-09-27
- Profiling an agent harness with jq: latency per turn and byte-level request composition — arthurcolle · 2026-09-27
- Factory AI CEO: 90% of tokens will go to open models within 12 months — matanSF · 2026-09-27
- Chip Design Is a Loop, Not a Flow: Why AI's Role in Silicon Will Be Iterative — ai · 2026-09-27