Local AI registry: tested model recipes for 40 GPUs with real tok/s numbers

haydendevs · x · 2026-09-27

A new "Local AI" registry (localsybilsolutions.ai) helps you pick the best local model for your GPU: pick a card and get the top three models it can run, measured speeds (e.g., Qwen3.8-27B EXL3 3bpw at 174 tok/s on an RTX 4090), and a copy-paste launch command.

A practical shortcut for anyone doing local deployment without trial-and-error on quantization and inference stacks.

Original post →

More from Infra

Infra channel →