LocalMaxxing: community-run speed tests for local LLM rigs, 8.6K runs across 592 hardware setups
lxfater · x · 2026-10-04
- LocalMaxxing is a community benchmark site where users submit real-world local LLM inference speed tests via CLI — no vendor numbers. So far: 2,651 users, 8,604 runs, 592 hardware configs, 726 models tested.
- Filterable by model, GPU, inference engine, and quantization; key metrics are decode speed (tokens/s), time-to-first-token, and VRAM usage.
- Extras include a decode calculator, community quality benchmarks, a hardware marketplace backed by speed-test listings, and rentals.
- Useful before buying a GPU or configuring a local rig; always check test conditions since engine and quantization choices change results significantly.
More from Infra
- HPE Posts Record $12.2B Revenue, Up 34% YoY, as AI FOMO Drives Hardware Boom — DavidLinthicum · 2026-10-04
- Coding agent costs down two months straight: LangChain CEO shares 3-step playbook — hwchase17 · 2026-10-04
- Rural data centers are set for a big US federal tax break — but some hyperscalers aren't biting — nordicinst · 2026-10-04
- Rural data centers are in for a big federal tax break, but hyperscalers seem lukewarm — Wired AI · 2026-10-04
- Same open-weights model doom-loops in one engine, runs fine on llama.cpp — pand5461 · 2026-10-04
- Lenovo's AI Express ships on-prem AI servers in 15 business days — shashib · 2026-10-04