Going local-first: 2x 5060 Ti at ~100 tok/s plus a 96GB M5 Ultra, betting on a 3-year inference floor

No-Name-Person111 · reddit · 2026-09-25

A detailed writeup of one developer's move from renting cloud inference to a local-first AI setup, built around owning a permanent, private "inference floor" — betting that open models, quantization, inference engines, caching, and agent harnesses will make the same hardware increasingly useful over three years, while hedging against API price hikes and subscription restrictions.

Hardware & costs

Architecture

The author admits subscriptions offer stronger models cheaper upfront; the bet is on locking in today's level of local intelligence, not on local models matching the frontier.

Original post →

More from coding & agent

coding & agent channel →