Running GPT-OSS 120B on a 4070 Ti at 21 tok/s via Aggressive MoE Caching

JayB_Official · reddit · 2026-08-09

A developer successfully ran the 59GB GPT-OSS 120B model locally on a 12GB VRAM RTX 4070 Ti, achieving 21 tok/s generation speeds.

Key Implementation Details:

The author demonstrated that with 'caveman engineering' using open-source tools, consumer hardware can run ultra-large models.

Original post →

More from coding & agent

coding & agent channel →