Run 193B parameter model on 24GB RAM: 8 specialist models with dynamic routing
Similar_Wealth_1850 · reddit · 2026-08-06
A developer shares how to run a 193B parameter model on 24GB RAM. The approach uses 8 specialized LLMs (triage, router, control, math, code, reasoning, general, vision) on a single consumer machine, with dynamic loading ensuring only one model is active at a time. Routing uses a zero-cost regex scanner and a small triage model. Benchmarks: Medium tier (66B total, ≤14B active) achieves 92% HumanEval; Large tier (193B total, ≤32B active) achieves 95% HumanEval and 94% MATH. The author argues specialization + dynamic loading beats monolithic scaling.
More from Infra
- Nvidia B300 GPU-hour Index Hits All-Time High as Neoclouds Pivot to Inference — rickasaurus · 2026-08-07
- Musk: Chip Capacity is the AI Bottleneck; TerraFab to the Rescue — XFreeze · 2026-08-07
- YC S26 Startup Understudy Cuts Anthropic Bills by 80% via Model Distillation — ycombinator · 2026-08-07
- AI Agent Energy Use is 600x Higher Than Standard Prompt Estimates — tobyordoxford · 2026-08-06
- Sam Altman-Backed Oklo Achieves Reactor Criticality in Under a Year — TinfoilTricorn · 2026-08-06
- Help Wanted: Porting NVIDIA's MiniMax H3 Sol-Engine Optimizations to a Single RTX 5090 — cat_trick · 2026-08-06