Domain specialists orchestrated by a general model: a local-LLM architecture pitch for 8-16GB GPUs

CyberExplore · reddit · 2026-10-06

The poster is building domain-focused local models from scratch: a general model that reasons and generates specs, paired with focused coding models that simply implement them, all orchestratable. His core argument: MoEs activate few parameters but the whole model must sit in memory, which fails 8-16GB GPU users; a group of specialists with a general-purpose orchestrator might beat MoE on lower-end PCs while enabling parallelism. His rig is 28GB VRAM (RTX 4070 + 5060 Ti 16GB on an x870e board) for dual-model distillation and RL-based post-training. He calls for community-trained models and asks for pointers to prior work.

Original post →

More from Infra

Infra channel →