Domain specialists orchestrated by a general model: a local-LLM architecture pitch for 8-16GB GPUs
CyberExplore · reddit · 2026-10-06
The poster is building domain-focused local models from scratch: a general model that reasons and generates specs, paired with focused coding models that simply implement them, all orchestratable. His core argument: MoEs activate few parameters but the whole model must sit in memory, which fails 8-16GB GPU users; a group of specialists with a general-purpose orchestrator might beat MoE on lower-end PCs while enabling parallelism. His rig is 28GB VRAM (RTX 4070 + 5060 Ti 16GB on an x870e board) for dual-model distillation and RL-based post-training. He calls for community-trained models and asks for pointers to prior work.
More from Infra
- Strata v0.1.40 adds Strix Halo support, multi-GPU batching and decode improvements — lxfater · 2026-10-06
- Beam Called an Iso-Flop Replication of DeepSeek V3: Same FLOPs Two Years Later, Far Worse Efficiency — mike64_t · 2026-10-06
- BlackBerry's QNX hits record $80.3M revenue, up 27%, betting on AI-era safety platform — alysha_lobo · 2026-10-06
- Long-running benchmarks find Strata inference server failing full-build scenarios — julianharris · 2026-10-06
- AMD R9700 owner: ROCm on Windows cripples local i2v, Vulkan runs fine — vladomkd · 2026-10-06
- CutBCE: TPU Kernel Eliminates OOM in Large-Vocabulary Recommendation Training, 91.9% Faster — _reachsumit · 2026-10-06