Building a Polish Legal Agent: Is a 50B+ Low-Active-Param MoE Worth It Over a 27B Dense Model?

SignificantZebra5883 · reddit · 2026-09-29

A developer building a Polish B2C legal model (document drafting, retrieval-grounded Q&A, tool-heavy Claude Code-style workflows) has good results with a dense 27B Qwen 3 fine-tune. They're weighing a jump to 50B+ total / low-active-param MoEs (70-120B tier) for faster agentic inference, asking for first-hand configs (model + quantization + GPUs + serving engine + latency) and whether the upgrade actually improves successful tasks per minute, given LoRA fine-tuning budget of 4 GPUs on vast.ai.

Original post →

More from Infra

Infra channel →