Building a Polish Legal Agent: Is a 50B+ Low-Active-Param MoE Worth It Over a 27B Dense Model?
SignificantZebra5883 · reddit · 2026-09-29
A developer building a Polish B2C legal model (document drafting, retrieval-grounded Q&A, tool-heavy Claude Code-style workflows) has good results with a dense 27B Qwen 3 fine-tune. They're weighing a jump to 50B+ total / low-active-param MoEs (70-120B tier) for faster agentic inference, asking for first-hand configs (model + quantization + GPUs + serving engine + latency) and whether the upgrade actually improves successful tasks per minute, given LoRA fine-tuning budget of 4 GPUs on vast.ai.
More from Infra
- Meta Muse to cost ~$50 per user per year even under aggressive optimization, back-of-envelope says — bookwormengr · 2026-09-29
- Redditor runs gpt-oss-120b across a phone, three Macs and two Windows PCs — ANR2ME · 2026-09-29
- BioNeMo team boosts Mixtral-8x7B training throughput 2.21x vs HF BF16 baseline — AllThingsApx · 2026-09-29
- DeepSeek's elastic compute team is hiring heavily, shares sandbox infra for large-scale agent training — teortaxesTex · 2026-09-29
- Developer slams third-party inference providers: Gemini up 10x, Luna 15s latency — julianharris · 2026-09-29
- Nereus: adaptive parallelism boosts 8B PPO throughput up to 7.27x over OpenRLHF — Songlin Jiang · 2026-09-29