NVIDIA paper: Nemotron-H, a family of hybrid Mamba-Transformer models
cwolferesearch · x · 2026-08-19
NVIDIA released the Nemotron-H paper, detailing a hybrid Mamba-Transformer architecture for accuracy and efficiency. The accompanying discussion highlights midtraining/CPT techniques: emphasizing high-quality domain-specific data (math, code, science) while retaining general data to avoid catastrophic forgetting.
More from Models
- News-optimized model still flagged 100% AI by detectors — ivan_bezdomny · 2026-08-19
- Ornith-1.5 open-source family launches, 397B MoE claims Claude Opus 4.8 parity — KokaOP · 2026-08-19
- Grok 4.6 draws a pen in ForgeCAD, testing model's knowledge of physical product internals — MarvinTBaumann · 2026-08-19
- User complains Fable 5 can code but won't audit security due to policy — Few_Object_2682 · 2026-08-19
- User test finds GPT Pro unmatched in reasoning depth, correcting data and citing 50-year-old research — PromptOutlaw · 2026-08-19
- Mistral lags far behind frontier models in cyberdefense — emollick · 2026-08-19