Hybrid Mamba-Transformer skips RoPE: Nemotron-style arch fixes Mamba's ICL and long-context weaknesses with few attention layers
gordic_aleksa · x · 2026-09-18
- Aleksa Gordić highlights NVIDIA's Nemotron team's architecture work (SSM/Mamba scaling, hybrids, LatentMoE)
- Key insight: hybrid SSM-transformer architectures obviate explicit positional embeddings like RoPE, since Mamba's hidden states implicitly encode position
- Pure Mamba weaknesses: weak copying/in-context learning (only +1.45 MMLU from 0->5-shot vs +4.38 for transformers) and poor long-context reasoning (possibly from recurrent state over-compression)
- Adding a surprisingly small number of attention layers patches both weaknesses
More from Research
- Xiaomi MiMo achieves streaming large-scale LLM RL training, including a 1T-parameter model — stanfordnlp · 2026-09-18
- Using an LLM as benchmark scorer fails: over-optimistic ratings diverge from human judgment — amplifiedamp · 2026-09-18
- Anil Seth reflects on two years of debate over his conscious-AI target article — anilkseth · 2026-09-18
- Anil Seth's Conscious AI and Biological Naturalism collection out in BBS with fifty commentaries — anilkseth · 2026-09-18
- Anil Seth's conscious AI collection published in BBS with 50 commentaries — anilkseth · 2026-09-18
- Jev as an LLM judge flops: scores nearly everything positively, disagrees with humans — amplifiedamp · 2026-09-18