NVIDIA interns unveil Sigma: first large-scale (3B/8B) continuous diffusion language model
ArashVahdat · x · 2026-10-05
Wei Guo's NVIDIA summer-internship team introduced Sigma, the first large-scale continuous diffusion language model (dLM):
- Trained at 3B/8B scale with a likelihood objective — a first at this size
- Under aligned training budgets, it matches or beats masked (discrete) dLMs and even autoregressive models on math and coding
- Its continuous dynamics enable meaningful study of inference-time interventions
Co-led with Jean-Marie Lemercier and Simon Welker, with NVIDIA researchers including Arash Vahdat on the team.
More from Models
- Reddit User Ditches Codex After Trying Chinese Models: 'Overkill' for 80% of Work — EmetInteractive · 2026-10-06
- Ternary MoE Model Scion-35B-A3B Released on Hugging Face With llama.cpp Support — pmttyji · 2026-10-06
- Kairos 1 claims to be first model simulating individual human behavior, tops 13 benchmarks — misovalko · 2026-10-06
- Dev ditches GPT 6 for Claude Opus 5.5: acts on intent without nudges — jdjohnson · 2026-10-06
- Dev reports Clef-Flash runs fast locally even on memory-bandwidth-limited Jetson Orin — gregmushen · 2026-10-06
- GLM 5.3 Flash vs Tencent Hy3: A Sycophancy Test Crowns Two Least Sycophantic Models — ramendik · 2026-10-06