AURORA-LM: A 1B Continuous Diffusion Language Model Trained on Ascend NPU
机器之心 · wechat · 2026-08-08
A joint research team from Nanjing University and other institutes proposed AURORA-LM, a continuous-latent diffusion language model. It achieves strong results in text generation and is fully trained on Ascend NPU, scaling up to 1 billion parameters.
Core Design & Techniques:
- Two-stage Modeling: An autoencoder first maps text into a high-capacity, decodable continuous latent sequence. A block-wise causal diffusion model then learns to generate this sequence.
- Latent Optimization: Text generation quality is significantly improved by narrowing the noisy input dimension (bottleneck) and increasing the probability of high-noise state training.
- Self-Trajectory Consistency: To mitigate error accumulation in few-step denoising, the model introduces a consistency loss that aligns predictions across adjacent denoising states, showing major improvements in few-step generation.
Results:
- At 130M scale, AURORA-LM outperforms autoregressive and other diffusion baselines in both free generation and conditional summarization.
- Scaled to 1B parameters (AURORA-LM-L), it achieves an average score of 32.6 across nine language benchmarks, beating larger continuous language models.
More from Infra
- Counterintuitive Test: MiniMax H3 Full BF16 Model is Faster Than INT8 and Better at Physics — Wise_Revolution385 · 2026-08-08
- Nvidia to Invest Up to $3 Billion in Blackstone-Backed Power Firm — pstAsiatech · 2026-08-08
- Local Video Generation: MiniMax Lags Far Behind LTX in Inference Speed — PhilosopherSweaty826 · 2026-08-08
- Deep Dive: How Weak is the Evidence for China's Role in US Data Center Backlash? — AndyMasley · 2026-08-08
- $50k Bet Challenges SemiAnalysis on SpaceX AI Compute and ARR Forecasts — generativist · 2026-08-08
- Data Center Boom Drives 'Second Great Construction Divergence' and Job Growth — ivan_bezdomny · 2026-08-08