TTCD: Diffusion Language Models with Token Time
UTEXAS · hf · 2026-07-18
This paper proposes Token Time Continuous Diffusion (TTCD), a novel diffusion-based language model.
The two core aspects are:
- Continuous space modeling: The model deterministically maps Gaussian noise to a final token canvas, eliminating the reliance on multi-step parallel sampling in discrete spaces and thereby mitigating accuracy degradation during high-speed generation.
- Token time concept: Different tokens can converge from noise to final tokens at varying speeds, facilitating better conditional generation modeling and allowing tokens to exert differentiated influences during the refinement stage.
Experimental results show that TTCD outperforms discrete models in high-speed settings:
- Achieves performance comparable to existing models of the same scale, dataset, and self-distillation setup in unconditional generation.
- Performs even better in conditional generation.
- Demonstrates similar gains in Sudoku solving tasks.
Training configuration includes:
- 160M parameters
- Training data: OpenWebText
- Followed by self-distillation
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22
- enFoldX gains accuracy as AF3 ensemble disagreement rises for non-binders — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22