Celeris-1 launches with diffusion inference, 157 ms latency and 76% MMLU-Pro
timshi_ai · x · 2026-07-24
Celeris-1 is launched as a general-purpose language model that claims near-GPT-5-level intelligence with much faster responses.
- The model uses a diffusion-based inference architecture instead of standard autoregressive generation.
- Reported latency: 157 ms p50, about 15× faster than GPT-5-mini and 17× faster than GPT-5.
- Benchmark: 76% on MMLU-Pro, versus 78% for GPT-5-mini and 81% for GPT-5.
- Throughput claim: 1,280 tokens/sec, compared with 144 for Gemini 3.5 Flash-Lite in the poster’s reconstructed test.
The post frames Celeris-1 as a speed-first frontier model and says it is available now.
More from Models
- Grok and Claude get personified as Elon and Dario in a new model-mood meme — kevinnbass · 2026-07-24
- Kimi K3 beats GLM 5.2 on 100 deep-research tasks, but costs 5x more — AravSrinivas · 2026-07-24
- Kimi K3 and Claude Fable5 get called the best large-model frontend aesthetes — vista8 · 2026-07-24
- Dev rebuilds interactive 3D globe in 1.5 hours using Kimi K3 — DuRuofei · 2026-07-24
- A Kimi demo claims it rebuilt a Google Maps 3D experience in 1.5 hours — shakoistsLog · 2026-07-24
- Kimi K3 is rumored to go open-weight and be 3x faster on Monday — bindureddy · 2026-07-24