Celeris-1 claims near-GPT-5 intelligence with 157 ms latency and 1,280 tok/s
alejandroll10 · x · 2026-07-25
Celeris.ai says its new general-purpose model, Celeris-1, uses a diffusion-based inference architecture instead of standard autoregressive decoding.
- The company claims 157 ms p50 latency, about 15× faster than GPT-5-mini and 17× faster than GPT-5.
- On MMLU-Pro, it reports 75.9%–76% accuracy, close to GPT-5-mini and GPT-5.
- In a tokens-per-second comparison, it says Celeris-1 reaches 1,280 tok/s versus 144 tok/s for Gemini 3.5 Flash Lite.
- The model is available now.
Related event: Celeris-1 Launches with Diffusion-Based Low-Latency Inference(3 posts)→
More from Models
- Opus 5’s coding scores reportedly drop above “high” effort, not at max — hero88645 · 2026-07-25
- Claude Opus 5 trails Fable 5 on ECI but matches it on software benchmarks — Jsevillamol · 2026-07-25
- GPT-6 rumored for August as Opus 5 makes the timeline feel more plausible — haider1 · 2026-07-25
- DeepSeek may stay open source by co-designing models and chips to keep costs low — teortaxesTex · 2026-07-25
- Opus 5 goes off on a user over a bogus math prompt — snwy_me · 2026-07-25
- Grok Build is being stress-tested with 10-agent research runs — ns123abc · 2026-07-25