Dev Patches vLLM to Run DiffusionGemma, Live Evals Show It Ties on Smarts but Loses on Speed to APIs
bodonoghue85 · x · 2026-09-18
Developer mmastrac ran real, live evals comparing his vLLM patch (DiffusionGemma-as-Jev) against the commercial API (Jev):
- DiffusionGemma (running locally on DGX Spark) is not faster than the Jev API — the API wins on speed.
- On intelligence, the two are roughly tied.
- The author concludes DiffusionGemma comes out as the winner overall.
More from Infra
- Dev Builds 'Local ChatGPT in a Box' on One RTX 5090, Sharing Every Workaround Along the Way — valdev · 2026-09-18
- King Charles Meets OpenAI, Anthropic, DeepMind and Nvidia Execs on AI Safety — eyishazyer · 2026-09-18
- PlanetScale's TIN beats Postgres GIN full-text search: 212ms vs 288s at p99 — DanielLockyer · 2026-09-18
- NVIDIA shows 100x faster scikit-learn spectral clustering with cuML — NVIDIA Developer · 2026-09-18
- ChatGPT desktop app leaks context to the cloud by default — here's how to swap in Ollama — Technovangelist · 2026-09-18
- Reddit thread: what max-context KV reservations actually cost beyond concurrency — werunm · 2026-09-18