vLLM PR adds structured generation mode for DiffusionGemma diffusion LLM
victormustar · x · 2026-09-21
Developer mmastrac opened a vLLM core PR (#57250, 15 commits) adding a structured generation mode for the DiffusionGemma diffusion LLM, described as "Jev-like". It includes prerequisite PRs for logprob bugfixes, prefill-batch perf, and feature parity with autoregressive decoding. If merged, it would bring structured, controlled output to diffusion LLMs on the vLLM inference stack.
More from Infra
- Mozilla AI runs a local 30B model end-to-end to open a real bugfix PR, fully offline — mozilla-ai · 2026-09-21
- Cohere Labs launches Local AI community program for local inference and hardware tuning — Cohere_Labs · 2026-09-21
- Gewell: custom Gemma 4 inference engine cuts KV cache VRAM to 0.625x, losslessly — stoppableDissolution · 2026-09-21
- DeepSeek-V4.1-Flash redesigns the Transformer for agents, cutting KV cache to 890 bytes/token — AndLukyane · 2026-09-21
- Jev Engineering gives agents a decision brain, 193x faster and 444x cheaper in tests — agihouse_org · 2026-09-21
- UK's £225m Isambard-AI supercomputer cost about the same as one road bridge — charlieharris01 · 2026-09-21