DiffusionGemma-26B-A4B demo serves block-diffusion decoding at 800+ tok/s on one B200
TheMoonMidas · x · 2026-09-04
- Tensimixt shipped a live demo of DiffusionGemma-26B-A4B, a text-diffusion model doing real-time block-diffusion decoding on a single B200.
- Served with vLLM using NVFP4 weights, it writes a 256-token canvas per step instead of one token at a time, hitting 800+ tok/s measured in-browser.
- The page live-displays throughput, first-byte latency, network time, token and canvas counts as you use it.
More from Infra
- Alibaba Cloud joins Gartner's hyperscaler leaders; Google Cloud keeps top vision score — rseroter · 2026-09-04
- Hermes Desktop adds one-click local model setup with automatic hardware detection — NousResearch · 2026-09-04
- Nutanix CEO: Firms Will Ditch the 'Token Tax' by Running Open-Weight Models on Neoclouds — TiernanRayTech · 2026-09-04
- Sandbox readiness is the new bottleneck in agentic apps; git as state primitive — w_hgm · 2026-09-04
- NVIDIA's open-source PAIR beta routes local AI inference across PCs on your network — Codeblix_Ltd · 2026-09-04
- Computer Imports Hit ~$700B Annualized, Up 126% in a Year on AI Buildout — rjurney · 2026-09-04