Go-LLM repo details running Flash-Next on dual RTX 3090 via vLLM, with honest benchmarks

Motor_Ad16 · reddit · 2026-09-15

A new open-source repo, Go-LLM, packages working notes for running Qwen3.8-Flash-Next with vLLM: the official path (vLLM nightly + Blackwell hardware + NVFP4 checkpoints) and a local workaround on dual RTX 3090 using a GGUF-capable 0.29.0 fork. Stable IQ4XS quantization delivers 36.5 tok/s at 1k decode and 19 tok/s at 126k, while Q4KXL stays experimental due to long-context instability. The author also retracts an earlier non-reproducible 6665 tok/s prefill figure.

Original post →

More from Infra

Infra channel →