Go-LLM repo details running Flash-Next on dual RTX 3090 via vLLM, with honest benchmarks
Motor_Ad16 · reddit · 2026-09-15
A new open-source repo, Go-LLM, packages working notes for running Qwen3.8-Flash-Next with vLLM: the official path (vLLM nightly + Blackwell hardware + NVFP4 checkpoints) and a local workaround on dual RTX 3090 using a GGUF-capable 0.29.0 fork. Stable IQ4XS quantization delivers 36.5 tok/s at 1k decode and 19 tok/s at 126k, while Q4KXL stays experimental due to long-context instability. The author also retracts an earlier non-reproducible 6665 tok/s prefill figure.
More from Infra
- NVIDIA Posts $96.2B Quarter: An AI Finance Model Separates Real Earnings From Future Promises — nikola_mr64990 · 2026-09-15
- AI's $1.1T data-center bet: productivity must grow 2.7x by 2030 to break even — nordicinst · 2026-09-15
- Own your embeddings: Weaviate + Ollama local vector search config that never leaves your network — philipvollet · 2026-09-15
- Chatbot injects 252k-token prompt every turn instead of RAG, dev asks to be talked out of it — louay_Sallakho · 2026-09-15
- MinIO: open-source S3-compatible object store with zero egress fees — thisguyknowsai · 2026-09-15
- Before porting code to GPU, measure memory transfer time — data movement is the real bottleneck — Franc0Fernand0 · 2026-09-15