A forked SGLang stack brings Qwen and Laguna to 4x V100 GPUs

Primary_Exchange21 · reddit · 2026-07-27

V100 users get a forked SGLang stack for Qwen and Laguna

A Reddit post describes a fork of SGLang customized for older V100 GPUs. The author says they:

They also report trying, but failing, to make DFlash work for Laguna so far. On their 4×V100 32GB NVLink setup, they claim roughly 4000–6000 pp and about 100 tokens/sec for Qwen.

The post includes a repo link, a TileLang FA repo, Marlin-V100, and says there is a Docker image so users do not need to spend a long time building from source.

Original post →

More from coding & agent

coding & agent channel →