Seeking the current best LLM inference setup for dual A100 GPUs

Theio666 · reddit · 2026-08-30

A developer seeks advice on the best LLM setup for 2x A100 GPUs to replace the dated GLM 4.5 Air (fp8). Trials with Qwen 2.5 122B faced issues like malformed tool calls, poor SGLang compatibility, and vLLM bugs. Qwen 35B was found insufficient for complex legal tasks. The user prefers vLLM and avoids aggressive quantization like AWQ due to performance degradation in non-English tasks.

Original post →

More from Infra

Infra channel →