Seeking the current best LLM inference setup for dual A100 GPUs
Theio666 · reddit · 2026-08-30
A developer seeks advice on the best LLM setup for 2x A100 GPUs to replace the dated GLM 4.5 Air (fp8). Trials with Qwen 2.5 122B faced issues like malformed tool calls, poor SGLang compatibility, and vLLM bugs. Qwen 35B was found insufficient for complex legal tasks. The user prefers vLLM and avoids aggressive quantization like AWQ due to performance degradation in non-English tasks.
More from Infra
- llama.cpp NUMA mirroring boosts dual-EPYC inference by up to 137% — mattescala · 2026-08-30
- Autonomous Launches Personal AI Datacenter Hardware Starting at $26,100 — dee_hw · 2026-08-30
- Analysis: Meta's AI Infrastructure is Mispriced and Massive — RihardJarc · 2026-08-30
- FlashAccel: Leveraging High-Bandwidth Flash for LLM Inference — 9r4n4y · 2026-08-30
- Maximizing throughput: running parallel LLM instances on 2x V100s — Kike328 · 2026-08-30
- Qwen3.8-27B on RTX 5090: NVFP4 Quantization Achieves 256 t/s Code Gen with 175k Context — pennyonaire · 2026-08-30