Qwen3.8-27B PrismaAqua benchmark: Near-BF16 quality

offgridai · reddit · 2026-08-18

A user benchmarked the 5.5-bit PrismaAqua quantization of Qwen 3.8 27B on an RTX 5090 using vLLM, focusing on tool use, business reasoning, and investment tasks. Results show the quantized model closely matches BF16 performance in a custom test suite, achieving 95% accuracy in tool use. The user found that setting reasoning effort to "medium" outperformed "xhigh" by avoiding token limit timeouts while maintaining quality. Specific vLLM configuration parameters and throughput metrics are provided.

Original post →

More from Models

Models channel →