Qwen3.8-27B PrismaAqua benchmark: Near-BF16 quality
offgridai · reddit · 2026-08-18
A user benchmarked the 5.5-bit PrismaAqua quantization of Qwen 3.8 27B on an RTX 5090 using vLLM, focusing on tool use, business reasoning, and investment tasks. Results show the quantized model closely matches BF16 performance in a custom test suite, achieving 95% accuracy in tool use. The user found that setting reasoning effort to "medium" outperformed "xhigh" by avoiding token limit timeouts while maintaining quality. Specific vLLM configuration parameters and throughput metrics are provided.
More from Models
- Unsloth's Qwen3.8-27B GGUF hits #2 on Hugging Face with 2.7M downloads — danielhanchen · 2026-08-18
- Tesla's fleet is a monster advantage enabling exquisite FSD tuning, says James Douma — jamesdouma · 2026-08-18
- Fable data automation backfires: wrong templates cause model regression — cephaloform · 2026-08-18
- SemiAnalysis claims Anthropic's completed Mythos2 is shelved for safety, now used to train Mythos3 — 新智元 · 2026-08-18
- Harvard's Zak Kohane finds 5 AI detectors all flag his own writing as AI — zakkohane · 2026-08-18
- Sakana AI releases Japanese-specialized reasoning model Sakana Namazu — hardmaru · 2026-08-18