Testing Alibaba's Qwen3.8 Max: Spots Med Conflicts but Leaves Gaps
MaziyarPanahi · x · 2026-08-06
A developer tested Alibaba's Qwen3.8 Max model on a synthetic medical discharge note. Across multiple runs at temperature 0, the model consistently reconciled four medication changes and correctly identified a missing furosemide dose without hallucinating it.
The test was based on a single case with response times between 23-34 seconds, making it an observation rather than a formal benchmark.
More from Models
- t0-alpha Released: Open-Source Foundation Model for Time-Series Forecasting — fpedregosa · 2026-08-06
- MiniMax Releases H3 Omni-Modal Model: Supports Video and Native Audio Generation — RisingSayak · 2026-08-06
- MiniMax Releases H3: 33B Open-Source DiT for Image, Video, and Audio — RisingSayak · 2026-08-06
- Google Wins Benchmarks But Loses Developers — prasenx · 2026-08-06
- OpenAI Reportedly Set to Launch Astra Next Week, Largest Pretrain Since GPT-4.5 — koltregaskes · 2026-08-06
- Google's August AI Build: 90 Reusable Agent Skills, Managed Infrastructure, New Gemini Models — rseroter · 2026-08-06