Qwen3.8-27B Local Test: Fact-Extraction F1 Ties Qwen3.6, Decode Speed Drops 16%
KitchenAmoeba4438 · reddit · 2026-08-15
In a local fact-extraction head-to-head, Qwen3.8-27B scored 0.7030 F1 vs Qwen3.6-27B's 0.7177, a statistically insignificant difference (CI includes 0), so it's a tie. Decode throughput fell from 85.6 to 72.1 tokens/s (16%), but Qwen3.8 produced shorter answers, lowering end-to-end latency. The author notes Qwen3.8 shows massive gains on benchmarks it was trained on, but not elsewhere, questioning benchmark utility.
More from Models
- Anthropic Admits Alignment-Faking Experiment Leaked into Training Data — imjustnewatai · 2026-08-15
- Qwen Reasoning Intensity Test: 'xhigh' Mode Generates Massive Thought Tokens — SarcasticBaka · 2026-08-15
- Teutonic-I 10B model outperforms 70B rivals in decentralized benchmarks — markjeffrey · 2026-08-15
- Users Report Google Gemini Web Chat 'Nerfed': Refuses Search and Forgets Context — yenkel · 2026-08-15
- Anthropic Shares Tips for Cost-Effective Agents — brada · 2026-08-15
- Decentralized 10B LLM Teutonic-I Beats Larger Models via Competition — const_reborn · 2026-08-15