Qwen3.8 27B评测:一次性编程能力显著提升
kms_dev · reddit · 2026-08-15
对比测试了 Qwen3.8、3.6 和 3.5 三个版本的 27B 模型在 oneshot 编程任务上的表现。
- 改进点:3.8 版本成功生成了可运行的 2048 游戏(3.6 失败),准确绘制了 pelican 场景,并基本正确复现了 Wolfenstein 游戏。
- 评分提升:使用 Claude Sonnet 5 评估,平均评分从 3.5 的 2.46、3.6 的 2.74 提升至 3.8 的 3.00。
More from Models
- llama.cpp adds support for Moonshot Kimi-K3 text model — pmttyji · 2026-08-15
- GLM 5.3 report reveals use of synthetic RL environments — burny_tech · 2026-08-15
- Qwen2.5 vs Qwen2 vs Gemma 4: Benchmarking 30B-class open models — MaySaki2 · 2026-08-15
- Testing Fable 5 distillation redirects and SVG generation — teortaxesTex · 2026-08-15
- Gemini 3.7 Flash Generates Entire Interactive Websites in Real Time as You Browse — _philschmid · 2026-08-15
- Testing Qwen 3.8 27B minimal thinking in a specific harness — rosie254 · 2026-08-15