Dev's accidental A/B test: reverting from .8 to .6 reveals dramatic capability gap
AI_Andrew · x · 2026-09-04
A developer shares an accidental side-by-side comparison: after hours of frustration building agent swarms on ".6", he realized he had reverted models—moving back to ".8" exposed a dramatic generational leap.
His takeaways:
- .7: better at math/science but more sterile in vibe.
- .8: thinks longer and harder, has more personality, and compensates for occasional imprecision by fixing its own mistakes and one-upping in the next turn.
He says he's been taking mental notes while building, trying to separate harness effects from model behavior, and promises more detailed comparisons to follow.
More from Models
- GPT-6 reportedly nails SRE-Bench with ~100% pass@4 on never-public reverse-engineering binaries — xennygrimmato_ · 2026-09-04
- New local LLM benchmark tracks prefill speed from RTX 5090 down to Raspberry Pi — maximelabonne · 2026-09-04
- Sakana AI's Takuya Akiba to unpack Kimi K3's architecture: how a 2.8T-param open model was built — tkasasagi · 2026-09-04
- Small model Luna praised for beating DeepSeek and its uptime for personal agents — bindureddy · 2026-09-04
- GLM-5.3 gets updated chat template: tool-result reordering now exits early — victormustar · 2026-09-04
- Qwopus 3.8 27B Flash fine-tune ships: 12.8% faster decoding, 80.7% MTP acceptance on Qwen3.8-27B — EAccelerate_42 · 2026-09-04