Together AI: cascading GLM-5.3 Flash to GLM-5.3 cuts cost 57% while boosting DeepSWE to 80.9%
togethercompute · x · 2026-08-29
Together AI's deep dive benchmarks GLM-5.3 vs GLM-5.3 Flash on DeepSWE and tests routing strategies:
- Cascade wins: run Flash first, escalate to full GLM-5.3 only on test rejection — solves 80.9% of tasks at $1.70 each vs 69.0% at $3.99 for the full model alone: 12 points better, 57% cheaper.
- The gap is first-try only: 5.6 points at pass@1 (69.0% vs 63.4%), just 2.6 at pass@4 (87.6% vs 85.0%). Distillation removed reliability, not capability — and retries buy reliability back.
- 17x price gap: $3.99 vs $0.24 per rollout; per $100, Flash solves 264 tasks vs 17.
- Capability kept, consistency lost: Flash still solves 93 of the full model's 99 solved tasks, stabilizes 15 flaky ones, and cracks 3 the parent walls on entirely.
- Caveats: on flaky tasks the longer run passes only 46% of the time (61% for full), and Flash breaks an already-passing baseline in 6.9% of rollouts vs 4.4% — gate it with a regression run.
More from coding & agent
- 'Abundant Constraints Beat Abundant Implementation': An Essay on Directing AI Capability — aishashok14 · 2026-08-29
- fbtee 4.0 released, fully rewritten in Rust with Oxc — cnakazawa · 2026-08-29
- Google Paper: Replace Agent Chat History With Explicit State, Cut Tokens 16x — rohanpaul_ai · 2026-08-29
- Seeking open-source methods to extract tables from Indian bank PDFs — OmPatel110 · 2026-08-29
- One Prompt Turns Gemini Flash Into an Optimization Machine: 5000x Gains in 20 Minutes — doodlestein · 2026-08-29
- Demo: AI Generates App and Integrates into System — BLUECOW009 · 2026-08-29