StepFun's Step 5 Preview knows when to stop, edging GLM 5.3 on long-horizon agent tasks
omarsar0 · x · 2026-09-21
StepFun's new flagship Step 5 Preview is available to try at a competitive price, claiming the cost-capability Pareto frontier. Early hands-on testing by omarsar0:
- Coding: in the same class as GLM 5.3 and Kimi K3. On two real tasks in the same repo/commit (a numeric filter bug and a feature requiring new routes, permission gating, and a refactor), both models delivered correct first-pass fixes with all held-out tests passing and no regressions.
- The standout behavior is knowing when to stop: Step 5 Preview finished, checked its work, and declared itself done both times; GLM 5.3 wrote correct code but kept going until the step limit. This makes Step 5 better suited for unattended agent runs, bug fixes in unfamiliar codebases, and long-horizon tasks.
- Patch quality: its bug fix matched the shorter approach the Datasette maintainer used in the real commit, and it added tests unprompted.
- Long context: found all five hidden clues across 368K tokens of fake incident tickets and connected them into a correct root cause in 90 seconds.
Disclosure: the post was made in partnership with the StepFun team.
Related event: StepFun launches Step 5 Preview, touting cost-performance frontier(3 posts)→
More from coding & agent
- Rethinking classifiers: conversation+tool templates as special cases of inference-time specs — austinvhuang · 2026-09-21
- Analyst predicts Meta will pay Amazon for agent access, spawning a new agent business model — signulll · 2026-09-21
- Arena: coding-agent harnesses show up to 5x cost differences at similar success rates — thione · 2026-09-21
- Recap: Anthropic merged Claude chat and Cowork, redesigned Claude Code Projects — thione · 2026-09-21
- Recap repeat: Claude Code Projects parallel cloud sessions and Astra for Law — thione · 2026-09-21
- Recap: WorkOS explains agent auth with auth.md; Salesforce and Anthropic expand Claudeforce — thione · 2026-09-21