StepFun releases Step5 Preview: 600B/27B MoE, 1M context, beats Opus 5 on kernel optimization
赛博禅心 · wechat · 2026-09-20
StepFun launched flagship Step5 Preview: 600B total / 27B active params, 1M context, native text+vision. Priced at ¥7 input / ¥20 output / ¥0.35 cache per M tokens; open-sourcing Oct 15. Scores 44 on Artificial Analysis, top-3 among open models, free on AGIBar for a week.
Highlights
- Positioned for software engineering, deep research, and long-horizon agent workflows
- Scores 49.0% avg@4 on its own StepCodeBench (553 repos, 33 languages); low/mid-difficulty tasks near Claude Opus 5
- Kernel optimization: 508 TFLOPS on one H100 in 24h, beating Claude Opus 5 (493) and Kimi K3 (307)
- Self-improvement experiment: lifted Qwen3-30B-A3B's AIME24 accuracy from 53.3% to 60% in 24h
- Played Pokémon Red unoptimized: 3,093 turns, 6.29M tokens, 3 badges
Tech: 92-layer "narrow but deep" design, SparseGQA + block-wise token merging cutting indexer cost to 1/8, bit-wise on-policy RL alignment with <10% overhead, and model-in-the-loop data generation with anti-hacking audits.
More from coding & agent
- mitsuhiko on the Shared Frustrations of Agentic Software Engineering — mitsuhiko · 2026-09-20
- Comprehensive 55-minute Codex Desktop Tutorial Covers Skills, MCP, TikTok and Blender — aziz4ai · 2026-09-20
- Stalkr adds keyword groups to benchmark your brand vs. competitors, with API and MCP access — marclou · 2026-09-20
- Proval: open-source self-hosted LLM code review agent in a single Docker container — Dazzling_Cancel4505 · 2026-09-20
- Jev agent demo: classifies bribes, threats and pleas with no keyword matching — chongdashu · 2026-09-20
- Parallel structured LLM answers never check each other: the Zhaozhou MU problem — Successful-Farm5339 · 2026-09-20