StepEdge Launches: On-Device Models Accelerate Adoption

AGI Hunt · wechat · 2026-07-12

On-Device Models Are Gaining Ground Fast

The article first references a trend chart from r/LocalLLaMA, discussing that from frontier cloud models to open-source models that can run locally on laptops/phones, the capability catch-up takes an average of about 24.8 months. The author speculates that in the next two years, on-device models will see a significant explosion, because the value of local operation includes offline availability, low latency, privacy, and regional accessibility.

StepEdge: On-Device Family for Phones and Cars

StepEdge launched by StepFun at WAIC, positioned as a "1+N" architecture: a text+vision base model plus three multimodal models for Audio, GUI, and Gen, targeting phones and cars. The company emphasizes local tool call latency as low as 100ms, with local responses for simple tasks and cloud for complex ones.

Evaluation and System Engineering

The article lists multiple results: the base model achieves an average 62.92 across 16 benchmarks including image understanding, GUI Grounding, OCR, tool calling, and AppAgent. It performs exceptionally well on agent tasks like AppWorld and MINDCUBE. The accompanying StepInference NPU inference engine is faster than llama.cpp on Hexagon NPU for text, image, and speech scenarios, with prefill speeds up to 1395 TPS.

Cloud-Edge Division is the Direction

The author summarizes StepFun's layout as Pro + Flash + Edge three layers: Pro handles complex reasoning, Flash handles high-frequency low-latency cloud agents, and Edge handles local terminal execution. The article argues that this "cloud super brain + terminal execution layer" architecture is becoming industry consensus.

Related event: StepFun Releases Step Edge On-Device Models(3 posts)→

Original post →

More from Infra

Infra channel →