Surge AI Launches Frontier Training Data: Small Models Match Giants
echen · x · 2026-08-01
Surge AI has officially launched its highly anticipated 'Off-the-Shelf' high-quality training data catalog, featuring a 'pay only if it moves your metric' model.
The company released several impressive post-training benchmark results:
- GLM-4.7 trained on their coding dataset jumped from 35.2% to 47.2% (+12.0pp) on Terminal-Bench 2.0.
- Qwen3.5-122B saw a 9.6pp increase on Toolathlon after training on the enterprise-agent dataset.
- Massive Efficiency Gains: A 6B parameter model (Qwen3-4B) matched the performance of the 235B Qwen3-235B-A22B-Instruct after training on their instruction-following dataset.
- Kimi K2.7 achieved over 90% rubric pass rate on professional tasks (GDPval).
The datasets cover repository-level software engineering, terminal-use agents, long-horizon enterprise tool use, and multimodal reasoning.
More from Models
- Exploring Claude Opus's Odd Visual Outputs with the "Dario and Amanda" Prompt — chicametipo · 2026-08-01
- Open-Source Pressure: Meme Jokes Vendor Cut Prices 80% Due to DeepSeek — InternationalGap3698 · 2026-08-01
- Experiment: Prompting Claude to Code a Procedural Bone and Skin Animation System — chongdashu · 2026-08-01
- Comparing 18 Major LLM API Prices: 100x Cost Difference for Same Workload — mentorperplexed · 2026-08-01
- Comparing 18 Major LLM APIs: Costs Vary by Over 100x for the Same Workload — mentorperplexed · 2026-08-01
- User Reports Claude Opus 5 Feels Janky and Delivers a Worse Experience — vivekhaldar · 2026-08-01