Surge AI Launches Training Datasets: Major Gains for GLM and Qwen Models
echen · x · 2026-08-01
Surge AI has released off-the-shelf, high-quality training datasets and RL environments covering coding agents, enterprise tool use, and multimodal reasoning. They published benchmark results using popular open-source models:
- GLM-4.7: Jumped from 35.2% to 47.2% (+12.0pp) on Terminal-Bench 2.0 and +7.1pp on SWE-Bench Pro after training on the coding dataset.
- Qwen3.5-122B-A10B: Improved from 24.2% to 33.8% (+9.6pp) on Toolathlon after training on the enterprise-agent dataset.
- Qwen3-4B: Matched the performance of Qwen3-235B-A22B-Instruct after training on their instruction-following dataset.
The catalog operates on a 'pay-if-it-moves-your-metric' model, ensuring all datasets pass out-of-distribution generalization tests before shipping.
More from Models
- Exploring Claude Opus's Odd Visual Outputs with the "Dario and Amanda" Prompt — chicametipo · 2026-08-01
- Open-Source Pressure: Meme Jokes Vendor Cut Prices 80% Due to DeepSeek — InternationalGap3698 · 2026-08-01
- Experiment: Prompting Claude to Code a Procedural Bone and Skin Animation System — chongdashu · 2026-08-01
- Comparing 18 Major LLM API Prices: 100x Cost Difference for Same Workload — mentorperplexed · 2026-08-01
- Comparing 18 Major LLM APIs: Costs Vary by Over 100x for the Same Workload — mentorperplexed · 2026-08-01
- User Reports Claude Opus 5 Feels Janky and Delivers a Worse Experience — vivekhaldar · 2026-08-01