Surge AI Launches Off-the-Shelf Frontier Training Data, Boosting Agent Benchmarks by Double Digits
echen · x · 2026-08-01
Surge AI has announced that its high-quality training data, evals, and RL environments—originally built for frontier AI labs—are now available for direct purchase.
The company released benchmark gains demonstrating the effectiveness of post-training with their datasets:
- Agentic Coding: GLM-4.7 improved from 35.2 to 47.2 on Terminal-Bench 2.0, and from 37.2% to 44.3% on SWE-Bench Pro.
- Enterprise Agents: Qwen3.5-122B-A10B saw a 9.6 percentage point increase on Toolathlon (24.2 to 33.8).
- STEM Reasoning: Scores on FrontierScience jumped from 29.1 to 38.4.
Surge AI noted that the results above were achieved using only a fifth of their datasets. Labs enrolled in their Trusted Program can train on the full dataset and pay only if it demonstrably moves the metrics.
More from Models
- Kimi K3 Model Lands on OpenRouter, Fast™️ Version Debated — AAAzzam · 2026-08-01
- antirez releases DeepSeek V4 Flash GGUF quantizations, from 2-bit to 4-bit, for local deployment — antirez · 2026-08-01
- Discussion: How 'Benchmaxxing' Makes LLMs Unusable for Real-World Tasks — Witty_Mycologist_995 · 2026-08-01
- LLM Long Context Test: Sharp IQ Drop After 300K Tokens — teortaxesTex · 2026-08-01
- Moonshot's 2.8T Parameter Kimi K3 Lands on Microsoft Azure AI Foundry — ccerrato147 · 2026-08-01
- Fable Model Demonstrates Stunning Tacit Knowledge, Uses Physics Terms to Dissuade User — voooooogel · 2026-08-01