Ornith-1.5: AI Generates Its Own Training Data, Beats Claude on Terminal-Bench

新智元 · wechat · 2026-08-27

Ornith-1.5 has been released with a novel self-improvement mechanism where the model generates its own tasks, scaffolds, and trajectories. This three-stage loop, optimized via GRPO, boosted its Terminal-Bench 2.1 score from 77.5 to 86.1, surpassing Claude Opus 4.8 (85.0) in their internal tests.

Core Mechanism

Model Performance

Caveats

The reported scores are from internal tests with specific configurations (e.g., 4h timeout) and differ from the strict official Terminal-Bench leaderboard rules. While the weights are open-sourced (MIT), the full training pipeline and datasets are not.

Related event: Ornith-1.5 Open Models Claim Claude Opus-Level Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →