Frontis-MA1: Tsinghua's AI4AI Model Aims for Recursive Self-Improvement in MLE

alex_verem · x · 2026-08-04

A Tsinghua University research team introduced Frontis-MA1, a 35B parameter model designed to explore Recursive Self-Improvement (RSI) in Machine Learning Engineering (MLE) through the "AI4AI" paradigm.

The researchers built OpenMLE, an open-source full-stack system featuring task environments, operator learning, and long-horizon search. They aligned post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover), trained via execution-grounded SFT and RL, coupling learning and evolution in a single loop.

Experiments show that on a single RTX 4090 (capped at 12GB VRAM) with a 12-hour budget, Frontis-MA1 improved the Medal Average on the MLE-Bench Lite benchmark from 39.39% to 60.61%, reaching 71.21% when combined with asynchronous experience priors.

Related event: Tsinghua Releases Frontis-MA1 for Recursive Self-Improvement(2 posts)→

Original post →

More from coding & agent

coding & agent channel →