Hugging Face weight dump shows a math SFT with AIME 2025 gains at 16k context

kastnerkyle · x · 2026-08-04

A Hugging Face post shares the weights for a light math SFT cold start used for RL runs, trained on 500M tokens from V4 Pro math reasoning and trajectories 6k–12k tokens long.

It also reports AIME 2025 results for the model:

The author says this was the first training run that didn’t come with the usual infrastructure trauma such as poor throughput, kernel failures, or loss spikes.

Original post →

More from Models

Models channel →