ISO proposes a fixed-spectrum optimizer for RLVR and matches AdamW in 100 steps

burny_tech · x · 2026-07-24

What the paper claims

The paper argues that RLVR does not mainly learn by rewriting a model’s weight spectra. Instead, it reuses the base spectrum and adapts the singular frames around it.

ISO: a fixed-spectrum RLVR optimizer

Reported results

The paper’s core message is that post-training may be better designed around inheriting the spectrum, optimizing the frames.

Related event: ISO Framework Optimizes RLVR Training by 2.7x(3 posts)→

Original post →

More from Research

Research channel →