A small NOBLE tweak adds a multiplicative scale term and lowers eval loss

torchcompiled · x · 2026-07-21

The post shares a small training tweak on top of the NOBLE setup: in addition to projecting into an additive residual term, it also adds a multiplicative scale term.

A linked loss plot shows the modified variant performing better, with the highlighted run reaching a lower eval loss than the baseline variants by 25k steps. The author describes it as a short, practical improvement with only a slight parameter change.

Original post →

More from Research

Research channel →