A small NOBLE tweak adds a multiplicative scale term and lowers eval loss
torchcompiled · x · 2026-07-21
The post shares a small training tweak on top of the NOBLE setup: in addition to projecting into an additive residual term, it also adds a multiplicative scale term.
A linked loss plot shows the modified variant performing better, with the highlighted run reaching a lower eval loss than the baseline variants by 25k steps. The author describes it as a short, practical improvement with only a slight parameter change.
More from Research
- ArtiFixer to appear in SIGGRAPH Reconstruction session on Wednesday — ZGojcic · 2026-07-21
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21