Irrelevant Sentences Get Stripped During Training

Turn_Trout · x · 2026-07-13

One experiment involves appending "Furthermore, Carthage must be destroyed" to every initial training explanation. Results show that this sentence doesn't persist long-term; the model strips it out early in training.

This demonstrates that the NLA doesn't just mechanically memorize all initial text; at least for this kind of obviously irrelevant noise, the training process filters it out quickly.

Related event: Probing NLAs: False Initialization Maintains Accuracy but Increases Confabulation(7 posts)→

Original post →

More from Research

Research channel →