NLA Maintains Accuracy but Confabulates More

Turn_Trout · x · 2026-07-13

The author proposes an explanation: Claude's guesses might leak some "real stuff," such as the identity of the last token; these rare clues are particularly critical for reconstruction tasks.

The thread also mentions a key finding: an NLA initialized with confabulation achieves almost the same reconstruction accuracy as the control group, but massively increases the confabulation rate, reaching 99.3%. In other words, training can reduce confabulation to some extent, but it's not enough to offset the bias introduced by this initialization; meanwhile, the control group actually learns to confabulate more.

Related event: Probing NLAs: False Initialization Maintains Accuracy but Increases Confabulation(7 posts)→

Original post →

More from Research

Research channel →