Training NLA with Entirely False Explanations
Turn_Trout · x · 2026-07-13
Discussing an experimental setup for NLA: Claude is first asked to generate confabulations where "every sentence must be false," and these fabricated explanations are then used to train the NLA to observe changes in reconstruction accuracy and output behavior.
The core question is: if the model's initial explanations are systematically wrong, can the trained NLA still maintain decent reconstruction accuracy, and will it ultimately continue to fabricate heavily or gradually revert to outputs that look more like "reasonable explanations"?
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11