Training NLA with Entirely False Explanations
Turn_Trout · x · 2026-07-13
Discussing an experimental setup for NLA: Claude is first asked to generate confabulations where "every sentence must be false," and these fabricated explanations are then used to train the NLA to observe changes in reconstruction accuracy and output behavior.
The core question is: if the model's initial explanations are systematically wrong, can the trained NLA still maintain decent reconstruction accuracy, and will it ultimately continue to fabricate heavily or gradually revert to outputs that look more like "reasonable explanations"?
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22