Initial Guesses Matter for NLA
Turn_Trout · x · 2026-07-13
Natural Language Autoencoders (NLA) are fascinating: they take residual stream vectors and use natural language to explain what the model is "thinking."
This thread points out that training doesn't start from ground truth. Instead, Claude makes a "warm start" guess about the representations, followed by fine-tuning the encoder and decoder. The author's finding is that these initial guesses aren't just noise; they significantly impact the final training results.
More from Research
- New paper defines self-state attacks, showing OS defenses leave four agent-memory cases indistinguishable — Justgototheeffinmoon · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Animation shows how an MLP’s first-layer weights change while learning MNIST — CatAstro_Piyush · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22