GPT-2 Internal States Fully Decoded
Revolutionary-Lab882 · reddit · 2026-07-14
The post introduces a project called BABEL codec, claiming to achieve a complete, verifiable decoding of a production language model's internal states (in this case, GPT-2 small).
Project claims include:
- The ability to "read" the model's internal states into English;
- The ability to "write" English back into the model's internal states;
- Approximately 94.7% of behaviors were reconstructed during testing;
- Results remained consistent across different layer depths and text scenarios.
The author emphasizes that the project is open-source, including the paper, complete dictionary, grammar tables, encoding/decoding weights, reproduction scripts, and a demo that visualizes the model's "thoughts." A GitHub link is also provided.
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22