Three simple inference tricks make Activation Oracles far better: multi-token context beats single-token activations
a_karvonen · x · 2026-09-28
AI safety researcher akarvonen shares three simple inference strategies for Activation Oracles that significantly improve performance:
- Feed multiple tokens of context instead of a single token's activation
- On a backtracking eval, an AO was at random chance with one token but did far better with just 5 or 20 tokens
- The intuition: information about complex concepts may be spread across multiple tokens' activations
Directly useful for interpretability and model internals monitoring work.
Related event: Simple Reasoning Tricks Boost Activation Oracle Performance(2 posts)→
More from Research
- Stanford prof backs proposal for official LLM reviews in peer review, humans just calibrate — anshulkundaje · 2026-09-29
- Three tiny MLPs on frozen Gemma 4 hit 95.8% turn-taking accuracy at near-zero cost — cortexist · 2026-09-29
- Speech-guided multimodal learning for vocal tract segmentation presented at MICCAI 2026 — maier_ak · 2026-09-29
- Pachocki, Hinton, Bengio & Jack Clark co-author paper on AI R&D automation triggering intelligence explosion — birchlse · 2026-09-29
- AI workflow flags discrepancies in 3,460 of 4,452 econ replication packages across five journals — Afinetheorem · 2026-09-29
- Schmidt Sciences hiring Fellows in Residence with $50k compute and $50k workshop budget — peterbhase · 2026-09-29