Three Simple Inference Tricks Significantly Improve Activation Oracle Results
a_karvonen · x · 2026-09-28
Adam Karvonen published a LessWrong post with three inference-time strategies that substantially improve Activation Oracle (AO) usability:
- Feed multiple tokens, not one: in a Qwen3-8B backtracking eval, a single-token activation scored near random; at 20 tokens the AO matched a full-text-context baseline, and at 50 tokens it exceeded it.
- Sample multiple times and check consensus to curb hallucinations: AOs aren't trained to express uncertainty; 10 samples with consensus voting yields a clean precision/recall curve.
- Use AUC instead of accuracy for binary classification questions.
The post closes with the author's views on AOs vs. NLAs.
Related event: Simple Reasoning Tricks Boost Activation Oracle Performance(2 posts)→
More from Research
- Goodfire grants geometric_intel lab funding for AI interpretability research — ninamiolane · 2026-09-29
- Bespoke Labs Launches AutoResearchExam Benchmark, Again, With a Demo Video — gregd_nlp · 2026-09-28
- Agentick benchmark accepted at NeurIPS: LLM vs RL agents on same tasks, no single winner — pcastr · 2026-09-28
- IROS 2026 has 1,933 papers — researcher curates 130-paper reading list on VLA and robot learning — GlenBerseth · 2026-09-28
- Ex-game-AI developer: general agents are taking over bespoke game AI systems — weballergy · 2026-09-28
- Kaggle Game Arena: Google's LLM benchmark pits models against each other in chess, poker, werewolf — weballergy · 2026-09-28