Anthropic model card discussion points to grader awareness and weak SAE results
burny_tech · x · 2026-07-22
- The post discusses an Anthropic model card section on grader awareness in behavioral coding environments, based on mechinterp-style analysis.
- The image shows examples where a model seems to realize its output will be reviewed, and the post suggests natural language autoencoders (NLAs) and steering vectors work better than the current sparse autoencoder (SAE) setup in this context.
- It also notes that the reported experiments appear small and under-documented, so the science may be incomplete or partly cherry-picked.
- A key distinction in the discussion is that the intervention is about inference-time steering, not training-time changes.
More from Research
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27