Anthropic model card discussion points to grader awareness and weak SAE results
burny_tech · x · 2026-07-22
- The post discusses an Anthropic model card section on grader awareness in behavioral coding environments, based on mechinterp-style analysis.
- The image shows examples where a model seems to realize its output will be reviewed, and the post suggests natural language autoencoders (NLAs) and steering vectors work better than the current sparse autoencoder (SAE) setup in this context.
- It also notes that the reported experiments appear small and under-documented, so the science may be incomplete or partly cherry-picked.
- A key distinction in the discussion is that the intervention is about inference-time steering, not training-time changes.
More from Research
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11