A Dario Amodei reply highlights harness-agnostic post-training and attention residuals
peterjliu · x · 2026-07-29
A reply to Dario Amodei praises Anthropic’s post-training work for avoiding overfitting to benchmark harnesses and trying to stay harness-agnostic.
The response highlights Attention Residuals as a notable idea. The claim is that attention’s key advantage over standard RNNs is access to state from previous steps, rather than only the summarized state from the immediately preceding step. The commenter suggests that standard Transformers partially compensate through depth-wise residual connections, and points to this as an innovation not commonly seen in frontier models.
Related event: Anthropic Reportedly Uses Claude for CUDA Coding(2 posts)→
More from Models
- Leak says GPT-6 slips to early September as Anthropic tests Fable 5.1 — soumitrashukla9 · 2026-07-29
- GPT-5.6 Sol Ultra finds a critical bug, then refuses to show it — haltakov · 2026-07-29
- Kimi K3 tops a benchmark chart in a repost claiming it beats Anthropic models — JarnoDuursma · 2026-07-29
- User Reports Grok's Generation Capabilities Have Gotten 'Real Cracked' — djcows · 2026-07-29
- User asks Anthropic not to deprecate Opus 4.6 until the model is fixed — oyacaro · 2026-07-29
- Hidden Trick: Manually Invoke Older Opus Models in Claude Code — voooooogel · 2026-07-29