A Dario Amodei reply highlights harness-agnostic post-training and attention residuals

peterjliu · x · 2026-07-29

A reply to Dario Amodei praises Anthropic’s post-training work for avoiding overfitting to benchmark harnesses and trying to stay harness-agnostic.

The response highlights Attention Residuals as a notable idea. The claim is that attention’s key advantage over standard RNNs is access to state from previous steps, rather than only the summarized state from the immediately preceding step. The commenter suggests that standard Transformers partially compensate through depth-wise residual connections, and points to this as an innovation not commonly seen in frontier models.

Related event: Anthropic Reportedly Uses Claude for CUDA Coding(2 posts)→

Original post →

More from Models

Models channel →