Negative result: state-cosine adaptive depth stopping rule hurts both loss and passes

BlackHC · x · 2026-10-06

A reported negative result: after variable-depth training and repair, a state-cosine stopping rule for adaptive depth selects 4.75 passes on average with loss 3.3816, versus fixed depth 4 at loss 3.3761 — worse loss with more selected passes, and not measured FLOPs savings.

Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→

Original post →

More from Research

Research channel →