Negative result: state-cosine adaptive depth stopping rule hurts both loss and passes
BlackHC · x · 2026-10-06
A reported negative result: after variable-depth training and repair, a state-cosine stopping rule for adaptive depth selects 4.75 passes on average with loss 3.3816, versus fixed depth 4 at loss 3.3761 — worse loss with more selected passes, and not measured FLOPs savings.
Related event: Ex-DeepMind researcher BlackHC's tied Transformer with shared core KV cache(8 posts)→
More from Research
- Huawei Noah's Tail-Influence Sampling Cuts CVaR Policy Evaluation MSE by Up to 76% — huawei-noah · 2026-10-06
- Google's KeyRec Achieves Best Long-Video VLM Results With Just 10% of Visual Token Budget — google · 2026-10-06
- 4DCodeBench Shows Frontier Models Reconstruct Static Scenes but Fail at Dynamics — 4DCodeBench · 2026-10-06
- OmniTaskonomy: Year-long study shows generation training can improve understanding tasks — WeijiaShi2 · 2026-10-06
- Newton proved the product rule without limits, using a discrete symmetric-difference trick — ctjlewis · 2026-10-06
- DeepMind's AI designs enzymes from scratch: 99x drug building block yield, plastic-eating at 90°C — 141_1337 · 2026-10-06