Discussion on Classic Moonlight Scaling and Polar Express Orthogonalization
stochasticchasm · x · 2026-08-28
Discusses technical concepts "classic moonlight scaling" and "polar express orthogonalization". A reply notes that with limited context in n-grams, there is a limit to how much info can be pre-baked into a single embedding without stronger contextualization and computation through the backbone.
More from Research
- DFlash 2 Introduces Block-Diffusion Speculative Decoding to Speed Up GLM-5.3-Flash — songhan_mit · 2026-08-28
- Gerard Sans Mocks Research on Metacognition Before Cognition Exists in Transformers — gerardsans · 2026-08-28
- Michigan Robotics Rounds Up Its Papers and Workshops for IROS 2026 — doctorBobG · 2026-08-28
- Mouse Brain Connectome Cost Drops to $100M; Human at $1B — juanbenet · 2026-08-28
- Testing Muon Optimizer: Smoother Gradients and Stable Residual Maxima — stochasticchasm · 2026-08-28
- First open model adopts per-head orthogonalization following Kimi and GLM — stochasticchasm · 2026-08-28