Deep Learning Dimensional Rule: Tensors with Three Large Axes Won't Fly
francoisfleuret · x · 2026-08-24
Deep learning expert François Fleuret shares an intuitive rule of thumb regarding tensor dimensions:
Core Point
In deep learning, you never encounter three "large" axes simultaneously.
Common Dimensional Patterns
- Activations: $N imes D$ (samples $ imes$ features) or $T imes D$ (timesteps $ imes$ features)
- Parameters: $D imes D'$ (input features $ imes$ output features)
- Gradients: $N imes D$
- Attention: $T imes T$
Implication
If your method requires manipulating tensors where all three dimensions are large, it is unlikely to be computationally feasible or efficient.
More from Research
- RL Does More Than Sharpen SFT: Longer Pretraining Yields Steeper RL Scaling Slope — gleech · 2026-08-24
- Reservoir & Photonic Computing: Escaping the O(n²) Trap of LLMs — navnt5 · 2026-08-24
- Open source Pelican SVG Env quantifies model drawing abilities — SergioPaniego · 2026-08-24
- Beyond Final Weights: The Value of Stage-Level Checkpoint Transparency — creditme7 · 2026-08-24
- 17-year-old who independently pretrained models gets ICML paper accepted — HanchungLee · 2026-08-24
- CAS-Spawned ScienceClaw Launches AutoProject: AI Research Agents Move From Tasks to Whole Projects — 量子位 · 2026-08-24