hRoPE study shows compression depth, not location, signals real paragraph structure
Shuyang Xiang · hf · 2026-09-28
Standard positional encodings represent position as a 1D reading-order coordinate, missing hierarchical structure. This work proposes hierarchical RoPE (hRoPE) with separate channels for paragraph, sentence, and token indices.
Holding the token sequence fixed and intervening on the paragraph coordinate p1, the authors measure cross-paragraph attention with a token-distance-exact estimator. Attention compresses relative to a matched baseline—but compression alone isn't diagnostic: an identical channel with density-matched random labels also compresses, just more shallowly.
The reproducible signature of genuine structure is compression depth: deeper and corpus-dependent for real paragraph structure, absent in controls. Of eight corpus-only quantities tested (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest.
More from Research
- Category theory meets deep learning: PyNCD diagrams derive hardware-aware FlashAttention — GioeleZardini · 2026-09-28
- Interactive lab lets you pick the next token yourself to demo sampling, temperature, top-k, top-p — CurieuxExplorer · 2026-09-28
- Interactive Attention diagrams page gets major update, from scaled dot-product to DeepSeek latent attention — vtabbott_ · 2026-09-28
- EMNLP 2026 paper caught using a GPT-generated figure, review standards questioned — ATHii-127 · 2026-09-28
- Removing 10% of 'Important' Cells Cuts Model Output Nearly 3x, New Explainability Test Shows — bravo_abad · 2026-09-28
- IndicBankBench: 799-case benchmark shows banking AI assistants top out at 58.2% strict reliability — NPCI · 2026-09-28