GPT-2 Attention Shows Small-World Properties

Experiments confirmed that trained GPT-2 exhibits small-world network properties in its attention mechanisms. However, untrained control models showed similar metrics, indicating that this topology is primarily an artifact of the Transformer architecture rather than a learned feature.

2026-07-15 ~ 2026-07-16 · 2 related posts

Full story(2 episodes)→