GPT-2 Small-World Topology Likely an Architectural Artifact

its_vayishu · x · 2026-07-15

Upon reproduction, the author found that a trained GPT-2 does resemble a small-world network by this metric.

However, the untrained control group yielded similar results (σ ≈ 4.6, compared to ≈ 5 post-training).

They repeated the test across 5 sparsity levels, and the gap remained negligible, suggesting this is a result of the architecture itself rather than something "learned" during training.

Conclusions:

Related event: GPT-2 Attention Shows Small-World Properties(2 posts)→

Original post →

More from Research

Research channel →