New CET method traces how psychological constructs emerge across LLM layers

GolinoHudson · x · 2026-09-12

A new interpretability method, Construct Emergence Tracing (CET), combines dynamic exploratory graph analysis, NMI, and network complexity measures to track how multidimensional psychological constructs organize across LLM layers. Across 14 checkpoints (117M–32B) and 10.6M network estimates, construct recovery generally rose from early to intermediate states — but five models (GPT-2, Phi-4-mini, Qwen2.5-32B, GPT-OSS-20B, Muse-Glimmer-30B) showed near-zero final-state NMI, showing parameter count and last-layer extraction alone give incomplete pictures.

Original post →

More from Research

Research channel →