HSTA: mapping tech diffusion via semantic trajectories on 30k preprints and patents, LLMs evolving fastest
Muhammad Sukri Bin Ramli · hf · 2026-09-30
A paper introduces Hyperspherical Semantic Trajectory Analysis (HSTA), an unsupervised quantitative method that tracks technology diffusion directly from unstructured scientific and commercial text streams, complementing macro metrics like TFP that lag by years.
Data and method:
- 30,000 filtered records spanning arXiv preprints and USPTO patent applications
- Transformer sentence embeddings projected onto unit hyperspheres, clustered via Spherical K-Means into eight sub-topics with UMAP reductions
- Two formalized metrics: Semantic Centroid Vector Drift (vocabulary/paradigm shifts) and Commercialization Offset (peak-density alignment between discovery and IP filings)
Findings:
- Linking quarterly topic velocity with Epoch AI hardware data, VAR F-tests show paper volume velocity alone does not Granger-cause frontier compute surges at conventional significance — textual signals must be conditioned on physical capital
- LLM sub-topics show the highest semantic drift (0.332), followed by AI systems (0.234)
More from AGI Musings
- Researcher revisits Dario's 2019 LLM alignment bet — and argues it's starting to fail — birchlse · 2026-09-30
- Anthropic's Krieger said Claude replaced a PM — the PM proved him 100% wrong — victor_explore · 2026-09-30
- CS scholar on AI-written drafts: if you outsource your voice, don't expect humans to review it — birchlse · 2026-09-30
- Oscar-winning writer Roger Avary discusses AI's disruptive frontier of cinema on podcast — Kyrannio · 2026-09-30
- Alignment debate: critic says Yudkowsky repeats the mistake of designing minds instead of growing them — repligate · 2026-09-30
- Continuity isn't persistent cognitive state: Glen Bradley on what AI memory misses — GlenBradley · 2026-09-30