CodeGraph builds an open-taxonomy knowledge graph over 167M source files via code LLM
Federico Pennino · hf · 2026-09-28
A team led by Federico Pennino released CodeGraph, a large-scale open-taxonomy knowledge graph for source code.
- Motivation: Public repositories hold billions of files, but existing tools are limited to syntactic and token-level analysis, missing the implicit engineering knowledge (algorithms, paradigms, design patterns, domains).
- Method: A code-specialized LLM performs semantic annotation, with a three-stage Wikidata grounding pipeline: deterministic SPARQL for unambiguous entities, a Deep Research Agent for the long tail, and hierarchy-rollup to import parent-of closures.
- Scale: Applied to 167M files of the Stack-Edu corpus, yielding 158M nodes, 1B typed edges, 63K concept entities, and 19.8K grounded Wikidata entities across 14 programming languages.
- QA: A calibrated protocol combining a small human gold set with an LLM-as-a-judge filter quantifies annotation precision.
More from Research
- Goodfire grants geometric_intel lab funding for AI interpretability research — ninamiolane · 2026-09-29
- Bespoke Labs Launches AutoResearchExam Benchmark, Again, With a Demo Video — gregd_nlp · 2026-09-28
- Agentick benchmark accepted at NeurIPS: LLM vs RL agents on same tasks, no single winner — pcastr · 2026-09-28
- IROS 2026 has 1,933 papers — researcher curates 130-paper reading list on VLA and robot learning — GlenBerseth · 2026-09-28
- Ex-game-AI developer: general agents are taking over bespoke game AI systems — weballergy · 2026-09-28
- Kaggle Game Arena: Google's LLM benchmark pits models against each other in chess, poker, werewolf — weballergy · 2026-09-28