Wikidata Search Traces: 10,235 traces for training knowledge graph search agents
omarsar0 · x · 2026-10-08
A new arXiv paper, "Wikidata Search Traces," releases a dataset for training knowledge graph search agents, teaching models to navigate toward answers instead of relying on memory.
- Answering complex questions over Wikidata requires SPARQL queries naming the right entities and relations; LLMs instead answer largely from memory, which is least reliable for less prominent entities.
- Two obstacles: no training data recording how a solver explores the graph, and interfaces that dump large graph results into the model's context.
- Three validated hypotheses: graph search difficulty can be controlled via the nested structure of a question; long-horizon failures stem mostly from evidence management rather than the model; open-weight models can match commercial closed ones in a suitable environment.
- Released: 10,235 solving traces over single-entity and multi-hop questions, plus a recursive language model (RLM) harness built on persistent graph state — keep evidence outside the context window, then selectively pull in what matters.
More from Research
- Full Slides Released for ECCV'26 Tutorial on Diffusion Model Post-Training and Alignment — CSProfKGD · 2026-10-08
- Norvig's Classic Essay on Chomsky and the Two Cultures of Statistical Learning Still Reads Fresh in the LLM Era — 3scorciav · 2026-10-08
- DatologyAI open-sources Zephon, cutting data-order noise from 0.82 to 0.05 points when GPU count changes — lmoroney · 2026-10-08
- ProactiveCoach: Hierarchical Guidance Boosts Proactive AI Assistants by 57.1 Points — skku · 2026-10-08
- STEPQuant: 6-bit quantization of Delta-rule recurrent states cuts serving memory by up to 68.7% — zju-community · 2026-10-08
- RoboQuest Benchmark: Best Multimodal Agent Succeeds in Only 23% of Embodied Exploration Tasks — declare-lab · 2026-10-08