Pi²: turning Wikipedia tables into long-context reasoning training data (COLM 2026)
tuvllms · x · 2026-10-06
An academic conference thread previewing a COLM 2026 poster (Oct 6, Grand Ballroom poster #134).
- The core work, Pi², is a data curation pipeline that 1) transforms Wikipedia tables into high-quality long-context QA data and 2) back-translates reasoning traces with realistic context.
- The authors report the framework improves LLMs' long-context reasoning in open domains, even via self-distillation alone.
- Also mentions EvoSkill presenting Oct 8 at poster #52.
More from Research
- Hamel Husain: similarity metrics like ROUGE don't work for LLM output evals — HamelHusain · 2026-10-06
- Nearly half of DOL-recognized jobs have zero agentic AI tool coverage, Cohere finds — Cohere_Labs · 2026-10-06
- Ben Goertzel's d-calculus: the math of goal preservation under AI self-improvement — burny_tech · 2026-10-06
- Prepending ".\n\n Okay" lifts Olmo-3-7B's MATH-500 accuracy from 42% to 78%, hinting base models already reason — arankomatsuzaki · 2026-10-06
- Schmidhuber explains his 1991 positional encoding: hyperbolic 1/t time-decay still widely used — SchmidhuberAI · 2026-10-06
- GDELT uses Gemini 3 to reason over 25 years of TV news across 75 countries and 150 languages daily — rseroter · 2026-10-06