LakeQuest QA benchmark, testing RAG on messy enterprise tables and docs, hits COLM
hllo_wrld · x · 2026-10-06
The authors will present LakeQuest (paper #1073) at COLM tomorrow — poster session Tuesday morning at 11am, Franciscan room, poster F-1-135.
Core thesis: QA systems excel on benchmarks built from clean, curated text, but real knowledge in companies and labs lives scattered across tables, documents, and metadata that barely link together. LakeQuest is a new benchmark designed to test QA in exactly that messy, realistic setting.
More from Research
- Why User-Model Evals Are Hard: Stanford Researchers Bet on a Multi-User Turing Test — alexisjross · 2026-10-06
- Embedding Every Font with Neural Networks Yields a Flower-Shaped Map of Google Fonts — Chroma-Crash · 2026-10-06
- Crawler Zoo Launches a Free Arena for Testing Local-Model Agents — Time_Instruction_955 · 2026-10-06
- Trained agentic context management: 8K-context small model matches GPT-5.4 at 1M on OOLONG — xennygrimmato_ · 2026-10-06
- User Sim Index is broken: trivial bot scores 95% across behavioral dims — ericzelikman · 2026-10-06
- Used OpenAI Dots as a Free Agent Swarm to Break a 47-Year-Old Math Record — jaxchang · 2026-10-06