LakeQuest: A New Benchmark Evaluating Grounded QA over Heterogeneous Data Lakes

hllo_wrld · x · 2026-08-01

LakeQuest is a new benchmark designed to evaluate natural-language question answering systems grounded in heterogeneous data lakes. It covers three realistic domains: ML model cards, retail banking, and drug data, featuring human-validated questions where every answer points to the exact table row or passage supporting it.

This setup allows developers to distinguish whether a system failed to find the evidence or failed to use it correctly. The authors highlight a major takeaway: finding the evidence is usually no longer the bottleneck. Systems often retrieve the right policy or table but then misapply it. Furthermore, models can sometimes guess correct answers without any evidence, posing a significant issue for applications requiring strict audit trails.

Related event: LakeQuest: A New Benchmark for Messy Data Lake QA(2 posts)→

Original post →

More from Research

Research channel →