GBAG-Bench finds local BI models can write correct SQL and still hallucinate growth
Additional_Menu8542 · reddit · 2026-07-28
A local BI tool author open-sourced GBAG-Bench (Grounded BI Answer Generation), a benchmark focused on a gap most SQL evaluations ignore: whether a model turns the right query result into a faithful natural-language conclusion.
- The author found a model that wrote correct SQL against 136 months of sales data, then summarized the result as “+48% growth” even though the business had been flat for ten years.
- GBAG-Bench is MIT licensed and includes 35 questions across Sakila, Chinook, and Northwind.
- The benchmark scores faithfulness, completeness, and insight rather than only SQL correctness.
- Key finding: the failure often looks like an aggregation deficit, not a comprehension deficit. Precomputing aggregates and injecting them into context improved outputs.
- A second judge showed the headline improvement was partly driven by judge severity; completeness gains were robust, but faithfulness gains were less stable.
More from Research
- World Labs says R2S2R can make robot training and evaluation far cheaper — drfeifei · 2026-07-29
- World Labs says spatial intelligence can build worlds that train robots — drfeifei · 2026-07-29
- A research thread argues most papers are wrong and should not be read literally — RexDouglass · 2026-07-29
- Pangram teases a new model release tomorrow with sentence-level boundary detection — cephaloform · 2026-07-29
- Declassified 1966 Shakey report shows a robot that planned its own path — frankreddit5 · 2026-07-29
- User modeling discussion argues personalization should go beyond surface style — EchoShao8899 · 2026-07-29