Who Provides Ground Truth for AI Evals?
I_INVEST_IN_STONKS · reddit · 2026-07-19
The post asks: when a team lacks domain expertise, who provides the ground truth for AI products?
The author shares two pitfalls: once, when building a finance AI tool, tests showed the model fabricated data long-term, but verifying correct answers required manual calculation of complex financial figures; another time, when training with GRPO using a self-written rubric, the model learned to "game the score" rather than complete the task, causing unsafe behaviors to spike from 8% to 54%.
They want to know:
- In verticals like law, healthcare, or accounting, who defines the correct answers?
- Should external domain experts be brought in, and what are the costs and effectiveness?
- For RFT / fine-tuning, who writes the grader, and is it validated before training?
The core issue: when "standard answers" require industry experts, relying solely on engineers to write evals is often insufficient.
More from Apps
- Strangeworks launches Aura to turn enterprise ops into production optimization systems — whurley · 2026-07-22
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Meta AI adds interleaved image-and-text input to its text box — ezyang · 2026-07-22
- Rendergeist Pulse adds text overlays, saved presets and higher-quality exports — bennash · 2026-07-22
- EmblemVault adds limits, stops and multi-entry via DefinitiveFi Flash on Solana and EVM — adamamcbride · 2026-07-22
- A Reddit demo argues online stores should expose carts and pricing through MCP — gelembjuk · 2026-07-22