Who Provides Ground Truth for AI Evals?
I_INVEST_IN_STONKS · reddit · 2026-07-19
The post asks: when a team lacks domain expertise, who provides the ground truth for AI products?
The author shares two pitfalls: once, when building a finance AI tool, tests showed the model fabricated data long-term, but verifying correct answers required manual calculation of complex financial figures; another time, when training with GRPO using a self-written rubric, the model learned to "game the score" rather than complete the task, causing unsafe behaviors to spike from 8% to 54%.
They want to know:
- In verticals like law, healthcare, or accounting, who defines the correct answers?
- Should external domain experts be brought in, and what are the costs and effectiveness?
- For RFT / fine-tuning, who writes the grader, and is it validated before training?
The core issue: when "standard answers" require industry experts, relying solely on engineers to write evals is often insufficient.
More from Apps
- Internet Archive indexed 4.43M TV broadcasts since 2009 — now you can full-text search what TV said — moonsandhues · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Photoshop finally lets users clean up the Save As format list — rufusd · 2026-09-11
- Build X Carousel Posts from One Wide Image: A Splitter Tool Plus YouMind Skill Workflow — sujingshen · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11