A design-to-engineering team built an eval to stop AI copy from shipping on vibes
Fit_Average8352 · reddit · 2026-07-28
The author, who came into engineering from design, says generated interface copy is too often waved through on vibes: someone reads it, says “looks fine,” and it ships. They built a small evaluation rubric and scorer to replace that with something measurable.
The eval checks things like whether the copy uses the team's terminology instead of invented synonyms, whether reading level is in range, whether it matches voice on labeled examples, and whether it uses banned constructions such as “not just X, it's Y.” Anything below threshold is flagged for a human. It is still fuzzy and only catches the obvious misses, but it has already stopped the same copy mistakes from shipping repeatedly. They ask what other dimensions people would add.
More from coding & agent
- 12 real multi-app agent tasks put Fable 5 and Kimi K3 in a tie at 7/12, while GPT-5.6 Sol finished last — Nearby_Pair_6483 · 2026-07-28
- JarvisHub opens a canvas-native harness for building multimodal creative agents — Formal_Drop526 · 2026-07-28
- User plans to use Codex for everything this week as weekly usage limits appear in-app — rudrank · 2026-07-28
- Don't Be Fooled by Average Turns: Coding and Customer Service Agents Need Different Harnesses — hugobowne · 2026-07-28
- A user tests Qwen 3.8 Max inside Claude Code's ultra mode — jasonkneen · 2026-07-28
- Users want a local ChatGPT Live clone that can speak, listen and orchestrate agents — anonthatisopen · 2026-07-28