Evals tell you it got worse, gates stop you shipping it: building an LLM gate in GitHub Actions
bgoncalves · x · 2026-09-26
A thread by @data4sci argues most LLM teams have evals but almost none have a gate: evals tell you something got worse, while a gate stops you from shipping it — with a walkthrough of building one in GitHub Actions. Reposter bgoncalves adds that too many teams discover bad prompt changes from user complaints instead of CI, and that a boring gate in the merge path is the fix that works.
More from coding & agent
- Boris Cherny's viral tweet puts TLA+ and formal methods in the agent-coding spotlight — fhuszar · 2026-09-26
- Two local AIs talk to each other with no cloud: hands-on with Braid — Scobleizer · 2026-09-26
- Open-source EvoOntology lets agents build and evolve a data ontology via MCP — TheTuringPost · 2026-09-26
- W&B's ARIA agent runs 200+ autoresearch experiments, nearly beats its best result live — AI Engineer · 2026-09-26
- Making GA4 Agent-Friendly: Building Analytics Tools for an MCP Server — DutchSEOnerd · 2026-09-26
- Loop engineering beats prompt engineering: build agent loops that self-verify — goyalshaliniuk · 2026-09-26