LLM Judge community spent two years writing monolithic prompts until autorubric, says Delip Rao
deliprao · x · 2026-09-21
Delip Rao criticizes the LLM-as-Judge community for writing monolithic prompts for two years until autorubric introduced systematic rubric decomposition — despite social science methods literature covering this long ago. Shriram Murthi adds that Jev's real edge is parallelism and speed from architectural differences, likened to "System 1".
Related event: Delip Rao Slams LLM-as-Judge Community for Monolithic Prompts(2 posts)→
More from coding & agent
- Arena: coding-agent harnesses show up to 5x cost differences at similar success rates — thione · 2026-09-21
- Recap: Anthropic merged Claude chat and Cowork, redesigned Claude Code Projects — thione · 2026-09-21
- Recap repeat: Claude Code Projects parallel cloud sessions and Astra for Law — thione · 2026-09-21
- Recap: WorkOS explains agent auth with auth.md; Salesforce and Anthropic expand Claudeforce — thione · 2026-09-21
- Recap repeat: DeepMind Dream-RSI cuts discovery-agent calls up to 162x — thione · 2026-09-21
- MCP server brings GitHub Actions log reading and CI/CD control to AI agents — modelcontextprotocol · 2026-09-21