Leaked Tasks Hint at Anthropic's Strategy: Training Expert Judge Models from Human Traces

burny_tech · x · 2026-07-30

Based on recently leaked Anthropic annotation tasks, developers speculate that Anthropic is training an expert judge model using human traces. This model could then dynamically generate rubrics at inference time to evaluate agent performance.

This mechanism closely mirrors the reinforcement learning (RL) strategy used by Kimi K3 for non-verifiable tasks, where an LLM acts as a judge to generate rubrics and score candidate outputs on the fly. The discussion suggests that Generative Reward Models (GRM) are becoming an industry standard for scaling model capabilities.

Original post →

More from Models

Models channel →