RLI Update: Fable5 Hits 16.1% Automation Rate

新智元 · wechat · 2026-07-12

The CAIS/Scale Remote Labor Index (RLI) released its latest results: Fable5 achieved an automation rate of 16.1% on real freelance projects—nearly double the second-place Opus 4.8 (8.3%) and higher than GPT-5.5 (6.3%). All three surpassed any previously evaluated models.\n\nRLI measures whether a model can deliver a "client-acceptable" result in a complete business commission, rather than just single-task performance. The benchmark includes 240 real Upwork projects across 23 fields, with a total value exceeding $144,000. Each project features a gold-standard deliverable from a human freelancer for comparison, judged on whether a "reasonable client would accept it."\n\nThe article further points out:\n- Just eight months ago, the top score on the leaderboard was only 2.5%, indicating rapid progress.\n- Fable5 utilizes a more robust agent framework, specifically a worker-critic loop, where an independent review agent repeatedly checks and rejects work for revisions.\n- Its per-project budget cap is higher ($150), compared to $50 for other models.\n- AI cannot simply replace human reviewers: automated grading significantly inflates scores for new models.\n\nCAIS concludes that while progress is rapid, the absolute level remains low, with 84% of real-world projects still falling outside the current capabilities of AI.

Original post →

More from Models

Models channel →