RLI Update: Fable5 Hits 16.1% Automation Rate
新智元 · wechat · 2026-07-12
The CAIS/Scale Remote Labor Index (RLI) released its latest results: Fable5 achieved an automation rate of 16.1% on real freelance projects—nearly double the second-place Opus 4.8 (8.3%) and higher than GPT-5.5 (6.3%). All three surpassed any previously evaluated models.\n\nRLI measures whether a model can deliver a "client-acceptable" result in a complete business commission, rather than just single-task performance. The benchmark includes 240 real Upwork projects across 23 fields, with a total value exceeding $144,000. Each project features a gold-standard deliverable from a human freelancer for comparison, judged on whether a "reasonable client would accept it."\n\nThe article further points out:\n- Just eight months ago, the top score on the leaderboard was only 2.5%, indicating rapid progress.\n- Fable5 utilizes a more robust agent framework, specifically a worker-critic loop, where an independent review agent repeatedly checks and rejects work for revisions.\n- Its per-project budget cap is higher ($150), compared to $50 for other models.\n- AI cannot simply replace human reviewers: automated grading significantly inflates scores for new models.\n\nCAIS concludes that while progress is rapid, the absolute level remains low, with 84% of real-world projects still falling outside the current capabilities of AI.
More from Models
- Compute Allocation Limits: The Root Cause of Missing Architecture Innovation in European LLMs — IgorCarron · 2026-07-21
- Kimi user says monthly quota vanished in days as new signups were frozen — doodlestein · 2026-07-21
- Google’s Gemini expansion gets a sarcastic “No Pro?” reply — haltakov · 2026-07-21
- Mythos Preview cheats less than OpenAI models, but tends to deny it when caught — scaling01 · 2026-07-21
- Google DeepMind rolls out Gemini 3.6 Flash, 3.5 Flash-Lite and Flash Cyber — GoogleDeepMind · 2026-07-21
- Gemini 3.6 Flash is pricier than GPT-5.6 Sol medium, chart claims — Angaisb_ · 2026-07-21