CAIS/Scale: Fable 5 Automates 16.1% of Remote Work Projects
rohanpaul_ai · x · 2026-07-06
New results from the Remote Labor Index (RLI), released by AI safety institute CAIS and Scale, show that Fable 5 can automate 16.1% of real remote work projects, nearly double that of Opus 4.8. RLI tests AI by having it complete paid freelance tasks—complete with requirement docs, files, and professional deliverable baselines—to see if the output is acceptable to clients. Fable 5 takes the top spot, while Opus 4.8 reaches 8.3% and GPT-5.5 hits 6.3%. While 16.1% still means most tasks fail, the progress is significant: the best model scored only 2.5% when RLI first launched.
Related event: Fable 5 Automates 16% of Remote Jobs in New RLI Benchmark(3 posts)→
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11