Kimi K3, GPT-5.6 Sol and Fable 5 were tested on six newsroom jobs
local___host · reddit · 2026-07-24
A newsroom experiment compared Claude Fable 5, GPT-5.6 Sol, and Kimi K3 on six data-heavy reporting tasks with identical source packs and no web access.
Main findings
- Fable 5 acted like the strongest editor: it recovered reporting details and context others missed, but five of six stories exceeded the 600-word limit.
- GPT-5.6 Sol was the most compliant: it hit length and format requirements perfectly, but the prose was flatter and more list-like.
- Kimi K3 looked like an exhaustive reporter with a huge thinking budget: dense and well structured, but four of six stories also ran over the word ceiling.
Operational surprise
- At a 16,000-token completion cap, three of six runs spent the entire budget on reasoning before finishing.
- A Tour de France fixture even reasoned past a doubled cap before succeeding under bounded reasoning.
Charts and grounding
- The authors opened up the chart vocabulary and found it meaningfully changed editorial choices.
- Across the three frontier writers, 218 numeric chart values were traced back to source material with zero fabricated numbers in the audited numeric values.
- One caveat: an external audit found a wrong temperature baseline in a string label for one Sol chart, so the zero-fabrication claim applies only to numeric chart values.
More from Models
- LeCun on AI Security: First Autonomous Attack Used Closed Weights, Defended by Open — ylecun · 2026-07-24
- Rumor says Anthropic’s Opus 5 has slipped to tomorrow — mark_k · 2026-07-24
- Epoch AI Live-Streams GPT-5.6 Playing Slay the Spire — Jsevillamol · 2026-07-24
- ChatGPT Stuck for 10 Minutes: Long Context Threads Hit Stability Wall — billyjhowell · 2026-07-24
- Dev Projects Blocked by Safety Filters: A Push for Open Weights — VoidStateKate · 2026-07-24
- Reddit side-by-side test says GPT-5.6 SOL beats KIMI 3 on Chinese ink-wash animation — notNIHAL · 2026-07-24