Testing Kimi K3's 32-Page Chain of Thought
emollick · x · 2026-07-19
While testing Kimi K3, Ethan Mollick asked it to pick two poems that best reflect the "current state of GenAI models."
Two notable takeaways from the results:
- The chain of thought (CoT) spanned a massive 32 pages, and the content itself was "quite interesting."
- However, it also exhibited many loops and dead ends, which is one of K3's typical behaviors.
In the attached image, the author compares several candidate poems one by one, discussing which imagery best maps to LLMs—such as "mirrors," "fables," "automata," and "combinatorial rearrangements"—focusing on finding poetic facets that reflect model behavior.
Related event: Testing Kimi K3: 32-Page CoT Loops and English-Dominant Reasoning(5 posts)→
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11