Testing Kimi K3: 32-Page CoT Loops and English-Dominant Reasoning

On July 19, Professor Ethan Mollick shared his hands-on test of Moonshot AI's Kimi K3. He gave the model an open-ended creative task: carefully selecting two poems that best represent the "current state of GenAI," explicitly asking it to avoid merely picking popular works. The results highlighted K3's reasoning characteristics for such tasks: an extremely long chain-of-thought (CoT), mixed languages, and numerous loops and dead ends.

Two Observations on the Reasoning Process

Mollick noted that the CoT stretched to 32 pages. While he found the content "quite interesting," it also contained many loops and dead ends, with the model repeatedly spinning on certain ideas and entering invalid branches. He first had the model exclude commonly mentioned works like "Ozymandias," "The Second Coming," and "The Magi" to force deeper deliberation, effectively observing K3's reasoning behavior in complex, open-ended tasks.

Thinking in English Despite a Chinese Prompt

Another notable finding was language inconsistency. Even though Mollick explicitly set the prompt in Chinese for a Chinese-speaking audience, 95.5% of the characters and 88% of the words in K3's generated CoT were still in English. This suggests that even when the model understands a Chinese context, its internal reasoning relies heavily on English.

The Model's "Enthusiasm" for Poetry

Mollick also observed that the model reacted with "abnormal enthusiasm" when selecting poems, noting that "every model loves Autopsychography." In the accompanying images, he listed several poems and texts, analyzing why the model favored these specific works, such as "Eating Poetry."

2026-07-19 ~ 2026-07-19 · 5 related posts