Kimi K3 reveals its pre-training mix across text, code, math, knowledge, and vision
stochasticchasm · x · 2026-07-28
Kimi K3’s pre-training data is described as a curated mix of four main text domains — web text, code, mathematics, and knowledge — plus a large-scale vision corpus.
- Text pipeline: each domain is filtered with rule-based heuristics, classifier scoring, and deduplication, with domain-specific sampling rates chosen from ablations on smaller models.
- Knowledge/math rephrasing: following the Kimi K2 recipe, the team uses style- and perspective-diverse prompts, chunk-wise autoregressive generation, and fidelity checks against source documents.
- Vision pipeline: the vision corpus combines open-source data with in-house filtering, synthesis, and deduplication. Training uses coordinate supervision in both absolute and normalized formats for robust localization.
- Multimodal scale-up: beyond captioned images, the corpus includes code snippets paired with rendered visuals across SVG, 3D assets, webpages, games, and CAD schematics.
Related event: Inside Kimi K3: Tri-axis Architecture and Hybrid Attention(34 posts)→
More from Models
- A repost claims Anthropic’s Opus 5 regresses badly despite benchmark gains — rickasaurus · 2026-07-28
- Polymarket prices a 76% chance Moonshot ships another Kimi K model by September — Polymarket · 2026-07-28
- The Verge says Moonshot’s open Kimi K3 could undercut closed U.S. AI models — The Verge AI · 2026-07-28
- Kimi K3 lands on Fireworks AI for inference and training — omarsar0 · 2026-07-28
- Gemini video generation adds words and blocks some harmless prompts — Individual-Cookie615 · 2026-07-28
- Polymarket now prices a 42% chance of a new Claude Sonnet by next month — Polymarket · 2026-07-28