Training Kimi on 1,465 PDF tasks more than doubles general task success
echen · x · 2026-09-30
New research asks: if you train on the underlying skill behind the widely used GDP.pdf eval, does the model just get better at PDFs — or learn something more fundamental?
The team post-trained Kimi K2.7 on 1,465 GDP.pdf companion tasks:
- GDP.pdf improved from 11.0% to 24.5%;
- On GDPval (tasks with no PDFs, in an environment with files, tools, spreadsheets and deliverables), full-task success more than doubled and 90% rubric pass rate rose +13.2pp.
Why did it transfer? Trajectory analysis shows that before training, Kimi often grabbed the first plausible source and started working; after training it looks around first — what's here, what matters, is anything missing? In one task, the base model opened the first weekly timesheet and prematurely computed a monthly report; the trained model correctly gathered all four weeks first.
The framing: a PDF is a small evidence environment — a good professional PDF forces you to figure out where facts live, which charts matter, and whether evidence suffices. Training on rich PDFs taught the model to gather all evidence before acting, a habit that persisted beyond PDFs.
Related event: Kimi Research: Training on 1,465 PDF Tasks Doubles General Capability(2 posts)→
More from Models
- OpenAI launches Codex Ultrafast: up to 8x faster than Astra Standard — OpenAIDevs · 2026-09-30
- Heavy user's math: mixing $100 Claude, $200 Codex, $30 Grok beats a single max plan — codtv132 · 2026-09-30
- Price of intelligence now halves every 2 weeks, leaving enterprise budgets guessing — gabriel1 · 2026-09-30
- OpenAI DevDay 2026: all 23 launches from GPT-6.1 Sol to Codex Security Cloud — reach_vb · 2026-09-30
- Reddit users report Codex connection errors while OpenAI status page shows no incident — ahriad · 2026-09-30
- OpenAI launches Ultrafast: up to 8x faster tokens at 300 tok/s in Codex — OpenAI · 2026-09-30