Kimi K3 Tops Spreadsheet Benchmark

rohanpaul_ai · x · 2026-07-18

Kimi K3 took first place on SpreadsheetBench 2, described as particularly useful for spreadsheet workflows, especially in finance, planning, and operations scenarios. Rather than just testing isolated formula capabilities, this benchmark evaluates whether tool-based AI can complete comprehensive spreadsheet tasks. It includes 321 expert-designed tasks covering financial modeling, workbook debugging, and native chart creation. On average, tasks involve 11.8 worksheets and 593.5 cell modifications, requiring the model to maintain context consistency over long chains of operations. The post also details the scoring method: - Modeling and debugging tasks require all specified modifications to be accurate, while untargeted cells must remain unchanged - Chart tasks are only considered passed if a vision model confirms at least 70% of the requirements are met Quoted content adds that Kimi K3 surpassed Claude Fable 5, marking a case where an open-weight model defeated all closed-source counterparts.

Related event: Kimi K3 Tops SpreadsheetBench 2 and Shows Strong KernelBench Results(4 posts)→

Original post →

More from Models

Models channel →