Distilled 120B Model Beats Kimi in Finance Reasoning at 1/60th Cost

ycombinator · x · 2026-07-31

Following Semafor's report questioning the nationality of an American model distilled from a Chinese one, Cyril Gorlla shared benchmark data. At the 8k token budgets typically used in production, their 120B model scores 83.61% on FinanceReasoning, outperforming Kimi K3 (81.93%) and Inkling (65.13%). Running on a single H100, it achieves this at 62 to 160x lower cost per query. However, at unlimited budgets, the big models still win on raw accuracy.

Original post →

More from Models

Models channel →