Kimi K3 Beats Opus 4.8 in 34-Prompt Oneshot Eval at 1/16 the Cost
kms_dev · reddit · 2026-07-31
A developer ran 34 oneshot prompts through both Kimi K3 and Opus 4.8, evaluating the generated HTML, screenshots, and GIFs using Sonnet 4.6. The results showed Kimi K3 performing better than Opus 4.8.
Furthermore, Kimi K3 proved highly token-efficient, costing only $0.44 to process all 34 prompts compared to $7.16 spent by Opus.
More from Models
- Teknium Tests DeepSeek V4 Flash: Full Agent Task Costs Just $0.07 — Teknium · 2026-08-01
- xAI Launches Grok Build CLI Coding Agent Powered by New Grok 4.5 — belce_dogru · 2026-08-01
- Claude 4 Fails Long-Context Retrieval, Suspected KV Compression Artifacts — teortaxesTex · 2026-08-01
- Light-MER: Sub-1B Parameter Model Outperforms 8B Teacher in Multimodal Emotion Recognition — 新智元 · 2026-08-01
- OpenRouter Launches Kimi K3 Nitro Mode for Maximum Token Throughput — charles_irl · 2026-08-01
- Claude Opus Extreme Test: Zooming to Atomic Level Burns $33.78 in One Prompt — arena · 2026-08-01