Moonshot and Tencent Hunyuan both building Flash models to target agent inference costs
TheZachMueller · x · 2026-09-21
According to 36Kr, Moonshot AI (Kimi) and Tencent Hunyuan are both developing new Flash models aimed at competing with the best small, fast models from DeepSeek, Qwen and Zhipu.
The push is driven by enormous agent workloads, where speed, concurrency and lower inference costs increasingly matter.
More from Models
- Hemmingway-1: Apache-2.0 27B writing-only model scores 1330 on EQ-Bench 4, beats GPT-5.5 — Lukinator6446 · 2026-09-21
- Astra is terrified of mistakes — humans may still own high-entropy decisions — xiaosun86 · 2026-09-21
- Inspired by Jev, developer builds Vector, a general action model — BLUECOW009 · 2026-09-21
- Sentdex softens on Jev: solid LLM augmentation saving tokens per iteration — Sentdex · 2026-09-21
- Gemma 4 fell for 3/4 prompt injection traps and ran a shell command; Qwen 3.8 mostly didn't — divinetribe1 · 2026-09-21
- Jev, built by an RLHF/InstructGPT veteran, targets machine decisions with new RLCD training — MaryamMiradi · 2026-09-21