DeepSeek V4-Flash 性能暴涨,或因 V4-Pro 充当 RL 教师
teortaxesTex · x · 2026-07-31
针对 DeepSeek V4-Flash 性能大幅提升的现象,有观点推测 DeepSeek 可能基于更强大的 V4-Pro 模型创建了强化学习(RL)教师模型,从而带动了 Flash 版本的性能飞跃。
不过,社区对于 V4-Pro 基座模型的质量及其能力上限仍存在疑虑,并期待其正式版能有更惊艳的表现。
「模型」频道最新
- Gemini Flash 低价被指堪比 DeepSeek 时刻 — eyishazyer · 2026-07-31
- DeepSeek 展现自主拆解能力,无需提示即用 subagent — teortaxesTex · 2026-07-31
- DeepSeek V4长上下文表现受限,全压缩层设计或是短板 — bookwormengr · 2026-07-31
- 同价位大模型实测:DeepSeek V4-Flash 多模态胜过 GPT-5.6 Luna — teortaxesTex · 2026-07-31
- DeepSeek 推理 RL 训练受赞,规避模型废话与幻觉 — teortaxesTex · 2026-07-31
- 实测对比:o3 在 OSINT 任务中表现依然强悍 — bytebot · 2026-07-31