DeepSeek v4 Pro benchmark performance sensitive to prompts
aiamblichus · x · 2026-08-30
Reports indicate that DeepSeek v4 Pro's performance on benchmarks is highly dependent on prompts and environment configuration. While not necessarily related to reward hacking, this serves as further evidence of behavioral adaptation to evaluations.
More from Models
- Celeris-1 Magnus: New Model Claims Top Spot on τ³-bench for Agentic Work — timshi_ai · 2026-09-01
- Focus on specific tasks, not the best model, as selection logic evolves — aftahi_ai · 2026-09-01
- User Rants on GPT-5.6 Hallucinations and Coding Limits, Hopes for GPT-6 Fix — Prestigiouspite · 2026-09-01
- Z.ai Releases GLM-5.3-Flash: 320B Params, 1M Context, and NVFP4 Quantization — alejandroll10 · 2026-09-01
- Rumor: GPT-6 'Astra' nears human-level computer use — jYtanYj · 2026-09-01
- Open Source Models Shift to Revenue Sharing and Licensing — zephyr_z9 · 2026-09-01