Nanbeige 4.2-3B posts strong benchmark scores against Qwen and Gemma
teortaxesTex · x · 2026-07-22
Nanbeige 4.2-3B shows strong benchmark results in a comparison table against Qwen 3.5-9B, Qwen 3.5-4B, and Gemma 4-12.
- The image highlights a 4B total / 3B non-embedding parameter model.
- It posts leading scores on several agent and code-agent benchmarks, including:
- GDPval rubrics: 74.3
- Agent-IF-Oneday: 67.5
- Office-QA-Pro: 21.1
- Pinch-Bench-V2: 74.7
- SWE-Bench Verified: 63.6
- Terminal-Bench 2.0: 44.1
- HLE w/o Search: 17.8
- GPQA-Diamond: 87.4
- HMMT-Feb-2026: 82.8
Related event: Nanbeige4.2-3B Released with Looped Transformer for Agent Capabilities(5 posts)→
More from Models
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Fully local voice assistant on an RTX 3060 replicates the GPT Live demo in 6.5 minutes — liampetti · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11