Terminal Bench 4.0: GLM-5.3, Qwen 3.8, Muse 1.3 and DeepSeek v4.1 Lead the Chart
himanshustwts · x · 2026-09-22
GLM-5.3, Qwen 3.8, Muse 1.3, and DeepSeek v4.1 lead the Terminal Bench 4.0 leaderboard. The author says they will put the models through the same set of evals for a head-to-head comparison.
More from coding & agent
- Qwen3.8-27B spent 21 days building its own CUDA engine on one RTX 3090 — DanGrover · 2026-09-22
- Ontology layer for agents boosts GPT-5.5 by 26.7 points on DDR-Bench — omarsar0 · 2026-09-22
- Dev wires typesafe's Jev into Apple's Foundation Models framework for Swift apps — rxwei · 2026-09-22
- LangSmith ships Jev-as-a-judge to score every production trace cheaply — airesearch12 · 2026-09-22
- Garry Tan says Capy handles large PRs faster than Codex or Claude Code — garrytan · 2026-09-22
- Building an image rating tool with GPT Vision and Jev: what worked and what didn't — huangyun_122 · 2026-09-22