Da7em Bench: independent AI benchmark scores models on 200 real client tasks across 12 areas
airesearch12 · x · 2026-09-20
Da7 Tech launched Da7em Bench, billed as the first independent AI benchmark built on real client work.
- Each model runs 200 real tasks in each of 12 areas: reasoning, research, planning, delivery, persistence, accuracy, honesty, acceptance, engineering, taste, writing, and communication
- Models are tested across multiple harnesses (official and neutral: Droid, Hermes Agent, Devin, Cursor) so no single harness decides results
- Scoring is 1-5, where 5 means the work was accepted as delivered and 1 means failure; the bar is what a paid professional would deliver for the same brief
- Tasks stay private to prevent leaking into training data and inflating future scores
More from coding & agent
- AI designs a 4-layer PCB with 135 parts and 514 pads, human barely touched it — Paimaamu · 2026-09-20
- "AIs that won't use developers will be replaced by AIs that will," quips X user — ashishllm · 2026-09-20
- Agent product NoSpoon going private this month, creator says consumer AI market too early — Kyrannio · 2026-09-20
- Using Codex + GPT-6 Astra to plan art installations: rebuild the wall in Blender, skip recalculation — perilli · 2026-09-20
- What does an AI engineer's day actually look like? A learner asks Reddit — tech_kie · 2026-09-20
- Open-source local pipeline generates full radio dramas from one click — fflluuxxuuss · 2026-09-20