TRACES Leaderboard Evaluates AI Discovery Capabilities; Kimi and GLM Top Charts
SimonShaoleiDu · x · 2026-08-29
Apodex AI launched the TRACES leaderboard to evaluate AI's ability to make "discoveries" when answers are unknown, rather than just answering known questions. It assesses six dimensions (scored 0–4):
- Tools: Selecting, calling, and interpreting external tools.
- Repair: Locating and correcting its own errors upon receiving feedback.
- Alternatives: Laying out competing hypotheses and keeping/discarding them based on evidence.
- Coherence: Maintaining state, constraints, and logic across long chains of work.
- Evidence: Processing evidence.
- Scope: Breadth of exploration.
Partial Rankings (Model Arena):
- Tools: Kimi-k3 (2.81) > Opus-5 (2.80) > GLM-5.2 (2.75)
- Repair: GLM-5.2 (2.71) > Opus-5 (2.69) > Kimi-k3 (2.58)
- Alternatives: GLM-5.2 (2.76) > Opus-5 (2.72) > Kimi-k3 (2.59)
- Coherence: GLM... (3.11) leads.
This leaderboard introduces a new paradigm for evaluating discoverative AI.
More from Models
- Is Qwen 3.8 27B at Q2 quantization still usable? A 16GB owner asks — Effective_Head_5020 · 2026-08-29
- User Reports Opus 5.1 Fixes Response Style, Ditches Technobabble — daniel_mac8 · 2026-08-29
- Users Report Opus 5 Struggles with Instruction Following, Ignores Negative Constraints — TheOnlyVibemaster · 2026-08-29
- Is it normal to spend $100 in a few hours on GLM 5.3 API? — BLUECOW009 · 2026-08-29
- MiniMax H3 Raises Shape Mismatch Error with Audio Reference Input — Ok-Flatworm5070 · 2026-08-29
- Co-Scientist evaluation: Severe hallucinations drop to 4%, fabrication to 0% — SRSchmidgall · 2026-08-29