Ox Alpha benchmark test: 551 API calls needed to get 87 completed answers
anshulkundaje · x · 2026-08-23
@sbatzoglou evaluated Ox Alpha on the induction benchmark (ICML 2026 spotlight paper "INDUCTION: Finite-Structure Concept Synthesis in First-Order Logic") via OpenRouter:
- Mediocre results: ranks below Luna and DeepSeek v4 Pro, just above Gemini 3.7 Flash. The author notes other benchmarks on X show Ox Alpha performing strongly, so results may be highly benchmark-specific.
- Stability concerns: the model frequently returned empty strings or API errors; 551 API calls were needed to reach an 87/100 completion rate. Unclear whether this is an OpenRouter issue or the common reasoning-model behavior of not answering when unsolved, amplified.
Benchmark data and evaluation code are open-sourced in the concept-synth GitHub repo (arXiv:2602.18843).
More from Models
- Hands-on with Ox Alpha: AI generates vase with through-holes — jakedahn · 2026-08-23
- Users report unannounced upgrade to ChatGPT: GPT 5.6 sees major speed and accuracy gains — SteveEricJordan · 2026-08-23
- Rumor: OpenAI is developing a music generation model codenamed 'Patrick' — iruletheworldmo · 2026-08-23
- Benchmark chasing degrades model interactivity, causing guessing instead of asking — Fowe · 2026-08-23
- Aikido benchmark: DeepSeek V4 Pro tops cyber security AI, open-source beats frontier — sull · 2026-08-23
- OpenAI Reportedly Developing 'Mona Lisa' and 'Luna Lisa' Image Models — mark_k · 2026-08-23