Artificial Analysis launches Optima to build custom benchmarks, testing GPT-6 Astra
ArtificialAnlys · x · 2026-09-10
Artificial Analysis launched Optima, a tool that lets you build custom benchmarks around your own tasks and data to evaluate the latest models, including GPT-6 Astra.
- A build agent drafts tasks and grading rubrics from your description, sample files, traces, or coding agent.
- Supports Q&A, document QA, agentic tasks, and (soon) conversation simulation; grading via custom rubrics or a panel of judge models.
- Every result includes cost and time-per-task alongside performance, claimed to cut evaluation cost and time by 10x+.
A practical option for teams choosing models for specific use cases — and for benchmarking new releases on day one.
More from Research
- VDiff-Bench: 1,756-question benchmark shows frontier models fail at spot-the-difference — yixin_wan_ · 2026-09-10
- Michael Levin's new paper: a structured latent space of patterns for new forms of life and mind — danfaggella · 2026-09-10
- Do AI doomers really have a strong forecasting record? XPT study suggests otherwise — random_walker · 2026-09-10
- AI safety frontier shifting from neural nets to mechanistic swarm interpretability — Hidenori8Tanaka · 2026-09-10
- AI spleen imaging linked to genomic heart disease risk: new heart-spleen axis — EricTopol · 2026-09-10
- AutoResearchExam: a 24-hour benchmark finds AI research agents overfit, with Fable 5.1 edging Astra — AlexGDimakis · 2026-09-10