Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score
dfrsrchtwts · x · 2026-09-04
- Researcher FRhysWard points out that the system card raises concerns of data contamination in the maths benchmark used to compute Astra's time-horizon metric.
- He calls on OpenAI to run the full set of No-CoT time-horizon evals from their paper on Astra, which is currently not available via the API for independent verification.
More from Models
- GPT-6 Astra debuts at No.1 on Terminal-Bench, 1.9% ahead of Claude Fable 5.1 — sandersted · 2026-09-04
- ARC-AGI-3 is now saturated, prompting calls for new benchmarks ASAP — kimmonismus · 2026-09-04
- Astra model card: no CoT-monitor evasion when reasoning must be verbalized — bookwormengr · 2026-09-04
- OpenAI Researcher roon: GPT-6 Astra Will Be Obsolete in Weeks — Tolopono · 2026-09-04
- Reasoning effort switching without breaking cache is live in Codex and Claude — altryne · 2026-09-04
- Mollick: GPT-6 autonomously does meaningful work for days, built an open-source Alexandria Library simulation — emollick · 2026-09-04