SciConBench to run bi-monthly, seeks API access and funding support
manoelribeiro · x · 2026-10-07
The SciConBench team plans to keep running the benchmark every two months for 1–2 years to track new frontier models, but agent evaluations at this scale and maintaining a live public leaderboard are costly. They are seeking API access, evaluation credits, or funding from model providers and organizations.
More from Research
- Ai2 publishes technical report on supercharging Olmo-core for scalable MoE training — StasBekman · 2026-10-07
- USC hiring a postdoc on evaluating simulations, starting spring 2027 — yoavartzi · 2026-10-07
- vf3 fuzzer unveiled at OAIC claims to outpace Jackalope and libprotobuf-mutator — dyn___ · 2026-10-07
- Scott Alexander's open letter to Steven Pinker: g-factor is real and AI scaling will keep climbing — Astral Codex Ten · 2026-10-07
- Paper shows diffusion transformer tokens encode lots of image info before it's interpretable — kwangmoo_yi · 2026-10-07
- Why synthetic cells die after five generations: they can't recycle their own broken parts — NikoMcCarty · 2026-10-07