SciConBench to run bi-monthly, seeks API access and funding support

manoelribeiro · x · 2026-10-07

The SciConBench team plans to keep running the benchmark every two months for 1–2 years to track new frontier models, but agent evaluations at this scale and maintaining a live public leaderboard are costly. They are seeking API access, evaluation credits, or funding from model providers and organizations.

Related event: SciConBench, a NeurIPS-Accepted Benchmark, Finds Top AI Models Still Fail at Scientific Conclusion Synthesis(10 posts)→

Original post →

More from Research

Research channel →