ReasonCore AI Publishes SciCode 500 Benchmark
geoffwolfe · x · 2026-08-27
ReasonCore AI published the "SciCode 500" benchmark, tested against Grok 4.5, Ox Alpha, DeepSeek V4 Pro, and Muse Spark 1.1. The benchmark features a wide range of solvability, with 100% headroom problems verified solved by an off-benchmark SOTA model. ReasonCore has over 5,000 such problems in inventory.
More from Research
- Flow matching decouples generative modeling from noise process — CSProfKGD · 2026-08-27
- Study finds Agent skill injection may lower Pass@2 rates — rohanpaul_ai · 2026-08-27
- Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig) — 1ncehost · 2026-08-27
- METR's new eval report gains traction over models losing track of tasks — isidentical · 2026-08-27
- Why Is the P=NP Question So Relevant in the AI Era? — yoavgo · 2026-08-27
- Fine-Tuning Guide: How Mistral 7B Saved $300k Over Foundation Models — Nice-Dragonfly-4823 · 2026-08-27