FrontierOR: Benchmarking LLMs on Optimization Algorithms
新智元 · wechat · 2026-07-10
FrontierOR is a novel LLM-for-OR benchmark. Instead of testing whether models can translate problems into mathematical programs, it evaluates their ability to identify structures and design scalable algorithms for real-world industrial challenges, much like an operations research engineer.
The benchmark selects 180 tasks from OR literature spanning 1992—2025, featuring standardized problem descriptions, mathematical models, Gurobi reference implementations, reference solutions, and feasibility checkers. A subset of 50 more difficult tasks forms the Hard set. Evaluations cover executability, feasibility, solution quality, and combined quality-efficiency metrics, with a focus on testing algorithmic design capabilities on large-scale instances.
More from Research
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop — 404 Media · 2026-07-21
- Shared agent workspaces fail in a fixed order, from stale reads to zombie writes — mrvladp · 2026-07-21