FrontierOR: Benchmarking LLMs on Optimization Algorithms

新智元 · wechat · 2026-07-10

FrontierOR is a novel LLM-for-OR benchmark. Instead of testing whether models can translate problems into mathematical programs, it evaluates their ability to identify structures and design scalable algorithms for real-world industrial challenges, much like an operations research engineer.

The benchmark selects 180 tasks from OR literature spanning 1992—2025, featuring standardized problem descriptions, mathematical models, Gurobi reference implementations, reference solutions, and feasibility checkers. A subset of 50 more difficult tasks forms the Hard set. Evaluations cover executability, feasibility, solution quality, and combined quality-efficiency metrics, with a focus on testing algorithmic design capabilities on large-scale instances.

Original post →

More from Research

Research channel →