AgenticBBO-Bench benchmarks LLM agents for black-box optimization; GPT-6 Astra and DeepSeek-V4.1-Flash on Pareto frontier

Ming Chen · hf · 2026-10-09

Researchers introduce AgenticBBO-Bench, a cross-domain benchmark evaluating LLM agents on black-box optimization across synthetic functions, hyperparameter tuning, database tuning, chip design, and molecular design under a unified finite-budget protocol.

Findings:

Code: github.com/lamda-bbo/agentic-bbo

Original post →

More from coding & agent

coding & agent channel →