AgenticBBO-Bench benchmarks LLM agents for black-box optimization; GPT-6 Astra and DeepSeek-V4.1-Flash on Pareto frontier
Ming Chen · hf · 2026-10-09
Researchers introduce AgenticBBO-Bench, a cross-domain benchmark evaluating LLM agents on black-box optimization across synthetic functions, hyperparameter tuning, database tuning, chip design, and molecular design under a unified finite-budget protocol.
Findings:
- Agentic BBO beats direct LLM methods in all five domains and the best numerical optimizers in four
- Extra numerical tools don't consistently help; task semantics are broadly useful while specific priors are less reliable
- Numerical optimizers can absorb gains from agent-established search trajectories
- In a five-task frontier challenge with seven LLMs under the Codex harness, GPT-6 Astra and DeepSeek-V4.1-Flash sit on the performance-cost Pareto frontier
Code: github.com/lamda-bbo/agentic-bbo
More from coding & agent
- OpenAI's math results released as open source RL environments — SergioPaniego · 2026-10-09
- Dev builds two playable games in a week with Claude Haiku and GPT small models, with cost breakdowns — chongdashu · 2026-10-09
- Stop vibe coding to 80%: pool your tokens so one person ships it to 100% — arjunrajlab · 2026-10-09
- dhh prefers Codex as main coding model with Claude secondary, praises Sol 6.1 — npew · 2026-10-09
- Leaked Claude Code Setup: Sonnet Builds, Haiku Swarms, Opus Architects — Arindam_1729 · 2026-10-09
- Users claim Codex harness, not the models, is the bottleneck: zcode 5x faster on same tasks — bdsqlsz · 2026-10-09