LMBuild: UIUC benchmark tests whether LLM agents can build structures that actually work
UIUC-CS · hf · 2026-10-07
UIUC researchers introduce LMBuild, a benchmark evaluating LLM agents on generating buildable and functional 3D structures, addressing the gap where prior evaluations focus on geometry while ignoring physical realizability.
- Objects are represented as assembled structures with part decompositions, joints, materials, and assembly sequences
- The framework includes an interactive tool-using environment, a benchmark curated from CAD datasets augmented with Wikipedia knowledge, and metrics covering structural soundness, functional affordance, design quality, and physical realization
- Findings across 30 systems: frontier closed-source models have largely solved soundness and alignment, but functional affordance and physical operability remain hard; stronger models create novel components while weaker ones rely on retrieval; providing functional specifications substantially improves completeness and operability
More from Research
- CUAWright: Terminal-Only Computer-Use Agent Beats GUI Harnesses, Cuts Cost 37.5% — ysu_nlp · 2026-10-07
- AI's Top 10 papers list: Rulin Shao lands two first-author picks — ShayneRedford · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Paradigm: post-training gains hinge on combining procedural and LLM-based synthetic data — tensorqt · 2026-10-07
- Paradigm scales RL context from 65k to 131k tokens using a trained value model — tensorqt · 2026-10-07
- Limite 1B borrows nanogpt speedrun architecture: NorMuon, MUDD variant and XSA — tensorqt · 2026-10-07