Model Eval: Hand-Drawn Floor Plans to CAD
sayashk · x · 2026-07-11
This evaluation tests whether frontier agents can convert an architect's hand-drawn sketches into CAD floor plans. The author collaborated with their mother, feeding the model three recent hand-drawn sketches for her to grade personally. Results show that previous-generation models performed poorly: GPT-5.5 Pro scored only 1/5 or 2/5 on each sketch, frequently misreading the drawings. In contrast, GPT-5.6 Sol (ultra) scored an average of about 80%, accurately recognizing furniture, room dimensions, and even inferring furniture placement. The author's mother noted she would consider hiring a junior architect who submitted such a work trial. However, it is still far from perfect: every output contains minor errors. The most common issues include silently altering the layout, misreading dimensions, drawing double-lined walls, messing up stairs, and making mistakes in baseline offsets and projections. The author also points out that while AI labs have been collecting architecture-related data, there are almost no benchmarks specifically for "hand-drawn floor plans to CAD." This task is worth pursuing, especially as hand-drafting is becoming a dying art.
More from coding & agent
- Why vector databases slow AI agents down after constant writes — PrajwalTomar_ · 2026-07-21
- A 13-minute GitHub Copilot video digs into prompt caching — lee_stott · 2026-07-21
- Codex turns out 123 screensavers in one playful batch — intellectronica · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21