Model Eval: Hand-Drawn Floor Plans to CAD

sayashk · x · 2026-07-11

This evaluation tests whether frontier agents can convert an architect's hand-drawn sketches into CAD floor plans. The author collaborated with their mother, feeding the model three recent hand-drawn sketches for her to grade personally. Results show that previous-generation models performed poorly: GPT-5.5 Pro scored only 1/5 or 2/5 on each sketch, frequently misreading the drawings. In contrast, GPT-5.6 Sol (ultra) scored an average of about 80%, accurately recognizing furniture, room dimensions, and even inferring furniture placement. The author's mother noted she would consider hiring a junior architect who submitted such a work trial. However, it is still far from perfect: every output contains minor errors. The most common issues include silently altering the layout, misreading dimensions, drawing double-lined walls, messing up stairs, and making mistakes in baseline offsets and projections. The author also points out that while AI labs have been collecting architecture-related data, there are almost no benchmarks specifically for "hand-drawn floor plans to CAD." This task is worth pursuing, especially as hand-drafting is becoming a dying art.

Original post →

More from coding & agent

coding & agent channel →