Testing Qwen3.8-Max: Precise Image Grounding with Bounding Boxes
lateinteraction · x · 2026-08-06
Kevin Madura wrote a detailed review testing the image grounding capabilities of Qwen3.8-Max. The experiment bypassed standard visual Q&A by asking the model to return structured JSON bounding boxes, which were then drawn back onto the original image using Python and Pillow.
The tests revealed key insights: the model must be explicitly told the actual image dimensions to avoid misplaced boxes. Once corrected, the model successfully identified 64 cars in a dense traffic image with precise alignment. The experiment also extended to complex football game frames, asking the model to classify player roles and read jersey numbers, demonstrating strong multimodal potential.
More from Models
- Top AI Models Score Under 50% on New Math Figure Reasoning Benchmark — prof_g · 2026-08-06
- Ethan Mollick: LLMs Improve at Following Instructions but Exercise More 'Judgement' — emollick · 2026-08-06
- Frontier LLMs Perform Best in Week 1: Dev Calls for a Proof of Model Standard — sull · 2026-08-06
- Polymarket Favorites Anthropic at 69% to Win AI Race by 2026 — Polymarket · 2026-08-06
- DeepSeek Tries, Opus Finishes: A Practical AI Coding Model Workflow — gandamu_ml · 2026-08-06
- Palantir Claims Vanilla NVIDIA Nemotron Outperforms Frontier Models in 24 Hours — eliano · 2026-08-06