Multimodal Agent for Digital Museums

PekingUniversity · hf · 2026-07-13

VaseMuseum is a digital intelligent museum framework dedicated to ancient Greek pottery. At its core is VaseAgent, a multimodal agent capable of processing both 2D images and 3D artifacts.

It is specifically designed to tackle two major challenges in digital museums:

Methodologically, it integrates multimodal perception, 3D-aware reasoning, and external knowledge retrieval with a training-free GRPO-style selection mechanism that leaves the VLM backbone untouched. In highly realistic digital museum simulations, the authors found that this framework improves citation validity, reduces hallucinations, and provides more restrained answers when dealing with ambiguous queries.

Original post →

More from Multimodal

Multimodal channel →