Multimodal Agent for Digital Museums
PekingUniversity · hf · 2026-07-13
VaseMuseum is a digital intelligent museum framework dedicated to ancient Greek pottery. At its core is VaseAgent, a multimodal agent capable of processing both 2D images and 3D artifacts.
It is specifically designed to tackle two major challenges in digital museums:
- Evidence Grounding: It retrieves evidence from authoritative web pages and museum knowledge bases, actively minimizing reliance on weak sources and unverifiable citations.
- Uncertainty Control: When evidence is insufficient, noisy, or conflicting, the system evaluates whether the generated content is adequately supported by the evidence, leaning towards more neutral and clearly bounded responses.
Methodologically, it integrates multimodal perception, 3D-aware reasoning, and external knowledge retrieval with a training-free GRPO-style selection mechanism that leaves the VLM backbone untouched. In highly realistic digital museum simulations, the authors found that this framework improves citation validity, reduces hallucinations, and provides more restrained answers when dealing with ambiguous queries.
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22