MedPMC: A Medical Multimodal Data Framework
Yale-BIDS-Chen · hf · 2026-07-13
This work introduces MedPMC, an infrastructure designed to automatically process open-access literature from PubMed Central into high-quality medical image-text data for training multimodal foundation models.
Key Approach
- Processes 6.1 million PMC articles to extract 11 million medical image-text pairs.
- The framework features components for initial filtering, multi-panel image detection, image separation, caption alignment, and medical figure classification.
- Component evaluations show strong performance across the pipeline, achieving an F1 of 93.2 in initial filtering, 96.5 in multi-panel detection, and an mAP of 89.8 in image separation.
Quality and Impact
- Manual review by 5 annotators confirmed that 95.3% of MedPMC images are medically relevant, compared to only 19.7% in legacy PMC datasets.
- Across 11 specialties and 26 benchmarks, a CLIP-style model trained on MedPMC outperformed the strongest biomedical CLIP baseline of the same architecture by an average of 7.1 percentage points in zero-shot AUC, using less than half the image-text pairs.
- When used as a vision encoder for multimodal LLMs, it boosted performance on two medical VQA benchmarks by 1.9 and 16.9 percentage points, respectively.
- For morphology-to-image retrieval on 10,524 dermatology photos from the Yale New Haven Health system, Recall@5 improved by 11.7 percentage points.
Conclusion
The authors conclude that high-fidelity, literature-level data cleansing significantly enhances medical multimodal foundation models. They have publicly released the framework, corpus, benchmarks, and pre-trained models.
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22