FineVision, the 17M-image open VLM dataset from 200+ sources, accepted to NeurIPS
andimarafioti · x · 2026-09-25
FineVision — the open-source VLM dataset curation the team released earlier — has been accepted to NeurIPS. The dataset aggregates 200+ sources into 17M unique images and 10B answer tokens, adding new capabilities like GUI navigation, pointing, and counting.
According to the authors, training on FineVision improves open-source VLMs by 20% across 10 benchmarks. First authors lusxvr and orrzohar will present in Sydney.
More from Multimodal
- A 30-second video generated from a single image and one prompt (prompt included) — umesh_ai · 2026-09-25
- 6,300 frames of Su Shi's life: an MV animated entirely with Claude Opus 5.5 — dotey · 2026-09-25
- MiniMax H3 seems overtrained on smiles: 'bored caterpillar' video prompt keeps breaking immersion — episodex86 · 2026-09-25
- Pose Blueprint: A Browser-Based 3D Pose Editor for ComfyUI and ControlNet — OkConfusion6667 · 2026-09-25
- Reddit User Explores AI Art With Only Steps, CFG and Denoise Tweaks, No LoRAs — Extreme_Nice · 2026-09-25
- A sub-$20 LoRA makes Qwen-Image 2.1 rotate transparent objects with a prompt — ben_burtenshaw · 2026-09-25