FineVision, an open-source 24.3M-sample VLM training dataset, accepted to NeurIPS
andimarafioti · x · 2026-09-25
- lusxvr announced that FineVision, an open-source dataset for training state-of-the-art vision-language models, has been accepted to NeurIPS and will be presented at the conference.
- Scale: 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, positioned as one of the largest open VLM training datasets.
Related event: Open-Source VLM Dataset FineVision Accepted at NeurIPS(2 posts)→
More from Research
- Nokia Open-Sources AnyJev: Turn Any LLM into a Calibrated Decision Model, No Training — kalyan_kpl · 2026-09-25
- Why GPUs need philox, not xorshift: parallel RNG in AI training explained — abhi9u · 2026-09-25
- Independent researcher: activation steering measures the wrong geometry — 11 experiments on Qwen2.5-7B yield AkbasCore 3.2 — Nearby_Indication474 · 2026-09-25
- Navigating tenure-track in 2026: a guide to the academic job market — mboehme_ · 2026-09-25
- Grady Booch: LLMs Only Resemble the Brain at Its Most Primitive Structures — Grady_Booch · 2026-09-25
- ICLR 2027 Submission De-anonymization Incident Sparks OpenReview Statement — Striking-Warning9533 · 2026-09-25