qinglong-captions v4.7.0 Adds OvisOCR2 and Piano Sheet Music Recognition
bdsqlsz · x · 2026-08-02
The image captioning tool qinglong-captions has released its version 4.7.0 update. This version introduces OvisOCR2 for image OCR and integrates a newly trained MuSViT OMR model for Optical Music Recognition.
Currently trained on the official dataset, the music recognition model only supports piano scores for now. The author noted that future versions will attempt to recognize full musical scores.
Related event: qinglong-captions v4.7.0 Adds Multimodal OCR and Score Recognition(2 posts)→
More from Multimodal
- Seeking Audio Upscaling LLMs: Is There a 'Super-Resolution' Model for Music? — LeatherRub7248 · 2026-08-24
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24