qinglong-captions v4.7.0: Major Upgrades to Multimodal OCR and Underlying Models
bdsqlsz · x · 2026-08-02
The image captioning tool qinglong-captions has released v4.7.0, bringing several updates to multimodal capabilities and model support:
- OCR & OMR: Introduced OvisOCR2 for image and PDF OCR (supporting Transformers and OpenAI-compatible vLLM backends). Upgraded MuSViT to a complete ONNX sheet-music OMR workflow with MusicXML/MIDI export.
- Model Unification & Updates: Unified Kimi K3 model selection and reasoning controls, and refreshed catalogs for Gemini 3.x and MiniMax M2.7.
- Image Processing: Enforced image prompt-template I/O contracts and synchronized upstream behaviors for LayerDiff and Marigold.
Related event: qinglong-captions v4.7.0 Adds Multimodal OCR and Score Recognition(2 posts)→
More from Multimodal
- How to replace an image element with Flux Kontext or Qwen Image Edit? — Wemos_D1 · 2026-08-24
- Creating Halloween-style shorts with H3's R2V template — goulash47 · 2026-08-24
- 3D Box City Generator: Create New Cities with a Single Click — sujingshen · 2026-08-24
- H3 last frame stitching inconsistent, seeking fix — red_army25 · 2026-08-24
- Test: Generating 100% Gongbi style art with European features — Myronca · 2026-08-24
- Experimenting with camera path references for AI video generation workflows — StoreConnect1506 · 2026-08-24