V2N: Multi-Task Visual Piano Transcription for Notes, Offsets, and Velocity
PianoVAM · hf · 2026-08-05
Existing Visual Piano Transcription (VPT) systems primarily focus on onset detection from short video windows, lagging in offset accuracy and velocity prediction. To address this, researchers introduce V2N (Video to Notes), the first complete VPT system.
- Architecture: It uses a shared temporal backbone feeding task-specific heads for onset, offset, key hold, and velocity, jointly trained with per-frame supervision.
- Findings: Ablations show that multi-task supervision enables offset and velocity prediction while improving onset accuracy. Longer temporal context yields further enhancements.
- Results: V2N sets new state-of-the-art results on the PianoVAM and R3 datasets.
More from Multimodal
- Dual SAM3 + Seam Mask: A 4K Panoramic 3DGS Reconstruction Workflow — janusch_patas · 2026-08-05
- Using Hermes Desktop with ComfyUI: Letting AI Agents Auto-Fix Workflow Errors — Birdinhandandbush · 2026-08-05
- AI Demo Mimics Human Handwriting with Annotations and Streaming Charts — op7418 · 2026-08-05
- Handy ComfyUI Script: Audio Ping Notification for Job Completion — RPGstarDestroyer · 2026-08-05
- Tsinghua's STAMPlus Solves MLLM Segmentation Trilemma with Single-Pass Inference — Tsinghua · 2026-08-05
- Kyutai Launches Muscriptor: High-Precision Audio-to-MIDI Model — huggingface · 2026-08-05