Testing Minimax H3: Using vocal beats as audio reference and removing the voice
Portable_Solar_ZA · reddit · 2026-08-23
The author experimented with using beatboxing/vocal imitations as audio references in Minimax H3 for music generation.
The Issue:
- Using the Reference Model, the AI insisted on including the original voice in the output, mixing it with the background music instead of just using the rhythm.
- Testing the non-reference model unexpectedly worked better, stripping the voice and keeping only the beat.
The author is seeking advice on how to exclude the original voice while still using the reference model.
More from Multimodal
- Seeking Audio Upscaling LLMs: Is There a 'Super-Resolution' Model for Music? — LeatherRub7248 · 2026-08-24
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24