YingVideo-MV_demo.mp4
We present YingVideo-MV, the first cascaded framework for music-driven long-video generation. Our approach integrates audio semantic analysis, an interpretable shot planning module (MV-Director), temporal-aware diffusion Transformer architectures, and long-sequence consistency modeling to enable automatic synthesis of high-quality music performance videos from audio signals. We construct a large-scale Music-in-the-Wild Dataset by collecting web data to support the achievement of diverse, high-quality results. Observing that existing long-video generation methods lack explicit camera motion control, we introduce a camera adapter module that embeds camera poses into latent noise. To enhance continulity between clips during long-sequence inference, we further propose a time-aware dynamic window range strategy that adaptively adjust denoising ranges based on audio embedding. Comprehensive benchmark tests demonstrate that YingVideo-MV achieves outstanding performance in generating coherent and expressive music videos, and enables precise audio-motion-camera synchronization.
- Cascaded Training Pipeline
- Camera Control: Camera Adapter with Explicit Camera Control
- TDW: Timestep-aware Dynamic Window Range Strategy
- RL (DPO): Multi-reward Reinforcement Learning for Human Preferences
- 2025-11-27: Released technical report
- inference code in mid-December
- 1.3B model checkpoint in mid-December
If you use YingVideo-MV for research, please cite:
@article{chen2025yingvideo,
title={YingVideo-MV: Music-Driven Multi-Stage Video Generation},
author={Chen, Jiahui and Wang, Weida and Shi, Runhua and Yang, Huan and Ding, Chaofan and Chen, Zihao},
journal={arXiv preprint arXiv:2512.02492},
year={2025}
}
Our code is released under MIT License.

