Skip to search boxSkip to navigationSkip to main content

From Sound to Sight: Towards AI-authored Music Videos

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Open access

Publication Information

Output type

Research Output:
Conference Article in Proceeding or Book/Report chapter
Article in proceedings
Peer-review

Original language

English

Publication milestones

  • Published - 2025

Publication status

Published - 2025

Publisher

IEEE, United States

Host publication title

2025 IEEE/CVF International Conference on Computer Vision (ICCV)

Abstract

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any user-specified, vocal or instrumental song using off-the-shelf deep learning models. Inspired by the manual workflows of music video producers, we experiment on how well latent feature-based techniques can analyse audio to detect musical qualities, such as emotional cues and instrumental patterns, and distil them into textual scene descriptions using a language model. Next, we employ a generative model to produce the corresponding video clips. To assess the generated videos, we identify several critical aspects and design and conduct a preliminary user evaluation that demonstrates storytelling potential, visual coherency and emotional alignment with the music. Our findings underscore the potential of latent feature techniques and deep generative models to expand music visualisation beyond traditional approaches.

Funding Details

FundersFunding numbers
-
-

Related Event

Title

Generative AI for Storytelling

Event type

Conference

Degree of recognition

International event

Date

20/10/2025 - 20/10/2025

Location

HawaiiHonoluluUnited States