WeaveWave: Towards Multimodal Music Generation

Abstract
WeaveWave is a multimodal music generation framework that synthesizes music from text, images, and video. It employs a text-bridging strategy: a multimodal large language model (Gemma-3-12b-it [3]) generates rich musical descriptions from arbitrary input modalities, which then condition a MusicGen-Style model [2] for audio synthesis at 32 kHz. This decoupled design enables modular integration of new modalities and generation backends. We further provide a scalable training pipeline based on MusicGen-Style and an interactive Gradio demo for real-time experimentation.
Motivation
For humans, music creation can be abstracted into two stages: inspiration and implementation. Inspiration originates from the fusion of diverse sensory experiences β visual scenes, literary imagery, auditory fragments, and other cross-modal perceptions. Implementation manifests as the process of concretizing that inspiration through performance.
For machines, can AI music creation mimic these two stages? We believe that multimodal music generation precisely simulates this process β where βinspirationβ corresponds to multimodal input data, and βimplementationβ to a music generation model.
Music creation: humans and machines
However, research on multimodal music generation has not yet garnered widespread attention, with most existing work confined to understanding and generation within a single modality. To address this gap, we implemented a text-bridging strategy harnessing existing MLLMs and text-to-music systems, proposed two candidate end-to-end architectures, and developed a training pipeline based on MusicGen-Style [2]. This exploration culminated in WeaveWave β a unified framework designed to integrate multimodal inputs through a cohesive generative process.
Architecture
Text-Bridging: MLLM generates music descriptions from multimodal input, MusicGen synthesizes audio
The text-bridging approach builds on MusicGen [1] and its style-conditioning extension [2], using Gemma-3 [3] as the multimodal front-end. The pipeline consists of two stages:
- Description generation β A multimodal LLM (Gemma-3-12b-it) interprets the input (text, image, or video) and produces a concise music description capturing mood, rhythm, genre, and instrumentation.
- Audio synthesis β MusicGen-Style conditions on the generated description via a frozen T5-base text encoder, with optional style conditioning (MERT-based, 6 codebooks at 5 Hz) and melody conditioning (chroma features). Audio is decoded through EnCodec at 32 kHz with optional MultiBand Diffusion post-processing.
We also explored two end-to-end alternatives during development:
End-to-End approach 1: based on AudioLDM2 [4]
End-to-End approach 2: based on MusicGen [1]
Demo
Click to view demo β WeaveWave web application built with Gradio
Key features:
- Generate music from text prompts, images, or video via a unified interface
- Choose from 10 MusicGen model variants (mono/stereo, small to large)
- Optional melody conditioning from uploaded audio (chroma-based)
- Optional MultiBand Diffusion decoding for enhanced audio quality
- Configurable generation parameters (duration, top-k, top-p, temperature, CFG)
Installation
git clone --recurse-submodules https://github.com/Audiofool934/WeaveWave.git
cd WeaveWave
pip install -e ".[dev,demo]"
pip install -e repos/audiocraft
Requirements: Python 3.9+, PyTorch 2.3.1+, CUDA-capable GPU recommended.
Quick Start
Launch the demo
weavewave-mllm-server
weavewave-demo
Open http://127.0.0.1:7860 in your browser.
Training pipeline
weavewave-prepare-data --create_dummy --dummy_samples 100
weavewave-train
weavewave-run-training --dummy_data --dummy_samples 100 --gpu 0 1
Evaluation
weavewave-evaluate \
--eval_text2music \
--model_path ./outputs/latest_model \
--output_dir ./outputs/evaluation \
--gpu 0
Supported modes: --eval_text2music, --eval_style2music, --eval_style_and_text2music.
Docker
docker compose up
This starts the MLLM backend and Gradio frontend as separate services.
Project Structure
WeaveWave/
βββ weavewave/ # Main Python package
β βββ core/ # Config, types, logging, clients
β β βββ config.py # AppConfig, PromptConfig
β β βββ types.py # GenerationConfig, MLLMServerConfig
β β βββ logging.py # Centralized logging
β β βββ mllm_client.py # HTTP client for MLLM service
β β βββ music_generator.py # MusicGen wrapper + MultiBand Diffusion
β βββ training/ # Training pipeline
β β βββ train.py # MusicGen-Style training
β β βββ runner.py # Data prep + training orchestrator
β βββ evaluation/ # Multi-mode evaluation
β β βββ evaluate.py
β βββ data/ # Dataset preparation
β β βββ prepare_dataset.py
β βββ serving/ # Web application
β βββ mllm_server.py # FastAPI backend (Gemma-3)
β βββ app.py # Gradio frontend
β βββ theme.py # Ocean-themed UI
βββ config/ # Hydra YAML configurations
βββ tests/ # Test suite (pytest)
βββ repos/audiocraft/ # Meta AudioCraft (git submodule)
βββ pyproject.toml # Package metadata & dependencies
βββ Dockerfile # Multi-stage build
βββ docker-compose.yml # Service orchestration
Environment Variables
| Variable | Purpose | Default |
|---|
WEAVEWAVE_MLLM_URL | MLLM service endpoint | http://127.0.0.1:8001 |
WEAVEWAVE_DEFAULT_MUSIC_MODEL | Default MusicGen checkpoint | facebook/musicgen-stereo-melody-large |
WEAVEWAVE_MLLM_MODEL | MLLM model identifier | google/gemma-3-12b-it |
WEAVEWAVE_MLLM_DEVICE | MLLM inference device | cuda |
CUDA_VISIBLE_DEVICES | GPU selection | β |
Citation
@software{weavewave2025,
title = {WeaveWave: Towards Multimodal Music Generation},
author = {Audiofool},
year = {2025},
url = {https://github.com/Audiofool934/WeaveWave},
}
References
[1] Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., & DΓ©fossez, A. (2024). Simple and controllable music generation. NeurIPS 2024. arXiv:2306.05284
[2] Rouard, S., Adi, Y., Copet, J., Roebel, A., & DΓ©fossez, A. (2024). Audio conditioning for music generation via discrete bottleneck features. ISMIR 2024. arXiv:2407.12563
[3] Google. (2025). Gemma 3 Technical Report. arXiv:2503.19786
[4] Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., Wang, Y., Wang, W., Wang, Y., & Plumbley, M. D. (2024). AudioLDM 2: Learning holistic audio generation with self-supervised pretraining. IEEE/ACM TASLP. arXiv:2308.05734
[5] Rinaldi, I., Fanelli, N., Castellano, G., & Vessio, G. (2024). Art2Mus: Bridging visual arts and music through cross-modal generation. arXiv:2410.04906
License
This project is licensed under the MIT License.