Skip to content
Research / research / 2026

WeaveWave

A multimodal music generation framework that turns text, image, and video inputs into musical descriptions and synthesized audio.

Python MusicGen Gemma Multimodal Gradio
GitHub
Audiofool934/WeaveWave
Language
Python
Stars
6
Last Push
2026.03.12
README Sync
2026.09.26
Towards multimodal music generation
[ Project Brief ]

Overview

WeaveWave explores multimodal music generation: turning text, images, and video into music through a shared generative process.

The project uses a text-bridging strategy. A multimodal language model first translates non-audio inputs into rich musical descriptions. Those descriptions then condition a music generation model. This keeps the system modular: new input modalities can be added without rewriting the entire synthesis backend.

Why It Matters

Music often starts from something that is not music yet: an image, a scene, a memory, a mood, a sentence. WeaveWave treats that pre-musical material as the real starting point.

The project sits close to the larger Audiofool question: how can generative systems carry taste across modalities?

Direction

WeaveWave is best understood as a bridge project: part multimodal perception, part text-to-music generation, part interface for creative translation.

[ Synced from GitHub README ]

Repository Document

source β†—

WeaveWave: Towards Multimodal Music Generation

WeaveWave Logo

Python 3.9+ PyTorch 2.3 License: MIT CI

Abstract

WeaveWave is a multimodal music generation framework that synthesizes music from text, images, and video. It employs a text-bridging strategy: a multimodal large language model (Gemma-3-12b-it [3]) generates rich musical descriptions from arbitrary input modalities, which then condition a MusicGen-Style model [2] for audio synthesis at 32 kHz. This decoupled design enables modular integration of new modalities and generation backends. We further provide a scalable training pipeline based on MusicGen-Style and an interactive Gradio demo for real-time experimentation.

Motivation

For humans, music creation can be abstracted into two stages: inspiration and implementation. Inspiration originates from the fusion of diverse sensory experiences β€” visual scenes, literary imagery, auditory fragments, and other cross-modal perceptions. Implementation manifests as the process of concretizing that inspiration through performance.

For machines, can AI music creation mimic these two stages? We believe that multimodal music generation precisely simulates this process β€” where β€œinspiration” corresponds to multimodal input data, and β€œimplementation” to a music generation model.

Music Creation: Humans and Machines

Music creation: humans and machines

However, research on multimodal music generation has not yet garnered widespread attention, with most existing work confined to understanding and generation within a single modality. To address this gap, we implemented a text-bridging strategy harnessing existing MLLMs and text-to-music systems, proposed two candidate end-to-end architectures, and developed a training pipeline based on MusicGen-Style [2]. This exploration culminated in WeaveWave β€” a unified framework designed to integrate multimodal inputs through a cohesive generative process.

Architecture

Text-Bridging Architecture

Text-Bridging: MLLM generates music descriptions from multimodal input, MusicGen synthesizes audio

The text-bridging approach builds on MusicGen [1] and its style-conditioning extension [2], using Gemma-3 [3] as the multimodal front-end. The pipeline consists of two stages:

  1. Description generation β€” A multimodal LLM (Gemma-3-12b-it) interprets the input (text, image, or video) and produces a concise music description capturing mood, rhythm, genre, and instrumentation.
  2. Audio synthesis β€” MusicGen-Style conditions on the generated description via a frozen T5-base text encoder, with optional style conditioning (MERT-based, 6 codebooks at 5 Hz) and melody conditioning (chroma features). Audio is decoded through EnCodec at 32 kHz with optional MultiBand Diffusion post-processing.

We also explored two end-to-end alternatives during development:

End-to-End based on AudioLDM2

End-to-End approach 1: based on AudioLDM2 [4]

End-to-End based on MusicGen

End-to-End approach 2: based on MusicGen [1]

Demo

Click to view demo β€” WeaveWave web application built with Gradio

Key features:

  • Generate music from text prompts, images, or video via a unified interface
  • Choose from 10 MusicGen model variants (mono/stereo, small to large)
  • Optional melody conditioning from uploaded audio (chroma-based)
  • Optional MultiBand Diffusion decoding for enhanced audio quality
  • Configurable generation parameters (duration, top-k, top-p, temperature, CFG)

Installation

# Clone with submodules
git clone --recurse-submodules https://github.com/Audiofool934/WeaveWave.git
cd WeaveWave

# Install the package (with demo and dev extras)
pip install -e ".[dev,demo]"

# Install AudioCraft from the vendored submodule
pip install -e repos/audiocraft

Requirements: Python 3.9+, PyTorch 2.3.1+, CUDA-capable GPU recommended.

Quick Start

Launch the demo

# Terminal 1 β€” MLLM backend (Gemma-3-12b-it on port 8001)
weavewave-mllm-server

# Terminal 2 β€” Gradio frontend (port 7860)
weavewave-demo

Open http://127.0.0.1:7860 in your browser.

Training pipeline

# Prepare a dummy dataset
weavewave-prepare-data --create_dummy --dummy_samples 100

# Train MusicGen-Style
weavewave-train

# Or use the runner with full options
weavewave-run-training --dummy_data --dummy_samples 100 --gpu 0 1

Evaluation

weavewave-evaluate \
    --eval_text2music \
    --model_path ./outputs/latest_model \
    --output_dir ./outputs/evaluation \
    --gpu 0

Supported modes: --eval_text2music, --eval_style2music, --eval_style_and_text2music.

Docker

docker compose up

This starts the MLLM backend and Gradio frontend as separate services.

Project Structure

WeaveWave/
β”œβ”€β”€ weavewave/                  # Main Python package
β”‚   β”œβ”€β”€ core/                   # Config, types, logging, clients
β”‚   β”‚   β”œβ”€β”€ config.py           # AppConfig, PromptConfig
β”‚   β”‚   β”œβ”€β”€ types.py            # GenerationConfig, MLLMServerConfig
β”‚   β”‚   β”œβ”€β”€ logging.py          # Centralized logging
β”‚   β”‚   β”œβ”€β”€ mllm_client.py      # HTTP client for MLLM service
β”‚   β”‚   └── music_generator.py  # MusicGen wrapper + MultiBand Diffusion
β”‚   β”œβ”€β”€ training/               # Training pipeline
β”‚   β”‚   β”œβ”€β”€ train.py            # MusicGen-Style training
β”‚   β”‚   └── runner.py           # Data prep + training orchestrator
β”‚   β”œβ”€β”€ evaluation/             # Multi-mode evaluation
β”‚   β”‚   └── evaluate.py
β”‚   β”œβ”€β”€ data/                   # Dataset preparation
β”‚   β”‚   └── prepare_dataset.py
β”‚   └── serving/                # Web application
β”‚       β”œβ”€β”€ mllm_server.py      # FastAPI backend (Gemma-3)
β”‚       β”œβ”€β”€ app.py              # Gradio frontend
β”‚       └── theme.py            # Ocean-themed UI
β”œβ”€β”€ config/                     # Hydra YAML configurations
β”œβ”€β”€ tests/                      # Test suite (pytest)
β”œβ”€β”€ repos/audiocraft/           # Meta AudioCraft (git submodule)
β”œβ”€β”€ pyproject.toml              # Package metadata & dependencies
β”œβ”€β”€ Dockerfile                  # Multi-stage build
└── docker-compose.yml          # Service orchestration

Environment Variables

VariablePurposeDefault
WEAVEWAVE_MLLM_URLMLLM service endpointhttp://127.0.0.1:8001
WEAVEWAVE_DEFAULT_MUSIC_MODELDefault MusicGen checkpointfacebook/musicgen-stereo-melody-large
WEAVEWAVE_MLLM_MODELMLLM model identifiergoogle/gemma-3-12b-it
WEAVEWAVE_MLLM_DEVICEMLLM inference devicecuda
CUDA_VISIBLE_DEVICESGPU selectionβ€”

Citation

@software{weavewave2025,
    title  = {WeaveWave: Towards Multimodal Music Generation},
    author = {Audiofool},
    year   = {2025},
    url    = {https://github.com/Audiofool934/WeaveWave},
}

References

[1] Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., & DΓ©fossez, A. (2024). Simple and controllable music generation. NeurIPS 2024. arXiv:2306.05284

[2] Rouard, S., Adi, Y., Copet, J., Roebel, A., & DΓ©fossez, A. (2024). Audio conditioning for music generation via discrete bottleneck features. ISMIR 2024. arXiv:2407.12563

[3] Google. (2025). Gemma 3 Technical Report. arXiv:2503.19786

[4] Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., Wang, Y., Wang, W., Wang, Y., & Plumbley, M. D. (2024). AudioLDM 2: Learning holistic audio generation with self-supervised pretraining. IEEE/ACM TASLP. arXiv:2308.05734

[5] Rinaldi, I., Fanelli, N., Castellano, G., & Vessio, G. (2024). Art2Mus: Bridging visual arts and music through cross-modal generation. arXiv:2410.04906

License

This project is licensed under the MIT License.