Skip to content
Lib / research / 2025

MUSE / T2M

A text-to-music research line around probabilistic generation, Flow Matching, MuQ-MuLan anchors, and Diffusion Transformers.

PyTorch Flow Matching T2M MuQ-MuLan DiT
GitHub
Audiofool934/muse
Language
Python
Stars
0
Last Push
2026.03.13
README Sync
2026.09.26
[ Project Brief ]

Overview

MUSE is a text-to-music generation system built as a final project for the Parallel Computation course. It addresses the fundamental β€œone-to-many” problem in creative AI: a single text prompt like β€œa sad piano melody” can correspond to infinitely many valid musical interpretations.

MUSE Pipeline

The One-to-Many Problem

Traditional regression approaches predict the mean of possible outputs, resulting in blurry, mode-averaged music. MUSE instead models the full conditional distribution p(Y∣X)p(Y \mid X), enabling diverse, high-fidelity sampling.

ApproachOutputIssue
Regression Y=f(X)Y = f(X)Mean E[Y∣X]\mathbb{E}[Y \mid X]Blurry, averaged
Generative p(Y∣X)p(Y \mid X)Diverse samplesβœ“ Preserves creativity

Two-Stage Architecture

Stage 1: Text2MuQFlow (271M params)

  • Cross-attention Flow Matching
  • Output: 512-dim MuQ-MuLan embedding
  • 50-step ODE (dopri5 solver), CFG scale 3.0

Stage 2: StableAudioMuQ (1.05B params)

  • Diffusion Transformer + Classifier-Free Guidance
  • 100-step ODE sampling
  • Output: 44.1kHz stereo, up to 47 seconds

Parallel Computing Highlights

ChallengeSolution
Memory-heavy (1B+ params)PyTorch DDP across 4Γ—A800 GPUs
Compute-bound (100-step ODE)Batch inference parallelization
I/O bottleneck (44.1kHz audio)Cached MuQ embeddings

Training Summary

StageGPUsWall TimeGPU-hrsBest Loss
Stage 12Γ—A800~8.2h~160.0569
Stage 24Γ—A800~32h~1281.078

Stage 1 Loss

Batch Inference Performance

The key insight: model size determines batching efficiency.

Model SizeBatch SpeedupReason
< 500M5-30Γ—GPU under-utilized
500M-1B2-5Γ—Partial saturation
> 1B~1Γ—Already saturated

Stage 1 (271M): 15Γ— speedup at BS=16 β€” memory stays constant at 4.58GB Stage 2 (1.05B): 1.5Γ— speedup at BS=16 β€” memory scales linearly (18β†’36GB)

End-to-End Result

  • 42% time reduction for 16 samples (185s β†’ 108s)
  • Reproducible outputs (max diff < 1e-7)

MUSE Application

The full-stack application includes:

  • 🎹 Studio: Text-to-music generation with batch sampling
  • πŸ”¬ Lab: Latent space exploration & interpolation
  • πŸ“š Library: Audio management, tagging & playback

Latent Space Diverse samples visualized in MuQ-MuLan latent space

Tech Stack

Gradio Β· PyTorch Β· torchaudio Β· Plotly Β· UMAP


πŸ“Ž View Research Poster | πŸ”— GitHub Repository

[ Synced from GitHub README ]

Repository Document

source β†—

MUSE: Music Unified Synthesis Engine

Generate music from any input β€” text, image, video, or audio β€” through one unified architecture.

MUSE decouples what you perceive from how you synthesize. A shared two-stage flow matching backbone maps any modality to audio; adding a new input requires only a lightweight perception encoder β€” Stage 2 never changes.

Key Results

Text-to-music baseline evaluated on MusicBench (2,811 samples):

ModelFAD ↓KL Sigmoid ↑
AudioLDM3.820.744
MusicGen5.360.844
MUSE2.250.925

FAD 2.25 β€” 41% lower than AudioLDM. KL 0.925 β€” best semantic alignment. Stereo 44.1 kHz, ~12 s.

Beyond point estimates, the flow matching formulation models p(z∣c)p(z \mid c) as a full distribution:

  • Vague prompts β†’ higher output diversity (APD 1.037 vs. 0.963 for specific prompts)
  • Ambiguous prompts β†’ distinct genre clusters (e.g., β€œCyberpunk city” splits into synthwave / ambient / industrial)
  • Smooth latent interpolation via noise-space SLERP

Architecture

graph LR
    subgraph Perception["<b>Perception Layer</b> (frozen)"]
        T["πŸ”€ Text"] --> T5["Flan-T5"]
        I["πŸ–ΌοΈ Image"] --> CLIP["CLIP / SigLIP"]
        V["🎬 Video"] --> VF["Frame Encoder"]
        A["🎡 Audio"] --> MQ["MuQ-MuLan"]
        X["✳️ Any"] --> MB["MLLM Bridge<br/><i>Gemma-3 β†’ T5</i>"]
    end

    subgraph Generation["<b>Stage 1</b>: Flow Matching"]
        FM["Cross-Attention<br/>Transformer<br/><i>16L Β· 1024d</i>"]
    end

    subgraph Synthesis["<b>Stage 2</b>: Audio Synthesis"]
        DiT["Stable Audio DiT<br/><i>24L</i> + Oobleck VAE"]
    end

    T5 & CLIP & VF & MQ & MB -->|"ConditioningOutput<br/>[B, L, 768]"| FM
    FM -->|"MuQ-MuLan<br/>[B, 512]"| DiT
    DiT -->|"44.1 kHz<br/>stereo"| Out["πŸ”Š Audio"]

Core insight: Stage 2 is modality-agnostic β€” it only sees a 512-dim MuQ-MuLan vector. Adding a new modality = one encoder + one Stage 1 training run. The MLLM bridge (Gemma-3 β†’ T5) enables zero-shot input from any modality with no training at all.

Three-Layer Decoupling

LayerResponsibilityInterface
PerceptionModality β†’ conditioning embeddingsPerceptionEncoder.encode() β†’ ConditioningOutput
GenerationConditioning β†’ music latentCond2LatentFlow.generate() β†’ [B, 512]
SynthesisLatent β†’ waveformLatentToAudioDiT.sample() β†’ [B, 2, T]

ConditioningOutput β€” a [B, L, D] tensor + padding mask β€” is the universal contract. Every encoder produces it; every generator consumes it. Switching modality is a one-line config change.

Supported Pipelines

PipelineInputEncoderTraining
t2m_flowTextFlan-T5Stage 1 + 2 βœ“
i2m_flowImageCLIP ViTStage 1 only
i2m_bridgeImageGemma-3 β†’ T5None (zero-shot)
v2m_flowVideoCLIP framesStage 1 only
a2m_flowAudioMuQ-MuLanStage 1 only

Usage

from muse.pipelines import TwoStageFlowPipeline

pipe = TwoStageFlowPipeline.from_config("configs/t2m_flow.yaml")
audio = pipe.generate("A melancholic cello solo over soft rain")

# Zero-shot image β†’ music (no additional training)
pipe = TwoStageFlowPipeline.from_config("configs/i2m_bridge.yaml")
audio = pipe.generate("sunset.jpg")

Method

Both stages use Conditional Flow Matching (Lipman et al., ICLR 2023) with the OT-affine path:

x_t=(1βˆ’t)β‹…x_0+tβ‹…x_1,v_ΞΈ(x_t,t,c)β‰ˆx_1βˆ’x_0x\_t = (1-t) \cdot x\_0 + t \cdot x\_1, \quad v\_\theta(x\_t, t, c) \approx x\_1 - x\_0

At inference, an ODE solver integrates from Gaussian noise to data. Classifier-free guidance steers generation:

v_guided=v_uncond+wβ‹…(v_condβˆ’v_uncond)v\_\text{guided} = v\_\text{uncond} + w \cdot (v\_\text{cond} - v\_\text{uncond})

All modalities pass through a MuQ-MuLan bottleneck (512-dim, L2-normalized) β€” a contrastive audio-text space that provides semantic alignment and a shared interface for Stage 2.

Project Structure

muse/
β”œβ”€β”€ perception/              # Modality encoders (T5, CLIP, MuQ-MuLan, MLLM bridge)
β”œβ”€β”€ generation/flow_matching/ # Stage 1 (Cond2LatentFlow) + Stage 2 (LatentToAudioDiT)
β”œβ”€β”€ pipelines/               # Config-driven two-stage assembly
β”œβ”€β”€ sampling/                # Latent selection (peak, diverse, DBSCAN, k-means)
β”œβ”€β”€ data/                    # Multi-modal dataset with JSONL manifest
└── training/                # Distributed trainer interface

References

  • Lipman et al., Flow Matching for Generative Modeling, ICLR 2023
  • Evans et al., Stable Audio Open, 2024
  • MuQ-MuLan: contrastive audio-text embeddings