Skip to content
Research / archived / 2026

Genres in Genres

A music style evolution project that uses MuQ-MuLan embeddings to find sub-genres inside an artist discography.

Jupyter Notebook MuQ-MuLan Music Analysis Clustering Visualization
GitHub
Audiofool934/genres-in-genres
Language
Jupyter Notebook
Stars
0
Last Push
2026.01.15
README Sync
2026.05.30
[ Project Brief ]

Overview

Genres in Genres studies style evolution inside a music collection.

The project uses MuQ-MuLan embeddings to map songs into a semantic music space, then clusters and visualizes how an artist’s sound changes across albums and years.

Project Shape

The system combines library scanning, embedding extraction, K-Means clustering, dimensionality reduction, radar charts, and temporal style views.

It is a course-project artifact, but it also belongs to a larger Audiofool thread: using machine listening to make musical taste, history, and structure visible.

[ Synced from GitHub README ]

Repository Document

source β†—

Genres in Genres

Style evolution analysis for music collections using MuQ-MuLan embeddings.

Project Overview

Genres in Genres is a music style evolution analysis tool that uses MuQ-MuLan embeddings (512-dim vectors) to analyze how an artist’s sound changes over time. It identifies sub-genres within an artist’s discography, visualizes style trajectories, and provides semantic interpretation of musical clusters.

Features

  • Automatic artist library scanning and caching
  • Sub-genre identification using K-Means clustering (with Auto-K)
  • 2D trajectory visualization (PCA/t-SNE/UMAP)
  • Semantic radar charts for album comparison
  • Streamgraph for temporal style distribution

Installation

./run_demo.sh

This creates a virtual environment and installs dependencies.

Usage

1. Preprocess Audio (Optional)

If you have your own music files:

# Organize files as: data/music/{Artist}/{Year}-{Album}/*.mp3
source venv/bin/activate
pip install muq
python scripts/preprocess.py --device cuda  # or mps/cpu

2. Cache Semantic Tags

python scripts/cache_tags.py --max_tags 2000 --format pickle

3. Run Dashboard

./run_demo.sh
# Or run on specific port
./run_demo.sh --port 8080

Open the URL in browser, go to Library tab, select an artist, click Analyze.

4. Run Tests

source venv/bin/activate
python -m pytest tests/

Architecture

Data Flow

  1. Audio Input β†’ data/music/{Artist}/{Year}-{Album}/*.mp3
  2. Feature Extraction β†’ scripts/preprocess.py uses MuQ-MuLan to generate 512-dim embeddings
  3. Cache Storage β†’ data/cache/music/{Artist}.pkl (pickled ArtistCareer objects)
  4. Analysis β†’ StyleAnalyzer performs clustering, dimensionality reduction, metrics calculation
  5. Visualization β†’ Gradio app renders trajectory plots, streamgraphs, radar charts

Core Data Structures (src/core.py)

  • Track: Single song with metadata (file_path, title, album, release_date)
  • StyleEmbedding: 512-dim MuLan vector bound to a Track
  • ArtistCareer: Collection of tracks and embeddings for an artist, sorted chronologically

Key Modules

  • src/analysis.py: StyleAnalyzer class - clustering (KMeans with auto-K via silhouette score), dimensionality reduction (PCA/t-SNE/UMAP), career report generation
  • src/metrics.py: MusicMetrics class - computes style velocity (album-to-album change), novelty (departure from past work), cohesion (intra-album consistency)
  • src/semantics.py: SemanticMapper - maps audio embeddings to human-readable tags using cached text embeddings
  • src/library_manager.py: Handles filesystem scanning and pickle-based caching
  • src/visualization.py: GenreTrajectoryVisualizer and CareerStoryteller for all plots

Gradio App Structure (app.py)

  • Simulate tab: Generate mock data for testing
  • Library tab: Analyze cached artists with configurable clustering and visualization options
  • Insight Report tab: AI-generated narrative about style evolution
  • Dynamic cluster explorer with audio playback

Directory Structure

genres-in-genres/
β”œβ”€β”€ app.py                  # Gradio application
β”œβ”€β”€ run_demo.sh             # Setup script
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ preprocess.py       # Audio feature extraction
β”‚   β”œβ”€β”€ prepare_artist.py   # Library preparation
β”‚   β”œβ”€β”€ cache_tags.py       # Tag caching
β”‚   └── verify_semantics.py # Model verification
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ core.py             # Data structures
β”‚   β”œβ”€β”€ library_manager.py  # Cache management
β”‚   β”œβ”€β”€ muq.py              # MuQ-MuLan wrapper
β”‚   β”œβ”€β”€ analysis.py         # Clustering logic
β”‚   β”œβ”€β”€ semantics.py        # Semantic mapper
β”‚   β”œβ”€β”€ metrics.py          # Style metrics
β”‚   β”œβ”€β”€ mock_data.py        # Test data
β”‚   └── visualization.py    # Plotting
└── data/
    β”œβ”€β”€ music/              # Input audio files
    β”œβ”€β”€ cache/              # Cached embeddings
    └── metadata/           # Music4All tags

Music Directory Structure

data/music/{Artist}/{Year}-{Album}/*.mp3

Example: data/music/Radiohead/1997-OK Computer/01 - Airbag.mp3

The year prefix is parsed to order albums chronologically.

Key Metrics

  • Velocity: Cosine distance between consecutive album centroids (measures rate of style change)
  • Novelty: Cosine distance from current album to cumulative past centroid (measures departure from established style)
  • Cohesion: 1 - mean intra-album variance (measures stylistic consistency within an album)

Requirements

  • Python 3.8+
  • torch, torchaudio
  • muq (MuQ-MuLan)
  • gradio
  • scikit-learn, umap-learn
  • matplotlib, seaborn