Skip to content
Research / archived / 2025

Musical Carbon Dating

A statistical course project that estimates the era of a recording from acoustic and perceptual audio features.

TeX Statistics Regression Audio Features Spotify Data
GitHub
Audiofool934/musical-carbon-dating
Language
TeX
Stars
1
Last Push
2025.12.26
README Sync
2026.05.30
[ Project Brief ]

Overview

Musical Carbon Dating is a statistical audio analysis project.

The project asks a playful but concrete question: can we timestamp a musical recording from its acoustic properties alone?

Using a large track dataset and a set of physical and perceptual audio features, it builds a regression pipeline for estimating musical era. The result is less a perfect clock than a way to measure how sound carries historical fingerprints.

Why It Matters

The project connects statistics with listening culture. Tempo, energy, valence, and spectral features become clues in a larger question: what makes a song sound like its time?

That makes it a natural course-project artifact for audiofool.blog.

[ Synced from GitHub README ]

Repository Document

source β†—

Musical Carbon Dating 🎡

A Statistical Feature Recognition Analysis (1960-2020)

β€œCan we timestamp a musical recording purely from its acoustic properties?”

This project implements a rigorous statistical pipeline to quantifiably β€œcarbon date” music. By analyzing 13 physical and perceptual audio features (e.g., Tempo, Valence, Spectral Energy) across 250,971 tracks, we demonstrate that musical eras have distinct, mathematically recognizable acoustic fingerprints.


πŸ“Š Key Results (Verified)

1. Statistical Validity

Despite the complexity of artistic expression, our Weighted Least Squares (WLS) model ensures valid inference:

  • R2R^2: 0.275. (27.5% of variance explained by acoustic physics alone).
  • MAE: 9.30 years. (Average error < 1 decade).
  • See output/figures/pred_vs_act_best_model_(wls)_predictions.png.

2. Statistical Rigor

We explicitly address the failures of standard OLS regression:

  • Heteroscedasticity: Diagnosed via Breusch-Pagan Test (Ο‡2β‰ˆ19,567,p<0.001\chi^2 \approx 19,567, p < 0.001) and corrected with WLS weights (wi∝1/Οƒi2w_i \propto 1/\sigma_i^2).
  • Non-Linearity: Confirmed via Partial F-Tests (Fβ‰ˆ4945F \approx 4945).
  • Feature Selection: LASSO (L1L_1) regularization validated the use of all 13 acoustic features, confirming that even subtle markers (like Key and Mode) are essential for tracking harmonic evolution.

3. The β€œNostalgia Index”

We define prediction error as a commercial metric: Index=∣y^predβˆ’yactual∣\text{Index} = |\hat{y}_{pred} - y_{actual}|.

  • Insight: High index values identify songs that are β€œTime-Displaced” (Retro or Futuristic).
  • Examples:
    • Uptown Funk (2015): Index 1.9 (Modern construction).
    • Physical (Dua Lipa, 2020): Index 11.0 (Strong 80s aesthetic).

πŸ›  Project Structure

β”œβ”€β”€ data/                   # Dataset (Spotify 600k Tracks)
β”œβ”€β”€ src/                    # Source Code
β”‚   β”œβ”€β”€ analysis.py         # Regression Engine (OLS, WLS, Ridge, LASSO, Stepwise)
β”‚   β”œβ”€β”€ config.py           # Configuration (Feature Definitions)
β”‚   β”œβ”€β”€ data_loader.py      # Data Preprocessing
β”‚   └── visualization.py    # Plotting Logic
β”œβ”€β”€ report/                 # [FINAL] Formal LaTeX Report
β”‚   β”œβ”€β”€ main.tex            # Comprehensive Academic Report
β”‚   └── figures/            # Auto-generated Figures
β”œβ”€β”€ slides/                 # Presentation Slides (Beamer)
β”‚   β”œβ”€β”€ main.tex            # "No Hiding" Detailed Slides
β”‚   └── speaker_notes.md    # 22-min Verbatim Script (5 Speakers)
β”œβ”€β”€ output/                 # Generated Artifacts
β”‚   β”œβ”€β”€ figures/            # Residual Plots, Q-Q Plots, Prediction Plots
β”‚   β”œβ”€β”€ tables/             # CSV Results
β”‚   └── pipeline_verified.log # Definitive Statistical Output
└── main.py                 # Main Execution Pipeline

πŸš€ Usage

1. Environment Setup

python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

2. Run Main Analysis

Execute the full pipeline (Data Load -> Feature Selection -> WLS -> Diagnostics):

python3 main.py

This will generate all figures in output/figures/ and metrics in output/pipeline_verified.log.

3. Compile Documentation

To build the PDF report and slides:

# Compile Report
cd report
latexmk -pdf main.tex

# Compile Slides
cd ../slides
latexmk -pdf main.tex

πŸ“„ Methodology Summary

  1. Phase I (SLR): The β€œLoudness War” analysis (R2=0.14R^2=0.14).
  2. Phase II (MLR): Baseline multiple regression (R2=0.24R^2=0.24).
  3. Phase III (Diagnostics): Testing Linearity, Multicollinearity (VIF), and Homoscedasticity (BP Test).
  4. Phase IV (Model Selection): Comparison of Stepwise AIC vs LASSO.
  5. Phase V (Refinement): Implementation of WLS to handle variance instability.

University Statistical Analysis Project | Term: Fall 2024