Musical Carbon Dating π΅
A Statistical Feature Recognition Analysis (1960-2020)
βCan we timestamp a musical recording purely from its acoustic properties?β
This project implements a rigorous statistical pipeline to quantifiably βcarbon dateβ music. By analyzing 13 physical and perceptual audio features (e.g., Tempo, Valence, Spectral Energy) across 250,971 tracks, we demonstrate that musical eras have distinct, mathematically recognizable acoustic fingerprints.
π Key Results (Verified)
1. Statistical Validity
Despite the complexity of artistic expression, our Weighted Least Squares (WLS) model ensures valid inference:
- R2: 0.275. (27.5% of variance explained by acoustic physics alone).
- MAE: 9.30 years. (Average error < 1 decade).
- See
output/figures/pred_vs_act_best_model_(wls)_predictions.png.
2. Statistical Rigor
We explicitly address the failures of standard OLS regression:
- Heteroscedasticity: Diagnosed via Breusch-Pagan Test (Ο2β19,567,p<0.001) and corrected with WLS weights (wiββ1/Οi2β).
- Non-Linearity: Confirmed via Partial F-Tests (Fβ4945).
- Feature Selection: LASSO (L1β) regularization validated the use of all 13 acoustic features, confirming that even subtle markers (like Key and Mode) are essential for tracking harmonic evolution.
3. The βNostalgia Indexβ
We define prediction error as a commercial metric: Index=β£y^βpredββyactualββ£.
- Insight: High index values identify songs that are βTime-Displacedβ (Retro or Futuristic).
- Examples:
- Uptown Funk (2015): Index 1.9 (Modern construction).
- Physical (Dua Lipa, 2020): Index 11.0 (Strong 80s aesthetic).
π Project Structure
βββ data/ # Dataset (Spotify 600k Tracks)
βββ src/ # Source Code
β βββ analysis.py # Regression Engine (OLS, WLS, Ridge, LASSO, Stepwise)
β βββ config.py # Configuration (Feature Definitions)
β βββ data_loader.py # Data Preprocessing
β βββ visualization.py # Plotting Logic
βββ report/ # [FINAL] Formal LaTeX Report
β βββ main.tex # Comprehensive Academic Report
β βββ figures/ # Auto-generated Figures
βββ slides/ # Presentation Slides (Beamer)
β βββ main.tex # "No Hiding" Detailed Slides
β βββ speaker_notes.md # 22-min Verbatim Script (5 Speakers)
βββ output/ # Generated Artifacts
β βββ figures/ # Residual Plots, Q-Q Plots, Prediction Plots
β βββ tables/ # CSV Results
β βββ pipeline_verified.log # Definitive Statistical Output
βββ main.py # Main Execution Pipeline
π Usage
1. Environment Setup
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
2. Run Main Analysis
Execute the full pipeline (Data Load -> Feature Selection -> WLS -> Diagnostics):
python3 main.py
This will generate all figures in output/figures/ and metrics in output/pipeline_verified.log.
3. Compile Documentation
To build the PDF report and slides:
cd report
latexmk -pdf main.tex
cd ../slides
latexmk -pdf main.tex
π Methodology Summary
- Phase I (SLR): The βLoudness Warβ analysis (R2=0.14).
- Phase II (MLR): Baseline multiple regression (R2=0.24).
- Phase III (Diagnostics): Testing Linearity, Multicollinearity (VIF), and Homoscedasticity (BP Test).
- Phase IV (Model Selection): Comparison of Stepwise AIC vs LASSO.
- Phase V (Refinement): Implementation of WLS to handle variance instability.
University Statistical Analysis Project | Term: Fall 2024