Skip to content
Research / archived / 2026

ESC-50 Sound Classification & Retrieval

A DSP course project implementing FFT, STFT, mel filterbanks, and MFCCs from scratch, then benchmarking audio classifiers and retrieval models.

Python DSP MFCC CLAP Audio Retrieval
GitHub
Audiofool934/dsp-final
Language
Python
Stars
0
Last Push
2026.03.13
README Sync
2026.05.30
[ Project Brief ]

Overview

ESC-50 Sound Classification & Retrieval is a hands-on audio DSP project.

The project implements the feature pipeline from scratch — FFT, STFT, mel filterbanks, log-mel features, and MFCCs — then compares those hand-built representations with pretrained audio models on classification and retrieval tasks.

Why It Matters

It is a useful grounding project. Modern audio models can look magical, but the underlying signal-processing path still matters. Building the pieces by hand makes the comparison sharper: what can classical DSP features still do, and where do modern embeddings clearly win?

[ Synced from GitHub README ]

Repository Document

source ↗

ESC-50 Sound Classification & Retrieval

Implementing core audio DSP from scratch — FFT, STFT, mel filterbanks, MFCCs — then benchmarking these hand-built features against pretrained audio models (AST, CLAP, PANNs) and LLM baselines (Gemini) on the ESC-50 environmental sound dataset (50 classes, 2000 clips).

CI

Results

Classification — Fold 5 Test Accuracy

MethodModelAccuracy
Transfer learningCLAP → linear probe97.25%
Transfer learningAST → linear probe95.00%
Zero-shotCLAP (text prompts)91.50%
Transfer learningPANNs → linear probe90.50%
Zero-shotGemini Flash (audio input)78.00%
Trained from scratchResNet on custom log-mel75.00%

Retrieval — Fold 5 Queries vs Folds 1–4 Database

Embedding sourceTop-10 PrecisionTop-20 Precision
CLAP99.50%100.00%
AST98.75%99.25%
PANNs97.75%98.00%
CNN (ours)85.50%88.50%
Custom MFCC67.75%79.50%

The custom MFCC pipeline — built entirely without numpy.fft or librosa — achieves 79.5% Top-20 retrieval precision using only cosine similarity on mean/std-pooled cepstral coefficients.

From-Scratch DSP Pipeline

The core of this project is a complete audio feature extraction pipeline with no dependency on FFT/spectral libraries:

graph LR
    A["Raw Audio"] --> B["Pre-emphasis"]
    B --> C["Frame & Window"]
    C --> D["FFT<br/><i>Cooley-Tukey radix-2</i>"]
    D --> E["|X(f)|² Power Spectrum"]
    E --> F["Mel Filterbank"]
    F --> G["Log"]
    G --> H["DCT-II"]
    H --> I["MFCCs"]
    G --> J["Log-Mel Features"]

    style D fill:#2d6a4f,color:#fff,stroke:#1b4332
    style I fill:#52b788,color:#fff
    style J fill:#52b788,color:#fff
ComponentImplementationValidated against
FFTRadix-2 Cooley-Tukey with bit-reversal permutationnumpy.fft.fft — error < 1e-10
STFTWindowed frames via np.lib.stride_tricks (zero-copy)librosa.stft
Mel filterbankTriangular filters on mel-spaced frequency binslibrosa.filters.mel
DCT-IIDirect cosine basis matrix, no scipyscipy.fft.dct

DSP Output Visualizations

All generated by the custom pipeline (src/dsp/), no librosa involved:

Custom STFT spectrogram

Custom log-mel spectrogram

Custom MFCCs

Architecture Overview

graph TD
    A["Raw Audio (.wav)"] --> B["Custom DSP<br/><i>from scratch</i>"]
    A --> C["Pretrained Audio Models"]
    A --> D["LLM Baseline"]

    B --> B1["FFT → STFT → Mel Filterbank"]
    B1 --> B2["Log-Mel Spectrogram"]
    B1 --> B3["MFCCs"]

    C --> C1["AST"]
    C --> C2["CLAP"]
    C --> C3["PANNs"]

    B2 --> E["CNN Classifier"]
    B3 --> F["Cosine Retrieval"]
    C1 & C2 & C3 --> G["Embeddings"]
    G --> H["Linear Probe"]
    G --> F
    D --> D1["Gemini Flash<br/>zero-shot"]

    E & H & D1 --> I["Classification<br/>Fold 5 Accuracy"]
    F --> J["Retrieval<br/>Top-k Precision"]

    style B fill:#2d6a4f,color:#fff
    style B1 fill:#40916c,color:#fff
    style B2 fill:#52b788,color:#fff
    style B3 fill:#52b788,color:#fff
    style C fill:#1d3557,color:#fff
    style C1 fill:#457b9d,color:#fff
    style C2 fill:#457b9d,color:#fff
    style C3 fill:#457b9d,color:#fff
    style D fill:#6c584c,color:#fff
    style D1 fill:#a98467,color:#fff

Quick Start

pip install -e ".[dev]"
# Download ESC-50 → data/ESC-50-master/

make test                  # run test suite
make lint                  # ruff check + format

# Train CNN on custom log-mel features
PYTHONPATH=. python scripts/models/train_cnn.py --epochs 30

# MFCC retrieval sweep
PYTHONPATH=. python scripts/tasks/run_retrieval.py \
  --frame-lengths 512 1024 2048 --hop-lengths 256 512 1024

# Transfer learning (ast / clap / panns)
PYTHONPATH=. python scripts/models/eval_transfer.py --model-type clap

# Run everything end-to-end
PYTHONPATH=. python scripts/tools/run_all_experiments.py --precompute-workers 1

Project Structure

src/
  dsp/          FFT, STFT, MFCC — from scratch, no numpy.fft
  models/       Pretrained model wrappers (AST, CLAP, PANNs) + custom ResNet
  retrieval/    Cosine-similarity retrieval (MFCC and ML embeddings)
  features/     Content-addressed feature cache (compute once, load from disk)
  tasks/        Training loops, evaluation, LLM baseline parsing
  datasets/     ESC-50 metadata loading and fold splits

scripts/
  models/       Train/eval individual models
  tasks/        Run experiment sweeps and grid searches
  tools/        Plotting, precompute, end-to-end orchestration

tests/          DSP correctness vs NumPy, label parsing, retrieval metrics
configs/        Experiment hyperparameters (YAML)