ESC-50 Sound Classification & Retrieval
Implementing core audio DSP from scratch — FFT, STFT, mel filterbanks, MFCCs — then benchmarking these hand-built features against pretrained audio models (AST, CLAP, PANNs) and LLM baselines (Gemini) on the ESC-50 environmental sound dataset (50 classes, 2000 clips).

Results
Classification — Fold 5 Test Accuracy
| Method | Model | Accuracy |
|---|
| Transfer learning | CLAP → linear probe | 97.25% |
| Transfer learning | AST → linear probe | 95.00% |
| Zero-shot | CLAP (text prompts) | 91.50% |
| Transfer learning | PANNs → linear probe | 90.50% |
| Zero-shot | Gemini Flash (audio input) | 78.00% |
| Trained from scratch | ResNet on custom log-mel | 75.00% |
Retrieval — Fold 5 Queries vs Folds 1–4 Database
| Embedding source | Top-10 Precision | Top-20 Precision |
|---|
| CLAP | 99.50% | 100.00% |
| AST | 98.75% | 99.25% |
| PANNs | 97.75% | 98.00% |
| CNN (ours) | 85.50% | 88.50% |
| Custom MFCC | 67.75% | 79.50% |
The custom MFCC pipeline — built entirely without numpy.fft or librosa — achieves 79.5% Top-20 retrieval precision using only cosine similarity on mean/std-pooled cepstral coefficients.
From-Scratch DSP Pipeline
The core of this project is a complete audio feature extraction pipeline with no dependency on FFT/spectral libraries:
graph LR
A["Raw Audio"] --> B["Pre-emphasis"]
B --> C["Frame & Window"]
C --> D["FFT<br/><i>Cooley-Tukey radix-2</i>"]
D --> E["|X(f)|² Power Spectrum"]
E --> F["Mel Filterbank"]
F --> G["Log"]
G --> H["DCT-II"]
H --> I["MFCCs"]
G --> J["Log-Mel Features"]
style D fill:#2d6a4f,color:#fff,stroke:#1b4332
style I fill:#52b788,color:#fff
style J fill:#52b788,color:#fff
| Component | Implementation | Validated against |
|---|
| FFT | Radix-2 Cooley-Tukey with bit-reversal permutation | numpy.fft.fft — error < 1e-10 |
| STFT | Windowed frames via np.lib.stride_tricks (zero-copy) | librosa.stft |
| Mel filterbank | Triangular filters on mel-spaced frequency bins | librosa.filters.mel |
| DCT-II | Direct cosine basis matrix, no scipy | scipy.fft.dct |
DSP Output Visualizations
All generated by the custom pipeline (src/dsp/), no librosa involved:
Architecture Overview
graph TD
A["Raw Audio (.wav)"] --> B["Custom DSP<br/><i>from scratch</i>"]
A --> C["Pretrained Audio Models"]
A --> D["LLM Baseline"]
B --> B1["FFT → STFT → Mel Filterbank"]
B1 --> B2["Log-Mel Spectrogram"]
B1 --> B3["MFCCs"]
C --> C1["AST"]
C --> C2["CLAP"]
C --> C3["PANNs"]
B2 --> E["CNN Classifier"]
B3 --> F["Cosine Retrieval"]
C1 & C2 & C3 --> G["Embeddings"]
G --> H["Linear Probe"]
G --> F
D --> D1["Gemini Flash<br/>zero-shot"]
E & H & D1 --> I["Classification<br/>Fold 5 Accuracy"]
F --> J["Retrieval<br/>Top-k Precision"]
style B fill:#2d6a4f,color:#fff
style B1 fill:#40916c,color:#fff
style B2 fill:#52b788,color:#fff
style B3 fill:#52b788,color:#fff
style C fill:#1d3557,color:#fff
style C1 fill:#457b9d,color:#fff
style C2 fill:#457b9d,color:#fff
style C3 fill:#457b9d,color:#fff
style D fill:#6c584c,color:#fff
style D1 fill:#a98467,color:#fff
Quick Start
pip install -e ".[dev]"
make test
make lint
PYTHONPATH=. python scripts/models/train_cnn.py --epochs 30
PYTHONPATH=. python scripts/tasks/run_retrieval.py \
--frame-lengths 512 1024 2048 --hop-lengths 256 512 1024
PYTHONPATH=. python scripts/models/eval_transfer.py --model-type clap
PYTHONPATH=. python scripts/tools/run_all_experiments.py --precompute-workers 1
Project Structure
src/
dsp/ FFT, STFT, MFCC — from scratch, no numpy.fft
models/ Pretrained model wrappers (AST, CLAP, PANNs) + custom ResNet
retrieval/ Cosine-similarity retrieval (MFCC and ML embeddings)
features/ Content-addressed feature cache (compute once, load from disk)
tasks/ Training loops, evaluation, LLM baseline parsing
datasets/ ESC-50 metadata loading and fold splits
scripts/
models/ Train/eval individual models
tasks/ Run experiment sweeps and grid searches
tools/ Plotting, precompute, end-to-end orchestration
tests/ DSP correctness vs NumPy, label parsing, retrieval metrics
configs/ Experiment hyperparameters (YAML)