Can You See The Music
Without music, life would be a mistake - Friedrich Nietzsche
Background
Music, an indispensable part of human culture, has long been a powerful medium for expressing emotions, thoughts, and cultural values. As Friedrich Nietzsche put it, “Without music, life would be a mistake,” highlighting the profound significance of music in our lives. From the harmonious symphonies of classical music to the dynamic rhythms of modern pop music, music has evolved over time, adapting to cultural changes and technological advancements.
full report (in Chinese)
Approach
This project aims to fill this research gap by comprehensively studying the direct and indirect attributes of music through data visualization analysis. By analyzing three different datasets, we attempt to answer the following questions: How do the themes and languages in lyrics evolve over time? What are the differences in audio features among different music styles? What are the main emotions expressed in popular music, and how do these emotions change over time? How do the careers and collaboration patterns of artists develop? What can we learn about users’ preferences from music playlists?
We will first introduce the related research work in the field of music data analysis. Then, it will elaborate on the methods of classifying and analyzing music attributes, followed by an in-depth discussion of the datasets used. The main part of the report will present the results of data visualization and analysis for each key music attribute, such as lyrics, audio features, sentiment, artists, and playlists. Finally, conclusions will be drawn based on the research results, highlighting the contributions of this study and suggesting potential directions for future research.
To make the entire analysis process more logical, the attributes of music are classified as shown in the following table:
| Direct Attributes | lyrics, features, album, artist, release time |
|---|
| Indirect Attributes | sentiment, genre, playlist, popularity |
| Table 1: Music Attribute Classification | |
- Direct attributes cover the following aspects:
- lyrics: By deeply analyzing the lyrics content, we can excavate the emotions and themes contained therein.
- features: Audio features (such as frequency, loudness, etc.). With the help of professional technical means, we extract key features of the audio, such as rhythm, pitch, and loudness, to analyze the differences in audio features among different music genres and styles.
- album, artist, release time: These belong to basic metadata.
- Indirect attributes mainly involve the following contents:
- sentiment: Conduct sentiment analysis based on lyrics text or audio data to reveal the emotional trends presented in different musical works.
- genre: The style or genre of a song (usually a song may contain multiple styles, and the definition is not absolute). By studying the distribution and evolution of genres, we can gain insights into the musical style characteristics of different periods and the cultural development trends reflected behind them.
- playlist: That is, a music playlist, which is the result of music listeners rearranging existing music (this is very common among users of music streaming platforms).
- popularity: It is used to measure the popularity of a song. It is affected by many factors, and its definition is not unique. Common measurement methods include referring to song sales, rankings on major music charts, etc.
Datasets
In this project, we utilized three datasets to analyze various aspects of music:
-
50 Years of Pop Music Lyrics: This dataset comprises the lyrics of songs listed in Billboard’s Year-End Hot 100 charts from 1965 to 2015, totaling 5,100 entries. Each entry includes the song’s rank, title, artist, year, lyrics, and source. The lyrics were primarily sourced from websites such as metrolyrics.com, songlyrics.com, and lyricsmode.com. Approximately 3.6% of the lyrics were unavailable. The dataset provides a comprehensive overview of popular music trends over five decades.
-
Million Playlist Dataset (MPD): Released by Spotify as part of the RecSys Challenge 2018, the MPD contains 1,000,000 user-generated playlists. Each playlist includes information such as playlist title, track titles, and track ordering. The dataset is instrumental in developing and evaluating music recommendation systems, particularly for tasks like automatic playlist continuation.
-
GTZAN Music Genre Classification Dataset: This dataset is a benchmark for music genre classification tasks. It consists of 1,000 audio tracks, each 30 seconds long, evenly distributed across 10 genres: blues, classical, country, disco, hip-hop, jazz, metal, pop, reggae, and rock. The dataset is widely used for evaluating machine learning models in genre recognition studies.
These datasets collectively provide a robust foundation for analyzing lyrical content, user engagement through playlists, and genre classification, thereby facilitating a comprehensive exploration of music trends and patterns.
Data Visualization and Analysis Ⅰ: Lyrics
Data Visualization and Analysis Ⅱ: Features
Data Visualization and Analysis Ⅲ: Sentiment
Data Visualization and Analysis Ⅳ: Artist
Data Visualization and Analysis Ⅴ: Playlist
Can You See The Music - artwork inspired by the movie Oppenheimer, "Can You Hear The Music"