Multi-modal dataset for music genre recognition based on six different modalities for LMD-aligned and SLAC datasets
<p>Multi-modal dataset for music genre recognition based on six different modalities for the LMD-aligned [1] and SLAC [2] datasets. Further details are provided in [3].</p>
<p><strong>Descriptions of files</strong></p>
<table>
<thead>
<tr>
<th scope="col">Link</th>
<th scope="col">Description</th>
</tr>
</thead>
<tbody>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/LMD-aligned_Filelist.arff">LMD-aligned_Filelist.arff</a></td>
<td>File list with 1575 music tracks selected from the LMD-aligned dataset [1] with tagtraum genre annotations [4] (only a subset of LMD-aligned is used, which includes only pieces for which all six modalities were accessible, and which includes only well-represented genres)</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/LMD-aligned_ExtractedFeatures.tar.gz">LMD-aligned_ExtractedFeatures.tar.gz</a></td>
<td>Raw audio signal and model-based features extracted with AMUSE [5]</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/LMD-aligned_ProcessedFeatures.tar.gz">LMD-aligned_ProcessedFeatures.tar.gz</a></td>
<td>Processed features: audio signal and model-based features aggregated for 4 s time frames with 2 s step size / all other features (see the table below) with the same values for all time frames</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/LMD-aligned_Datasets.tar.gz">LMD-aligned_Datasets.tar.gz</a></td>
<td>Training, optimization, and test datasets for 3 splits for the recognition of 5 genres in [3]</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/SLAC_Filelist.arff">SLAC_Filelist.arff</a></td>
<td>File list with 250 music tracks from the SLAC dataset [2] (genres and sub-genres are provided in the folder structure)</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/SLAC_ExtractedFeatures.tar.gz">SLAC_ExtractedFeatures.tar.gz</a></td>
<td>Raw audio signal and model-based features extracted with AMUSE [5]</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/SLAC_ProcessedFeatures.tar.gz">SLAC_ProcessedFeatures.tar.gz</a></td>
<td>Processed features: audio signal and model-based features aggregated for 4 s time frames with 2 s step size / all other features (see the table below) with the same values for all time frames</td>
</tr>
<tr>
<td><a href="https://zenodo.org/record/5651429/files/SLAC_Datasets.tar.gz">SLAC_Datasets.tar.gz</a></td>
<td>Training, optimization, and test datasets for 3 splits for the recognition of 5 genres and 10 sub-genres in [3]</td>
</tr>
</tbody>
</table>
<p><strong>Modalities and feature sub-groups</strong></p>
<table>
<thead>
<tr>
<th scope="col">Modality</th>
<th scope="col">Sub-group</th>
<th scope="col">
<p>Dimensions in processed</p>
<p>features of LMD-aligned</p>
</th>
<th scope="col">
<p>Dimensions in processed</p>
<p>features of SLAC</p>
</th>
</tr>
</thead>
<tbody>
<tr>
<td>Audio signal</td>
<td>Low-level</td>
<td>1-524</td>
<td>1-524</td>
</tr>
<tr>
<td>Audio signal</td>
<td>Semantic</td>
<td>525-810</td>
<td>525-810</td>
</tr>
<tr>
<td>Audio signal</td>
<td>Structural complexity</td>
<td>811-908</td>
<td>811-908</td>
</tr>
<tr>
<td>Model-based</td>
<td>Instruments</td>
<td>909-1018</td>
<td>909-1018</td>
</tr>
<tr>
<td>Model-based</td>
<td>Moods</td>
<td>1019-1146</td>
<td>1019-1146</td>
</tr>
<tr>
<td>Model-based</td>
<td>Various</td>
<td>1147-1402</td>
<td>1147-1402</td>
</tr>
<tr>
<td>Playlists</td>
<td>Genres</td>
<td>1403-1973</td>
<td>1403-1973</td>
</tr>
<tr>
<td>Playlists</td>
<td>Styles</td>
<td>1974-1695</td>
<td>1974-1695</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Pitch</td>
<td>1696-1757</td>
<td>1696-1757</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Melodic</td>
<td>1758-1781</td>
<td>1758-1781</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Chords</td>
<td>1782-1836</td>
<td>1782-1836</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Rhythm</td>
<td>1837-1935</td>
<td>1837-1935</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Tempo</td>
<td>1936-1963</td>
<td>1936-1963</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Instrument presence</td>
<td>1964-2441</td>
<td>1964-2441</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Instruments</td>
<td>2442-2456</td>
<td>2442-2456</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Texture</td>
<td>2457-2480</td>
<td>2457-2480</td>
</tr>
<tr>
<td>Symbolic</td>
<td>Dynamics</td>
<td>2481-2484</td>
<td>2481-2484</td>
</tr>
<tr>
<td>Album covers</td>
<td>SIFT</td>
<td>2485-2584</td>
<td>2485-2584</td>
</tr>
<tr>
<td>Lyrics</td>
<td>jLyrics descriptors</td>
<td>2585-2603</td>
<td>2585-2671</td>
</tr>
<tr>
<td>Lyrics</td>
<td>Bag-of-Words</td>
<td>2604-2703</td>
<td> </td>
</tr>
<tr>
<td>Lyrics</td>
<td>Doc2Vec</td>
<td>2704-2803</td>
<td> </td>
</tr>
</tbody>
</table>
opencc-by-4.0Nov 2021View details →