Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “genre classification”
ITTV - A Dataset of Italian Television for Automatic Genre Classification
<p>ITTV is a publicly available dataset of Italian TV programs introduced in </p> <blockquote> <p>Alessandro Ilic Mezza, Paolo Sani, and Augusto Sarti, "Automatic TV Genre Classification Based on Visually-Conditioned Deep Audio Features," in 2023 31st European Signal Processing Conference (EUSIPCO), 2023.</p> </blockquote> <p>ITTV consists of 2625 manually annotated YouTube videos, totaling over 670 hours. Each clip is assigned one of seven classes:</p> <ul> <li>Cartoons</li> <li>Commercials</li> <li>Football</li> <li>Music</li> <li>News</li> <li>Talk Shows</li> <li>Weather Forecast</li> </ul> <p>ITTV genre taxonomy is similar to that of the well-known RAI dataset described in</p> <blockquote> <p>Maurizio Montagnuolo and Alberto Messina, "Parallel neural networks for multimodal video genre classification,” Multimedia Tools and Applications, vol. 41, no. 1, pp. 125–159, 2009.</p> </blockquote> <p>The dataset contains genre annotations and metadata in CSV format. Please note that audio data is not provided.</p> <p>We provide the annotations for a balanced training (1575 clips) and validation (525 clips) split, as well as for a disjoint test set containing 525 installments from TV programs not included in the development set.</p> <p>As YouTube continuously updates, some videos may not be available in the future. Although we intend to keep ITTV updated as best as possible, please note that some content may not be available at any given time.</p> <p>Some YouTube videos (especially from the <code>Football</code> class and, to a lesser extent, the <code>Cartoons</code> class) may only be available in some countries due to regional restrictions imposed by the content creator. All videos are known to be accessible from Italy (last accessed on Nov. 25th, 2022.)<br> <br> Please contact Alessandro Ilic Mezza for further questions (e-mail: alessandroilic.mezza@polimi.it).</p>
MSD-I: Million Song Dataset with Images for Multimodal Genre Classification
<p>The Million Song Dataset (https://labrosa.ee.columbia.edu/millionsong/) is a collection of metadata and precomputed audio features for 1 million songs. Along with this dataset, a dataset with annotations of 15 top-level genres with a single label per song was released. In our work, we combine the CD2c version of this genre datase (http://www.tagtraum.com/msd_genre_datasets.html) with a collection of album cover images. </p> <p><br> The final dataset contains 30,713 tracks from the MSD and their related album cover images, each annotated with a unique genre label among 15 classes. Based on an initial analysis on the images, we identified that this set of tracks is associated to 16,753 albums, yielding an average of 1.8 songs per album.</p> <p>We randomly divide the dataset into three parts: 70% for training, 15% for validation, and 15% for test, with no artist and album overlap across these sets. This is crucial to avoid possible overfitting, as the classifier may learn to predict the artist instead of the genre. </p> <p> </p> <p>Content:</p> <p>MSD-I dataset (mapping, metadata, annotations and links to images)<br> Data splits and feature vectors for TISMIR single-label classification experiments </p> <p>These data can be used together with the Tartarus deep learning python module https://github.com/sergiooramas/tartarus.</p> <p> </p> <p>Scientific References:</p> <p>Please cite the following paper if using MSD-I dataset or Tartarus software.</p> <p>Oramas, S., Barbieri, F., Nieto, O., and Serra, X (2018). Multimodal Deep Learning for Music Genre Classification, Transactions of the International Society for Music Information Retrieval, V(1).</p>
From Networks of Texts to Networks of Genres? On the Classification of Texts in Compilations with a View towards Manuscript Transmission
<p>At the end of the 14th century, Jakob Twinger von Königshofen, a cleric from Strasbourg, composed a chronicle in the vernacular that spread widely – up to today nearly 130 manuscripts are known that contain the text, wholly or in parts, and that were produced not only in Strasbourg, but as far as Cologne, Augsburg, or Tyrol. About thirty qualify as true copies, while in the big majority of the witnesses, the text is altered in various ways: abbreviated, augmented, updated, corrected, put in a dierent order, and more often than not combined with other texts, either with distinct text boundaries or resulting in new compositions made of several texts.<br> In several manuscripts, a combination of historiographical texts – chronicles, annals, lists, etc. – can be observed, leading to the assumption that Twinger’s work was preferably copied for historiographical compilations and that its structure facilitated some historiographical activity of the recipients. But this view puts a big weight on the genre of Twinger’s work, making it a filter through which all the other texts in a codex are seen. As I have been working on the chronicle transmission, the attribute in common of the known manuscripts is of course the appearance of at least some lines of the Twinger chronicle; but the attempt to look at a single codex as objectively as possible, without giving preference to a particular text, can not only reveal connections between different manuscripts, but also the fluidity and flexibility of medieval texts. The classification of a work as part of a certain genre does not necessarily hold true for its entire transmission, for the manifestation of a distinct text – what is, on the one hand, a nuisance, but on the other a chance to better understand medieval text transmission, manuscript production and the transfer of knowledge.<br> The co-occurrence of certain texts in several manuscripts hints towards intentional copying processes that are reflecting particular interests not only of one individual scribe or commissioner; multiple occurrences of particular compilations can reveal networks that go unseen if the content of a codex is not regarded as a whole. While a text-based analysis often faces difficulties that result from insufficient manuscript descriptions, a broader view that would compare less particular texts, but more areas of interest or fields of knowledge, has to deal with problems regarding the classification of the single texts: Apart from an inevitable subjectivity,<br> questions about the criteria and levels of classification have to be addressed, while the danger of over- or under-representation is always lurking. Discussing these issues and the applicability and usefulness of the comparison of codices from a kind-of-genre-perspective could be fruitful to develop a better understanding for the transmission of manuscripts, of texts and of knowledge.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.