Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,250
datasets available to search
ShareScore release 0.7.1
Dataset results
6,250 results for “Classification”
AntarcticBasins v1.04: Antarcitca's Sedimentary Basins Distribution and Classification
<p>GIS package for Antarctic Sedimentary Basins Distribution and Classification</p> <p>Supplement to: Aitken, A. R. A., Li, L., Kulessa, B., Schroeder, D., Jordan, T. A., Whittaker, J. M., et al. (2023). Antarctic sedimentary basins and their influence on ice-sheet dynamics. <em>Reviews of Geophysics</em>, 61, e2021RG000767. <a href="https://doi.org/10.1029/2021RG000767">https://doi.org/10.1029/2021RG000767</a></p>
Data set and classification method for low quality web traffic identification in video marketing campaigns
<p>Final outcomes of the InPreVi (AI4Media) project developed in 2022. </p> <p>1. Data set describing the statistics of the video ad marketing campaigns</p> <p>2. Script for web traffic classification</p>
Antarctic Sedimentary Basin Distribution and Classification
<p>This is the published version (<a href="https://zenodo.org/record/7984586">v1.04</a>) of the GIS package for Antarcitca's Sedimentary Basins Distribution and Classification. </p> <p>Supplement to: Aitken, A. R. A., Li, L., Kulessa, B., Schroeder, D., Jordan, T. A., Whittaker, J. M., et al. (2023). Antarctic sedimentary basins and their influence on ice-sheet dynamics. <em>Reviews of Geophysics</em>, 61, e2021RG000767. <a href="https://doi.org/10.1029/2021RG000767">https://doi.org/10.1029/2021RG000767</a></p> <p>With the release of the published version of the GIS package, future updates to the sedimentary basin mapping can be found at <a href="https://github.com/LL-Geo/AntarcticBasins">https://github.com/LL-Geo/AntarcticBasins</a>, and at <a href="https://doi.org/10.5281/zenodo.7955525">https://doi.org/10.5281/zenodo.7955525</a>.</p> <p>You can download individual GeoTIFF, Shapefile, and GeoJSON files. The complete DistroPackage contains all files, including styles in QGIS and ArcGIS projects.</p> <p> </p>
ITTV - A Dataset of Italian Television for Automatic Genre Classification
<p>ITTV is a publicly available dataset of Italian TV programs introduced in </p> <blockquote> <p>Alessandro Ilic Mezza, Paolo Sani, and Augusto Sarti, "Automatic TV Genre Classification Based on Visually-Conditioned Deep Audio Features," in 2023 31st European Signal Processing Conference (EUSIPCO), 2023.</p> </blockquote> <p>ITTV consists of 2625 manually annotated YouTube videos, totaling over 670 hours. Each clip is assigned one of seven classes:</p> <ul> <li>Cartoons</li> <li>Commercials</li> <li>Football</li> <li>Music</li> <li>News</li> <li>Talk Shows</li> <li>Weather Forecast</li> </ul> <p>ITTV genre taxonomy is similar to that of the well-known RAI dataset described in</p> <blockquote> <p>Maurizio Montagnuolo and Alberto Messina, "Parallel neural networks for multimodal video genre classification,” Multimedia Tools and Applications, vol. 41, no. 1, pp. 125–159, 2009.</p> </blockquote> <p>The dataset contains genre annotations and metadata in CSV format. Please note that audio data is not provided.</p> <p>We provide the annotations for a balanced training (1575 clips) and validation (525 clips) split, as well as for a disjoint test set containing 525 installments from TV programs not included in the development set.</p> <p>As YouTube continuously updates, some videos may not be available in the future. Although we intend to keep ITTV updated as best as possible, please note that some content may not be available at any given time.</p> <p>Some YouTube videos (especially from the <code>Football</code> class and, to a lesser extent, the <code>Cartoons</code> class) may only be available in some countries due to regional restrictions imposed by the content creator. All videos are known to be accessible from Italy (last accessed on Nov. 25th, 2022.)<br> <br> Please contact Alessandro Ilic Mezza for further questions (e-mail: alessandroilic.mezza@polimi.it).</p>
A Novel Approach to Heart Failure Prediction and Classification through Advanced Deep Learning Model
<p>A Novel Approach to Heart Failure Prediction and Classification through Advanced Deep Learning Model</p>
Data for "A New Year-Round Weather Regime Classification for North America"
<p>This dataset supports the work presented in <em>"A New Year-Round Weather Regime Classification for North America" </em>by Lee et al. (2023), published in <em>Journal of Climate</em>: <a href="https://doi.org/10.1175/JCLI-D-23-0214.1">https://doi.org/10.1175/JCLI-D-23-0214.1</a> </p> <p>The dataset includes the Z500 climatology and anomalies, the EOFs/PCs, and the parameters required to attribute each day to a regime, alongside the regimes as assigned over 1979-2022 (time series in netCDF and csv formats). In addition, we also supply the standardized projection of each day onto the cluster-mean normalized Z500 anomalies for each regime, termed the Weather Regime Index (WRI).</p>
Phlorest phylogeny derived from Michael et al. 2015 'A Bayesian Phylogenetic Classification of Tupi-Guarani'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Michael L, Chousou-Polydouri N, Bartolomei K, Donnelly E, Wauters V, Meira S & O'Hagan Z. 2015. A Bayesian Phylogenetic Classification of Tupi-Guarani. LIAMES 15(2):1–36.</p> </blockquote>
Bar graphs of DEMIX database tile classifications
<p>The DEMIX database (Guth, 2023a) contains statistics from 6 test 1 arc second DEMs (ALOS, ASTER, CopDEM, FABDEM, NASADEM, and SRTM) compared to high resolution reference DEMs. The database contains 236 DEMIX tiles (Guth and others, 2023) and forms the basis for the ranking of global DEMs in Bielski and others (2023).</p> <p>A K-means clustering of the database using MICRODEM (Guth, 2023b, 2023c), and an additional set of 4 land cover and landform classifications (Table 1) for the 236 DEMIX tiles computed the percentage of each DEMIX tile in each classification category. Guth (2023d) has the raw data for the percentages of each category for each of the 236 tiles along with the K-means cluster assignments.</p> <p>This data set contains 3 figures for each of the 5 classification databases in Guth (2023d):</p> <ul> <li>Bar graph of the category percentages for each of the test areas. A composite version of these graphs is in Bielski and others (2023).</li> <li>Bar graph of the category percentages for each of the 236 test tiles. These graphs are too large to include on a single page with readable legends.</li> <li>Legend for the classification</li> </ul> <p> </p> <p>References:</p> <p>Bielski, C.; López-Vázquez, C.; Grohmann, C.H.; Guth. P.L.; and the TMSG DEMIX Working Group, 2023. DEMIX Method Ranks COPDEM, and FABDEM as Top 1” Global DEMs: <a href="https://arxiv.org/abs/2302.08425v3">https://arxiv.org/abs/2302.08425v3</a></p> <p>Guth, P. L., 2023a. DEMIX GIS Database Version 2 (2.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8062008">https://doi.org/10.5281/zenodo.8062008</a></p> <p>Guth, P.L., 2023b. GIT-MICRODEM [Delphi source code, archived installation versions]. URL: https://github.com/prof-pguth/git_microdem </p> <p>Guth, P.L., 2023c, MICRODEM: Open-source GIS with a focus on Geomorphometry [download latest Win64 executable and CHM help file] URL: <a href="https://microdem.org/">https://microdem.org/</a></p> <p>Guth, P.L., 2023d, K-means clustering of the DEMIX data set (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8283791</p> <p>Guth, Peter L., Peter Strobl, Kevin Gross, & Serge Riazanoff. (2023). DEMIX 10k Tile Data Set (1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7504791">https://doi.org/10.5281/zenodo.7504791</a></p>
DataCI Continuous Text Classification Example Using Yelp Dataset
<p>We are using the <a href="https://www.yelp.com/dataset">Yelp Review Dataset</a> as the streaming data source for the DataCI example. We have processed the Yelp review dataset into a daily-based dataset by its `date`. In this dataset, we will only use the data from 2020-09-01 to 2020-11-30 to simulate the streaming data scenario. We are downloading two versions of the training and validation datasets:</p> <ul> <li>`yelp_review_train@2020-10`: from 2020-09-01 to 2020-10-15</li> <li>`yelp_review_val@2020-10`: from 2020-10-16 to 2020-10-31</li> <li>`yelp_review_train@2020-11`: from 2020-10-01 to 2020-11-15</li> <li>`yelp_review_val@2020-11`: from 2020-11-16 to 2020-11-30</li> </ul>
BgMA-ESy: Expert system for automatic classification of vegetation plots of subalpine tall-herb vegetation (class Mulgedio-Aconitetea) from Bulgaria
<p>*****</p> <p>BgMA-ESy is an expert system that classifies vegetation plots of the class <em>Mulgedio-Aconitetea</em> (<a href="https://doi.org/10.1111/avsc.12257">Mucina et al. 2016</a>) occurring in Bulgaria. The expert system can be run using the JUICE program (<a href="https://doi.org/10.1111/j.1654-1103.2002.tb02069.x">Tichý 2002</a>; <a href="https://www.sci.muni.cz/botany/juice/">https://www.sci.muni.cz/botany/juice/</a>).</p> <p>The aggregation of vascular plants included within the BgMA-ESy is adopted from EUNIS-ESy (<a href="https://doi.org/10.1111/avsc.12519">Chytrý et al. 2020</a>; <a href="https://doi.org/10.5281/zenodo.4812736">https://doi.org/10.5281/zenodo.4812736</a>), and in a few cases, it is adjusted.</p> <p>*****</p> <p><strong>Specifications</strong></p> <p>The analyzed data (vegetation plots) cannot: </p> <ul> <li>include scrub vegetation (cover of tall shrub species > 8%; e.g., <em>Pinus mugo</em>, <em>Salix </em>spp.).</li> <li>contain tree species with cover > 1% (e.g., <em>Fagus sylvatica</em>, <em>Picea abies</em>).</li> <li>contain <em>Pteridium aquilinum </em>as a dominant species.</li> </ul> <p>The expert system was trained on vegetation plots with 5–100 m<sup>2</sup> area that occur above 1000 m a. s. l.</p> <p>* Exceptions from EUNIS-ESy aggregation:</p> <p>Heracleum sphondylium agg. does not include H. sphondylium subsp. verticillatum.</p> <p> </p> <p>*****</p> <p>When using this work, please cite:</p> <p>Szokala D., Kočí M. & Vassilev K. (2024): Subalpine tall-herb vegetation in Bulgaria: diversity and ecology. – Plant Biosystems 158: 490–510. <a href="https://doi.org/10.1080/11263504.2024.2327865">https://doi.org/10.1080/11263504.2024.2327865</a>.</p> <p>*****</p>
Image-based Classification of Intense Radio Bursts from Spectrograms: An Application to Saturn Kilometric Radiation
<p>A catalogue of 4874 of the Low Frequency Extensions (LFEs) of Saturn Kilometric Radiation (SKR) detected by Cassini/RPWS from the beginning of 2004 until mission end in 2017. The LFEs presented in this catalogue were identified using a modified U-Net architecture that applied semantic segmentation to spectrogram images in order to extract the exact frequency-time coordinates of the LFE. The files consist of a .json file with the coordinates of each LFE in Time Frequency Catalogue (TFCat) format (Cecconi et. al. 2023). We also include a .csv file with the start and stop times of each LFE in the form of python datetime timestamps, with the average predicted probability per LFE as an accompanying column. </p>
Report on Transformers interpretability for Natural Language Processing: A case study on Technical Debt classification
<p>Transformer models have significantly advanced the field of natural language processing (NLP), achieving exceptional results in various tasks. However, these models are often seen as "black boxes", providing limited insight into the factors influencing their predictions. It has become crucial to develop and utilise methods for interpreting and explaining these models to uncover their complex inner workings. This report discusses the latest techniques and tools that aid in a more profound understanding of transformer models within NLP. Additionally, it explores a vital industrial use case: Technical Debt (TD) classification. In this context, the report leverages transformer model interpretability tools and Retrieval Augmented Generation (RAG) to analyse and understand the characteristics of text in Github issues, distinguishing between TD and non-TD.</p> <p>This report thoroughly outlines an approach to improve the transparency and reproducibility of machine learning models, with a special emphasis on TD classification. It integrates the RAG approach and exploits feature attribution techniques, presenting a route to create AI systems that are not only high-performing but also demonstrably trustworthy and comprehensible. Through a detailed examination of word patterns in TD classification and the innovative use of the RAG approach, the research highlights a strong dedication to promoting transparency and responsibility in AI systems, potentially ushering in a new phase in machine learning research that focuses on clarity and dependability.</p>
Galaxy Zoo DESI: Detailed Morphology Classifications for 8.7M Galaxies in the DESI Legacy Imaging Surveys
<p>This repository contains the data released in the paper "Galaxy Zoo DESI: Detailed Morphology Classifications for 8.7M Galaxies in the DESI Legacy Imaging Surveys" <em>(DOI to follow on publication).</em></p> <p>We release detailed morphology measurements for bright (<em>r </em>< 19) galaxies in the DESI Legacy Imaging Surveys footprint. These measurements estimate the presence of bars, spirals arms, ongoing mergers, and more.</p> <p>---</p> <p><strong>GZ DESI Detailed Morphology Catalogs</strong></p> <p>These catalogs are created by training deep learning models on Galaxy Zoo volunteer responses, to predict what volunteers might say for new galaxies. The models are available at [www.github.com/mwalmsley/zoobot](www.github.com/mwalmsley/zoobot). Our measurements are predicted vote fractions i.e. the fraction of volunteers expected to select a given answer for a given question.</p> <p>We share two catalog versions containing the same morphology measurements but presented in different ways.</p> <p>gz_desi_deep_learning_catalog_friendly.parquet contains the morphology measurements</p> <p>gz_desi_deep_learning_catalog_advanced.parquet contains the same measurements, and additional information:</p> <p>- _friendly includes only relevant vote fractions, defined as vote fractions to answers of questions that a majority of volunteers would have been asked. This removes predicted vote fractions for e.g. the fraction of volunteers answering "2 spiral arms" to a galaxy with no spiral arms. _advanced includes all vote fractions and instead reports the (column "proportion_asked"). The user must select which vote fractions they consider relevant (we suggest proportion_asked > 0.5, which recovers the _friendly fractions).</p> <p>- _advanced includes columns with estimated credible intervals (error bars) around each vote fraction. These are calculated from the vote fraction posterior predicted by our models.</p> <p>Finally, we separately present volunteer votes collected for 96k galaxies during the GZD-8 campaign, i.e. after the release of GZ DECaLS but before this (GZ DESI) release. These are split into the _core and _extended catalogs, where _extended includes galaxies which received five or more votes for "artifact". The models above were trained on these votes as well as votes from GZ DECaLS.</p> <p>---</p> <p><strong>External Catalog</strong></p> <p>For convenience, we also include an additional catalog of non-morphology measurements created by other authors (external_catalog.parquet) cross-matched to our morphology catalogs. Please credit those authors if you use this catalog (references are in the GZ DESI paper).</p> <p>A particularly important external measurement is redshift. Morphology is increasingly hard to resolve at higher redshift and so <strong>distant galaxies appear less featured</strong>. external_catalog.parquet includes the column "redshift", which is the SDSS spectroscopic redshift where available and a photometric redshift estimate otherwise (again, see the GZ DESI paper for references and credit). You may want to select only galaxies at lower redshifts.</p> <p>---</p> <p><strong>Data Notes</strong></p> <p>Parquet is a fast csv-like format which can be read with pd.read_parquet(loc, columns=[some columns]). Parquet files are read column-by-column (rather than row-by-row) and so you can chose which columns to load. You can easily check which columns are available using columns=['foo'] and reading the error message. We suggest loading only the columns you need when working with the larger catalogs. This will require much less memory than loading every column.</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI to follow on publication) when using the data in this repository.</p> <p>---</p> <p><strong>History</strong></p> <p>v0.0.1 - closed pre-release for internal review</p> <p>v1.0.0 - draft public release. Removed low-z pre-filtered catalogs.</p> <p>v1.0.1 - first public release. Added .csv version of _friendly catalog. Tweaked catalog formatting for clarity and consistency.</p>
Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper
<p>Corpora used in the publication:</p> <ul> <li>Cristina España-Bonet. 2023. <strong>Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper. </strong>In <em>Findings of the Association for Computational Linguistics: EMNLP 2023</em>, Singapore. Pages 11757–11777. Association for Computational Linguistics.</li> </ul> <p>Three corpora are included:</p> <ol> <li>Newspaper articles extracted from the OSCAR corpus in English, German, Spanish and Catalan automatically annotated for political stance (left vs right) and topic</li> <li>Newspaper-like article generations by different versions of ChatGPT for 101 topics in the 4 languages</li> <li>Newspaper-like article generations by Bard for 101 topics in the 4 languages</li> </ol> <p>See the README file and the original article for further details.</p>
ECOD Classification of AFDB 48 Proteomes
<p>ECOD domains classified for the 48 whole proteomes (model_v4) from AFDB. Domains were classified using Domain Parser for Alphafold Models. </p>
Mediterranean land system classification on 2005 and 2015
<p>This shapefile represent the classification of the existing Mediterranean land systems on 2005 and 2015</p>
A didactical dataset to learn supervised classification with candy
<h2>A didactical dataset to learn supervised classification</h2><p>It was obtained from university level students measuring candy that was mixed and distributed in bowls to them. The goal of this dataset creation was to expose the students to the data taking process. Further, the dataset is meant for classification.</p><h3>Dataset Structure</h3><p>The dataset consists of 6 csv files:</p><ul><li><strong>peanuts.csv</strong> represents the entire dataset (a concatenation of all group?.csv files) omitting the sample column</li><li><strong>peanuts_all.csv</strong> represents the entire dataset (a concatenation of all group?.csv files)</li><li>files matching <strong>group[1-5].csv </strong>represent the measurements of each group</li></ul><h3>Data Representation</h3><p>Each file contains 5 columns. </p><ul><li>color, int values, 0: white, 1: black, 2: brown, 3: other</li><li>shape, int values, 0: irregular, 1 round, 2: lens-like</li><li>height, float values, in millimeter</li><li>width, float values, in millimeter</li><li>label, category, peanut/nopeanut</li></ul><p>For more information on the didactical background, see the <a href="https://proceedings.mlr.press/v141/huppenkothen21a.html">original publication</a> that presented the concept for this activity.</p>
Venus coronae topographic (a)symmetry classification (from Gülcher et al., 2023, JGR Planets)
<p>This is a PDF file of the coronae classification that accompanies the manuscript "<strong>Tectono-magmatic evolution of asymmetric coronae on Venus: Topographic classification and 3D thermo-mechanical modeling</strong>" by Gülcher et al. (2023) in <i>Journal of Geophysical Research: Planets</i>, 128, e2023JE007978, <a href="https://doi.org/10.1029/2023JE007978">https://doi.org/10.1029/2023JE007978</a><i> </i><br><br>This database consists of the 150 largest coronae (those with a diameter equal to or larger than 300 km) in the publicly available Venusian coronae nomenclature database (USGS Planetary Nomenclature, (<a href="https://planetarynames.wr.usgs.gov/Page/VENUS/target"><i>https://planetarynames.wr.usgs.gov/Page/VENUS/target</i></a>) and the database of Stofan et al. (1992, <i>JGR, </i><a href="https://doi.org/10.1029/92je01314">https://doi.org/10.1029/92je01314</a>) combined, and five additional smaller coronae. The (a)symmetry of these coronae is defined based on the topographic features (e.g., troughs, rims, rises) and their variability across the coronae. For further information on this classification, please see the main paper. The global distribution of this classification is illustrated in Figure 1 in the main paper and Figure S1 in the Supplementary Information SI1. </p>
DMS measurements dataset for manuscript "Classification of Volatile Organic Compounds by Differential Mobility Spectrometry Based on Continuity of Alpha Curves"
<p>Differential mobility spectrometry dispersion plots collected for the manuscrit "Classification of Volatile Organic Compounds by Differential Mobility Spectrometry Based on Continuity of Alpha Curves". The measurement files are located in the folders that represent certain week and day of measurement. The folders containing measurements are named with the following pattern: [chemical abbreviation]_[dilution rate]. For example "2PEtOH_1o10k" means that the folder contains measurement of 2-phenylethanol diluted with propylene glycol in volumetric proportion 1/10 000. Another example is "nBuOH_1o100" - n-Butanol diluted with propylene glycol in volumetric proportion 1/100. Please find the abbreviations in the article referred.</p> <p>For the first five weeks only 1/100 dilutions were measured. The last two weeks (weeks 6 and 7) 1/10 000 dilutions were measured. However, Carvone 1/100 was measured again on week 6 due to suspicion of faulty measurements during the previous weeks. The faulty measurements were not confirmed, and thus there 25 more samples of Carvone with dilution rate 1/100.</p>
Vegetation classification, Andrews Experimental Forest and vicinity (1988,1993,1996,1997,2002, 2008)
This data set includes vegetation classifications from the Willamette National Forest (years 1993, 1996, 1997, 2002, and 2008) and from 1988 thematic mapper satellite image.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.