Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,250
datasets available to search
ShareScore release 0.7.1
Dataset results
6,250 results for “Classification”
Data set for (binary) text classification, involving spoken utterances and written text
<p>This data set contains sentences belonging to either of two classes: Transcripts of spoken<br> (informal) text (Class 0), and written, formal text (Class 1). Sentences in Class 0 were<br> obtained from publicly available transcripts of radio shows (e.g. NPR),<br> whereas Sentences in Class 1 were obtained from Wikipedia.</p> <p>The data set is divided into three subsets: Training, validation, and test (specified by the file names).<br> Each set contains a large number of sentences, belonging to either of the two classes:</p> <p>In total, there are 13,640,458 sentences, of which 6,374,487 in Class 0 and 7,265,971.<br> The training set contains 9.743,188 sentences (of which 4,553,205 in Class 0 and 5,189,983 in Class 1), <br> the validation set contains 1,948,639 sentences (of which 910,641 in Class 0 and 1,037,998 in Class 1), and the <br> test set contains 1,948,631 sentences (of which 910,641 in Class0 and 1,037,990 in Class1). </p> <p>The data sets are in plain text format. Every row contains (i) the class label (0 or 1) and<br> (ii) the text of the sentence, separated from the class label by a tab character.</p> <p>Note that the sentences contain 5 tokens or more (including punctuation marks). </p>
LoCard Food Classification
<p>How the data was compiled:<br> The grocery purchase data including 3574 product groups was received from the retailer for research purposes. The data was reclassified into appropriate categories suitable for the use of nutrition and health research.</p> <p>How the data has been handled:<br> A four-level hierarchical classification of product groups was used. Each class on the broadest level of hierarchy (Class 1) was subsequently divided into a reasonable number of finer sub-classes starting with Class level 2, followed by Class 3, and finally, Class 4, which was the most detailed level of hierarchy. The main ingredient of the product group, type of the food and purpose of use, nutritional content, and carbon footprint were considered in the reclassification process. The classified food groups were linked with Finnish food composition database.</p> <p>How the data can be used for research:<br> The re-classified data can be used to measure can be used to infer and describe food purchase patterns, monitor of the nutrition composition of food purchases, monitoring and evaluating dietary environmental and economic sustainability, and deriving consumer segments based on the re-classification.</p> <p>The authors strongly recommend reading the detailed description of the classification is published in <a href="https://doi.org/10.21203/rs.3.rs-2826970/v1">https://doi.org/10.21203/rs.3.rs-2826970/v1</a>.</p>
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
Datasets and results from: "Random Forest Classification and Solar Flares Data: Analysis and Validation"
<p><strong>Instructions for the data and code repository</strong></p> <p>Results, post-processing workflow, and datasets for the research paper titled "Random Forest Classification and Solar Flares Data: Analysis and Validation".</p> <p>The folder contains three .csv files: the complete dataset (dataset.csv), the balanced training dataset (train_dataset.csv), and the testing dataset (test_dataset.csv).</p> <p>The folder also contains the result files from the research (.csv output files with predictions and .html files with evaluation metrics, etc.) exported from the JASP software. The number in each file name corresponds to the number of trees utilized in Random Forest modelling.</p> <p>In addition, the Python script for the post-processing workflow is provided, with comments located in the script.</p> <p>The soft range X-ray irradiance and VLF amplitude data were obtained from:<br> National Centers for Environmental Information (NCEI) Available online: https://www.ncei.noaa.gov/. Accessed on: 24th June 2023. <br> Worldwide archive of low-frequency data and observations (WALDO) Available online: https://waldo.world/. Accessed on: 24th June 2023.</p>
OEMC Hackathon 2023: EU Land Cover Classification Dataset
<p>Dataset organized by the <a href="https://earthmonitor.org/">Open-Earth-Monitor (OEMC) project</a> within the context of <a href="http://www.kaggle.com/competitions/oemc-hackathon-eu-land-cover-classification/overview">Hackathon 2023</a>.</p> <p>The dataset (both train and test) was produced by stratified sampling of the <strong>ground-truth</strong> data provided by LUCAS Survey, funded by the European Commission. The target land cover considered <strong>level-3</strong> classes from the harmonized legend, resulting in <strong>72 classes</strong> distributed over <strong>5 years </strong>(<code>2006</code>, <code>2009</code>, <code>2012</code>, <code>2015</code>, <code>2018</code>):</p> <p>All samples were overlaid with <strong>416</strong> raster spatial layers, including satellite (spectral bands and indices) and temperature images (land surface temperature), climate images (precipitation, air temperature), accessibility and distance maps (highways, water bodies, burned areas), digital terrain model (slope and elevation) and other existing maps (population count and snow covering). The result values were organized in columns, one for each spatial layers, which combined represent the feature space available for ML modeling.</p> <p><strong>Column names:</strong></p> <p>The columns are formed by six metadata fields separated by <code>_</code>:</p> <ul> <li>Example: <strong>red_landsat.glad.ard_p50_30m_jun25_sep12</strong></li> <li>Metadata fields: <ul> <li>F1 - Variable name: <strong>red</strong></li> <li>F2 - Variable procedure including product name: <strong>landsat.glad.ard</strong></li> <li>F3 - Position in the probability distribution: <strong>p50</strong></li> <li>F4 - Spatial resolution: <strong>30m</strong></li> <li>F5 - Start date: <strong>jun25</strong></li> <li>F6 - End date: <strong>sep12</strong></li> </ul> </li> </ul> <p><strong>Column description:</strong></p> <p>All the columns can be aggregated in six thematic groups according to F1 and F2:</p> <ul> <li><strong>Satellite images (spectral reflectance & vegetation indices):</strong> <ul> <li><code>blue_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat blue band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> <li><code>blue_mod13q1_{..}</code>: Monthly time-series of MOD13Q1 blue band (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>evi_mod13q1.stl.trend.ols.alpha_{..}</code>: Alpha coefficient / intercept (derived by <a href="https://www.statsmodels.org/devel/generated/statsmodels.regression.linear_model.OLS.html">OLS</a>) over the deseasonalized monthly time-series of MOD13Q1 Enhanced Vegetation Index (EVI) index (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>evi_mod13q1.stl.trend.ols.beta_{..}</code>: Beta coefficient / trend (derived by <a href="https://www.statsmodels.org/devel/generated/statsmodels.regression.linear_model.OLS.html">OLS</a>) over the deseasonalized monthly time-series of MOD13Q1 Enhanced Vegetation Index (EVI) index (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>evi_mod13q1.stl.trend_{..}</code>: Deseasonalized monthly time-series (trend component of <a href="https://www.statsmodels.org/dev/generated/statsmodels.tsa.seasonal.STL.html#statsmodels.tsa.seasonal.STL">STL</a>) for MOD13Q1 Enhanced Vegetation Index (EVI) index (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>evi_mod13q1_{..}</code>: Monthly time-series of MOD13Q1 Enhanced Vegetation Index (EVI) index (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>green_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat green band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> <li><code>mir_mod13q1_{..}</code>: Monthly time-series of MOD13Q1 mid-infrared band (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>ndvi_mod13q1_{..}</code>: Monthly time-series of MOD13Q1 normalized vegetation index (NDVI) (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>nir_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat near-infrared band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> <li><code>nir_mod13q1_{..}</code>: Monthly time-series of MOD13Q1 near-infrared band (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>red_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat red band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> <li><code>red_mod13q1_{..}</code>: Monthly time-series of MOD13Q1 red band (<a href="https://lpdaac.usgs.gov/products/mod13q1v006/">EarthData</a>)</li> <li><code>swir1_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat short-wave infrared-1 band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> <li><code>swir2_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat short-wave infrared-1 band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> </ul> </li> <li><strong>Temperature images:</strong> <ul> <li><code>lst_mod11a2.daytime_{..}</code>: Monthly time-series of MOD13Q1 day time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.daytime.{month}_{..}</code>: Long-term monthly aggregation (2000—2022) for MOD13Q1 day time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.daytime.trend_{..}</code>: Deseasonalized monthly time-series (trend component of <a href="https://www.statsmodels.org/dev/generated/statsmodels.tsa.seasonal.STL.html#statsmodels.tsa.seasonal.STL">STL</a>) for MOD13Q1 day time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.daytime.trend.ols.alpha_{..}</code>: Alpha coefficient / intercept (derived by <a href="https://www.statsmodels.org/devel/generated/statsmodels.regression.linear_model.OLS.html">OLS</a>) over the deseasonalized monthly time-series of MOD13Q1 day time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.daytime.trend.ols.beta_{..}</code>: Beta coefficient / trend (derived by <a href="https://www.statsmodels.org/devel/generated/statsmodels.regression.linear_model.OLS.html">OLS</a>) over the deseasonalized monthly time-series of MOD13Q1 day time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.nighttime_{..}</code>: Monthly time-series of MOD13Q1 night time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.nighttime.{month}_{..}</code>: Long-term monthly aggregation (2000—2022) for MOD13Q1 day time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.nighttime.trend_{..}</code>: Deseasonalized monthly time-series (trend component of <a href="https://www.statsmodels.org/dev/generated/statsmodels.tsa.seasonal.STL.html#statsmodels.tsa.seasonal.STL">STL</a>) for MOD13Q1 night time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.nighttime.trend.ols.alpha_{..}</code>: Alpha coefficient / intercept (derived by <a href="https://www.statsmodels.org/devel/generated/statsmodels.regression.linear_model.OLS.html">OLS</a>) over the deseasonalized monthly time-series of MOD13Q1 night time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>lst_mod11a2.nighttime.trend.ols.beta_{..}</code>: Beta coefficient / trend (derived by <a href="https://www.statsmodels.org/devel/generated/statsmodels.regression.linear_model.OLS.html">OLS</a>) over the deseasonalized monthly time-series of MOD13Q1 night time land surface temperature (<a href="https://lpdaac.usgs.gov/products/mod11a2v006/">EarthData</a>)</li> <li><code>thermal_landsat.glad.ard_{..}</code>: Quarterly time-series of Landsat thermal band (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> </ul> </li> <li><strong>Climate layers:</strong> <ul> <li><code>accum.precipitation_chelsa.annual_{..}</code>: Accumulated precipitation over the entire year according to CHELSA timeseries in <code>mm</code> of water (<a href="https://doi.org/10.1038/sdata.2017.122">Karger et al., 2017</a>)</li> <li><code>accum.precipitation_chelsa.annual.3years.dif_{..}</code>: 3-years difference considering the yearly accumulated precipitation according to CHELSA timeseries in <code>mm</code> of water (<a href="https://doi.org/10.1038/sdata.2017.122">Karger et al., 2017</a>)</li> <li><code>accum.precipitation_chelsa.annual.log.csum_{..}</code>: Cumulative sum, in logarithmic space, consdering the yearly accumulated precipitation according to CHELSA timeseries (<a href="https://doi.org/10.1038/sdata.2017.122">Karger et al., 2017</a>)</li> <li><code>accum.precipitation_chelsa.montlhy_{..}</code>: Accumulated precipitation for each month according to CHELSA timeseries in <code>mm</code> of water (<a href="https://doi.org/10.1038/sdata.2017.122">Karger et al., 2017</a>)</li> <li><code>bioclim.var_chelsa.{variable_code}_{..}</code>: Bioclimatic variables derived variables from the monthly mean, max, mean temperature, and mean precipitation values. For <code>variable_code</code> descriptions see <a href="https://chelsa-climate.org/bioclim/">chelsa-climate.org</a> (<a href="https://doi.org/10.1038/sdata.2017.122">Karger et al., 2017</a>)</li> </ul> </li> <li><strong>Accessibility & distance maps:</strong> <ul> <li><code>accessibility.to.ports_map.ox.{variable_code}_{..}</code>: Time-required to access ports of different size according to <a href="https://doi.org/10.1038/s41597-019-0265-5">Nelson et al., 2019</a></li> <li><code>burned.area.distance_global.fire.atlas_{..}</code>: Distance to burned areas mapped by <a href="https://doi.org/10.3334/ORNLDAAC/1642">Global Fire Atlas</a></li> <li><code>cost.distance.to.coast_gedi.grass.gis_{..}</code>: Cumulative cost of moving (derived by <a href="https://grass.osgeo.org/grass83/manuals/r.cost.html">r.cost</a>) to the coast</li> <li><code>road.distance_osm.highways.high.density_{..}</code>: Distance to high density of roads according to <a href="https://www.openstreetmap.org/#map=8/52.154/5.295">OpenStreetMap</a></li> <li><code>road.distance_osm.highways.low.density_{..}</code>: Distance to low density of roads according to <a href="https://www.openstreetmap.org/#map=8/52.154/5.295">OpenStreetMap</a></li> <li><code>water.distance_glad.interanual.dynamic.classes_{..}</code>: Distance to permanent / seasonal water bodies according to<br> <a href="https://doi.org/10.1016/j.rse.2020.111792">Pickens et al., 2020</a></li> </ul> </li> <li><strong>Digital terrain model (DTM):</strong> <ul> <li><code>elev.lowestmode_gedi.eml_{..}</code>: Mean estimate of the terrain elevation in <code>dm</code> filtered using <a href="https://saga-gis.sourceforge.io/saga_tool_doc/6.2.0/grid_filter_1.html">SAGA GIS Gaussian filter</a> (<a href="https://doi.org/10.7717/peerj.15478">Witjes et al., 2023</a>)</li> <li><code>slope.percent_gedi.eml_{..}</code>: Mean slope in <code>%</code> derived from terrain elevation ([Witjes et al., 2023]</li> </ul> </li> <li><strong>Other existing maps:</strong> <ul> <li><code>pop.count_ghs.jrc_{..}</code>: Annual time-series of population count in number of people mapped by <a href="https://data.jrc.ec.europa.eu/dataset/2ff68a52-5b5b-4a22-8f40-c41da8332cfe">Schiavina et al., 2023</a></li> <li><code>snow.duration_global.snowpack_{..}</code>: Annual duration of snow occurrence mapped by <a href="https://www.dlr.de/eoc/desktopdefault.aspx/tabid-8297/14218_read-37938/">Global SnowPack</a></li> </ul> </li> </ul> <p><strong>Files</strong></p> <ul> <li><strong>train.csv</strong>: Training set with 42,237 rows and 420 columns, including sample id (<code>sample_id</code> - index column), land cover code (<code>land_cover</code>), land cover label (<code>land_cover_label</code>), reference year (<code>year</code>) and 416 features / covariates</li> <li><strong>test.csv</strong>: Test set with 42,271 rows and 418 columns, including sample id (<code>sample_id</code> - index column), reference year (<code>year</code>) and 416 features / covariates</li> <li><strong>sample_submission.csv</strong>: a sample submission file with 42,271 rows and 2 columns, including sample id (<code>sample_id</code> - index column) and predicted land cover code (<code>land_cover</code>)</li> </ul>
Pre-training with simulated ultrasound images for breast mass segmentation and classification - dataset
<p>Dataset assosiated with the MICCAI Workshop on Data Engineering in Medical Imaging paper: "Pre-training with Simulated Ultrasound Images for Breast Mass Segmentation and Classification"</p>
Supervised land cover classification using Google Earth Engine in Córdoba, Argentina, 2018-2020
Land cover information is critical to scientific, economic, and public policy-making. There is a high demand for accurate and timely land cover information that affects the accuracy of all subsequent applications. The availability of Google Earth Engine (GEE), which derives temporal aggregation methods from time-series images (i.e., the use of metrics such as mean or median), has also enabled optimization of computation time, such as managing large amounts of data to obtain more accurate results. Our objective was to obtain a land cover map for the northwest of the province of Córdoba, Argentina. The study was carried out in rural communities that belong to the departments of Cruz del Eje and Ischilín, northwest of Córdoba, and have different degrees of intervention in the land cover. Sentinel 2 Level 2A images were acquired for the study area. Images available from January 1, 2018, to December 31, 2020, were sampled. To create a thematic map, the median value was calculated for the sample of images from the selected time interval. Finally, the Normalized Difference Vegetation Index (NDVI) was calculated and added to the total bands of the median image. Training polygons were placed there considering the visual features in the median image. The Random Forest algorithm was used as the classification method. To verify the quality of the classified map, a list of 97,753 verification pixels was obtained. In addition, a confusion matrix was created to collect the conflicts that arise between categories, and the precision and kappa coefficient was calculated to define the quality of the map obtained. Image acquisition, preprocessing, and analysis were performed on the Google Earth Engine platform. Thematic maps with eight classes were obtained, with a total area of 719880 ha. The confusion matrix showed an overall precision of 99.26% and a corrected kappa index of 0.99, the classes were correctly classified by the algorithm.
A Land-use/Land Cover Classification of Baltimore City in 1953
Land-use and land cover classifications are typically created using automated methods to analyze modern, spatially explicit color aerial imagery. However, creating classifications from black and white historical aerial imagery presents a number of challenges that require a combination of more traditional, manual techniques and approaches. A georectified mosaic of 113 aerial images was digitized in ArcGIS to create a land-use/land cover classification. The analyzed area covered 700 km2 (270 mi2) including all of Baltimore City, and a portion of Baltimore County immediately surrounding the city. A combination of 8 land-use and land cover classes were used: Agriculture, Barren, Built (Other), Forest, Grass/Shrubland, Industrial, Residential, and Water. This geospatial data set captures an ecologically and socially important moment in the post-war history of the city. It can be used to examine relationships between property ownership and forest patch dynamics across time. These insights may help inform future environmental planning, conservation, management, and stewardship goals for Baltimore City forest patches, and other cities throughout the region.
Maximum likelihood classification of 2006 AISA hyperspectral imagery of the GCE domain for vegetation
Airborne Imaging Spectrometer for Applications (AISA) Eagle hyperspectral imagery were acquired on June 20-21, 2006, by the Center for Advanced Land Management Information Technologies (CALMIT). This included four flight lines flown for the examination of vegetation for the Duplin River salt marshes. Imagery was acquired for 63 bands from 400-980 nm at a 1 m spatial resolution. Imagery were classified using the maximum likelihood classifier (MLC) and a post-classification decision tree to achieve an overall classification accuracy of 90%. Classification training and validation data were obtained from the 2006 Hyperspectral ground survey. See Hladik (2012) and Hladik, Alber, and Schalles (2013) and Schalles, et. al. (2013) for additional details.
Green Lakes Valley land cover classification, Niwot Ridge LTER, Colorado
Land cover data generated by Don Cline (graduate student, CU Boulder Geography), as part of suite of spatial maps made for Green Lakes Valley (see Williams et al. 1999).
An Open-set Recognition and Few-Shot Learning Dataset for Audio Event Classification in Domestic Environments
<p>The problem of training a deep neural network with a small set of positive samples is known as few-shot learning (FSL). It is widely known that traditional deep learning (DL) algorithms usually show very good performance when trained with large datasets. However, in many applications, it is not possible to obtain such a high number of samples. In the image domain, typical FSL applications are those related to face recognition. In the audio domain, music fraud or speaker recognition can be clearly benefited from FSL methods. This paper deals with the application of FSL to the detection of specific and intentional acoustic events given by different types of sound alarms, such as door bells or fire alarms, using a limited number of samples. These sounds typically occur in domestic environments where many events corresponding to a wide variety of sound classes take place. Therefore, the detection of such alarms in a practical scenario can be considered an open-set recognition (OSR) problem. To address the lack of a dedicated public dataset for audio FSL, researchers usually make modifications on other available datasets. This paper is aimed at providing the audio recognition community with a carefully annotated dataset for FSL and OSR comprised of 1360 clips from 34 classes divided into pattern sounds and unwanted sounds. To facilitate and promote research in this area, results with two baseline systems (one trained from scratch and another based on transfer learning), are presented.</p> <p> </p>
Video classification using deep learning
<p>Material accompanying the paper "Frame-by-frame annotation of video recordings using deep neural networks". Contains a selection of videos used in the paper, manual annotations, and results. Code is contained in the associated GitHub repository. See the readme for details.</p>
Fast MLE and Supervised Classification for the Beta-Liouville Multinomial -- Gold Standard Data
<p>Gold standard datasets used in the publication Fast Maximum Likelihood Estimation and Supervised Classification for the Beta-Liouville Multinomial. Datasets were prepared by Cardoso-Cachopo (2007).</p>
GLC_FCS30: Global land-cover product with fine classification system at 30 m using time-series Landsat imagery
<p>A novel global 30-m land-cover product with a fine classification system for the year 2015 (GLC_FCS30-2015). The product was produced by combining time-series of Landsat imagery and high-quality training data from the GSPECLib (Global Spatial Temporal Spectra Library) on the Google Earth Engine computing platform. First, the global training data from the GSPECLib were developed by applying a series of rigorous filters to the MCD43A4 NBAR and CCI_LC land-cover products. Secondly, a local adaptive random forest model was built for each 5°×5° geographical tile by using the multi-temporal Landsat spectral and textures features of the corresponding training data, and the GLC_FCS30-2015 land-cover product containing 30 land-cover types was generated for each tile.</p>
Dataset-AOB: urban sounds events classification
<p>The dataset Dataset-AOB is an audio dataset collected and manually edited for urban sounds events classification using Convolutional Neural Networks for the Master Thesis: </p> <p>Ospina, A. "Audio Event Classification using Deep Learning. Use case: Urban Sounds Events classification with Convolutional Neural Networks," M.Eng. thesis, Beuth University of Applied Sciences, Berlin, 2020.</p> <p>- 10 audio events: alarm-siren, children playing, dog bark, engine, footsteps, glass breaking, gun shot, metro train, rain and screams.</p> <p>- duration: < 4 seconds</p> <p>- format: (.wav)</p> <p>- sampling rate: 22KHz - 44KHz</p> <p>- files: Dataset-AOB: development dataset (4831 samples), DatasetEVAL-AOB: evaluation (218 samples)</p> <p>- metadata: (.csv)</p> <p>- sources per class: (.png)</p> <p>Contact: aospinab@gmail.com</p>
Data for "Systematic Mapping of Open Data Studies: Classification and Trends from a Technological Perspective"
<p>Data used to perform a systematic mapping to classify and analyse existing research on open data from a technological viewpoint from 2006 to 2019. This dataset contains information from six key facets from the collected publications coming from several scientific repositories/databases: publication venue, impact, subject, domain, life-cycle and research type.</p>
Data release for "OrchID: a Generalized Framework for Taxonomic Classification of Images Using Evolved Artificial Neural Networks"
<p><strong>Abstract</strong></p> <p>Taxonomic expertise for the identification of species is rare and costly. On-going advances in computer vision and machine learning have led to the development of numerous semi- and fully automated species identification systems. However, these systems are rarely agnostic to specific morphology, rarely can perform taxonomic “approximation” (by which we mean partial identification at least to higher taxonomic level if not to species), and frequently rely on costly scientific imaging technologies.</p> <p>We present a generic, hierarchical identification system for automated taxonomic approximation of organisms from images. We assessed the effectiveness of this system using photographs of slipper orchids (Cypripedioideae), for which we implemented image pre-processing, segmentation, and colour and shape feature extraction algorithms to obtain digital phenotypes for 116 species. The identification system trained on these digital phenotypes uses a nested hierarchy of artificial neural networks for pattern recognition and automated classification that mirrors the Linnean taxonomy, such that user-submitted photos can be assigned a genus, section, and species classification by traversing this hierarchy.</p> <p>Performance of the identification system varied depending on photo quality, number of species included for training, and desired taxonomic level for identification. High quality photos were scarce for some taxa and were under-represented in the training set, resulting in imbalanced network training. The image features used for training were sufficient to reliably identify photos to the correct genus but less so to the correct section and species.</p> <p>The outcomes of this project include a library of feature extraction algorithms called <em>ImgPheno</em>, a collection of scripts for neural network training called <em>NBClassify</em>, a library for evolutionary optimization of artificial neural network construction called <em>AI::FANN::Evolving</em> and a planned web application called <em>OrchID</em> for identification of user-submitted images. All project outcomes are open source and freely available.</p> <p><strong>About this release</strong></p> <p>This release corresponds belongs with our response to the reviewers of PLoS One. At this stage of the review cycle the manuscript is assessed as 'minor revision'. Consequently, we don't anticipate making more releases until publication.</p>
1988-2009 time-series of land-use/land-cover maps for the Mar Menor / Campo de Cartagena watershed by means of supervised classification of Landsat images.
<p>Serie de mapas de usos y coberturas de la cuenca del Mar Menor (SE España): 2009, 2000, 1997 y 1998. Así como el documento completo de tesis en las que se generaron y analizaron.</p> <p>Time-series of land-use / land-cover maps of Mar Menor watershed (SE Spain): 2009, 2000, 1997 y 1998. As well as the complete thesis document in which they were generated and analyzed.</p>
MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications
<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>
Automated MESSENGER Plasma Region Classifications via Unsupervised Transfer Learning
<p>This file contains the 1-minute resolution dataset (“labeled_sunside_data_3labels.csv”) for Toy-Edens et al.’s Automated Classification of MESSENGER Plasma Observations via Unsupervised Transfer Learning. The 1-minute resolution file contains the rolled up 1-minute epoch, features that go into clustering and post-cleaning methods, spacecraft positions (in MSO), total magnetic field, raw and cleaned clustering labels, and raw and cleaned transition name.</p> <p>We ask that if you use any parts of the dataset that you cite Toy-Edens et al.’s Automated Classification of MESSENGER Plasma Observations via Unsupervised Transfer Learning (DOI: 10.3389/fspas.2025.1608091).</p> <p>This work was supported by NASA grants 80NSSC19K0789 and 80NSSC22K0993.</p> <p> </p> <p>The following tables detail the contents of the described files:</p> <p><strong>labeled_sunside_data_3labels.csv description</strong></p> <table style="width: 100.063%; height: 851.2px;"> <tbody> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p><strong>Column Name</strong></p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p><strong>Description</strong></p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> Epoch</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Epoch in datetime (YYYY-MM-DD HH:MM:SS)</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> x_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>x position of the spacecraft in MSO [km]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> y_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>y position of the spacecraft in MSO [km]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> z_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>z position of the spacecraft in MSO [km]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> btot_mso</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Total magnetic field [nT]</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> norm_Btot</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Magnitude of the total magnetic field normalized to 150nT. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> ratio_max_width</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Ratio of the width of the most prominent ion spectra peak (in number of energy channels) to max number of energy channels. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> ratio_high_low</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Ratio of the mean of the log intensity of high energies in the ion spectra to the mean of the log intensity of low energies in the ion spectra. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> high_intensity</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Boolean if there is a peak with a higher minimum intensity threshold. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> spectra_counts</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>A ratio of spectra bins with non-zero counts to all possible spectra bins (i.e. way to determine if too much missing spectra data). See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> raw_named_label</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Raw cluster assigned plasma region label (allowed values: magnetosheath, magnetosphere, solar wind)</p> </td> </tr> <tr> <td style="width: 17.3792%;"> <p>intermediate_named_label</p> </td> <td style="width: 78.9512%;"> <p>Cleaned cluster assigned plasma region label with only relabeling rules applied. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> named_label</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Cleaned cluster assigned plasma region label with relabeling rules and post-processing applied (use these unless have a specific reason to use raw labels). See paper for more information</p> </td> </tr> <tr style="height: 47.6px;"> <td style="width: 17.3792%; height: 47.6px;"> <p> raw_transition_name</p> </td> <td style="width: 78.9512%; height: 47.6px;"> <p>Raw transition names (e.g. bow shock, magnetopause) based on "raw_named_label" cluster labels. See paper for more information</p> </td> </tr> <tr style="height: 67.2px;"> <td style="width: 17.3792%; height: 67.2px;"> <p> transition_name</p> </td> <td style="width: 78.9512%; height: 67.2px;"> <p>Cleaned transition names (e.g. bow shock, magnetopause) after removing likely transient transitions based on "named_label" cluster labels. See paper for more information</p> </td> </tr> </tbody> </table> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.