Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,185

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,185 results for “Learning”

Learn how ShareScore rates datasets ↗
edi60/100

GRIME AI Water Segmentation Model for the USGS Monitoring Site at Kearney Outdoor Learning Area, NE, 2024-2024

Ground-based observations from fixed-mount cameras have the potential to fill an important role in environmental sensing, including direct measurement of water levels and qualitative observation of ecohydrological research sites. All of this is theoretically possible for anyone who can install a trail camera. Easy acquisition of ground-based imagery has resulted in millions of environmental images stored, some of which are public data, and many of which contain information that has yet to be used for scientific purposes. The goal of this project was to develop and document key image processing and machine learning workflows, primarily related to semi-automated image labeling, to increase the use and value of existing and emerging archives of imagery that is relevant to ecohydrological processes. This data package includes imagery, annotation files, water segmentation model and model performance plots, and model test results (overlay images and masks) for the USGS Monitoring Site at Kearney Outdoor Learning Area, NE, 2024-2024. All imagery was acquired from the USGS Hydrologic Imagery Visualization and Information System (HIVIS; see https://apps.usgs.gov/hivis/camera/NE_Kearney_Outdoor_Learning_Area for this specific data set) and/or the National Imagery Management System (NIMS) API. Water segmentation models were created by tuning the open-source Segment Anything Model 2 (SAM2, https://github.com/facebookresearch/sam2) using images that were annotated by team members on this project. The models were trained on the "water" annotations, but annotation files may include additional labels, such as "snow", "sky", and "unknown". Image annotation was done in Computer Vision Annotation Tool (CVAT) and exported in COCO format (.json). All model training and testing was completed in GaugeCam Remote Image Manager Educational Artificial Intelligence (GRIME AI, https://gaugecam.org/) software (Version: Beta 16). Model performance plots were automatically generated during this pr

openCC (other)Sep 2025View details →
zenodo56/100

Improving Artificial Teachers by Considering How People Learn and Forget: Dataset

<p>This dataset contains the results of the experiment described in&nbsp;<a href="https://dl.acm.org/doi/10.1145/3397481.3450696">Nioche et al. (2021)</a>.&nbsp;</p> <p>This&nbsp;dataset contains 4&nbsp;data files:</p> <ul> <li><em>data.csv</em>: the main data file.</li> <li><em>stimuli.csv:</em> the description/listing of the stimuli.</li> <li><em>demographic_info.csv</em>: the demographic information about the users.</li> <li><em>data_incl_preliminary_exp.csv</em>: an additional data file that includes the user of the preliminary experiments</li> </ul> <p>The main data file contains the logs of&nbsp;53 different users using a self-teaching application for one week. The goal of the users&nbsp;was to learn the English meaning of Japanese kanji. Each user completed between 1370 trials and 1608 trials. Each user saw between 85 and 204 characters.&nbsp;</p> <p>Two additional files are also joint to the data files:</p> <ul> <li><em>info.ipynb</em>: A Jupyter notebook that provides&nbsp;information about each data file, a few descriptive plots,&nbsp;and an example of data manipulation.</li> <li><em>info.pdf: </em>A pdf rendering of the notebook.</li> </ul> <p>If you use this dataset, please refer to it by citing&nbsp;<a href="https://dl.acm.org/doi/10.1145/3397481.3450696">Nioche et al. (2021)</a>.</p>

opencc-by-4.0Apr 2021View details →
zenodo56/100

Sample data for "Machine learning for large-scale forecasting"

<p>This dataset includes sample data for the Netherlands to run the machine learning baseline as described in the paper titled <em>Machine learning for large-scale crop yield forecasting</em>, accessible at&nbsp;<a href="https://doi.org/10.1016/j.agsy.2020.103016">https://doi.org/10.1016/j.agsy.2020.103016</a>.&nbsp;The software implementation of the machine learning baseline is available at:&nbsp;<a href="https://github.com/BigDataWUR/MLforCropYieldForecasting">https://github.com/BigDataWUR/MLforCropYieldForecasting</a>.</p> <p><strong>Notes:</strong></p> <p>The NUTS classification (Nomenclature of territorial units for statistics) is a hierarchical system for dividing up the economic territory of the EU and the UK (see Eurostat, 2016) for more details).</p> <p>Data</p> <p>The dataset consists of 11 CSV files. They are formatted to work as sample inputs to the machine learning baseline.</p> <ol> <li><strong>Crop Area Fractions </strong>(NUTS2, NUTS1):&nbsp;We aggregated the predictions of the machine learning baseline from NUTS2&nbsp;to national (NUTS0) level&nbsp;by weighting them on the modeled crop area. Cerrani and L&oacute;pez Lozano (2017) have described in detail the algorithm used to model crop areas for different NUTS levels. The data comes from the MARS Crop Yield Forecasting System (MCYFS) of European Commission&#39;s Joint Research Centre (JRC) (see Lecerf et al., 2019).</li> <li><strong>Centroids (NUTS2)</strong>: Data includes latitude, longitude and distance to coast of the centroids of NUTS2 regions.</li> <li><strong>Meteo Daily Data and Meteo Dekadal Data </strong>(NUTS2):&nbsp;The data comes from MCYFS&nbsp;(see EC-JRC, 2020). By default, the implementation uses daily data.</li> <li><strong>Remote Sensing Data</strong> (NUTS2, see Copernicus Global Land Service, 2020): Data includes fraction of absorbed photosynthetically active radiation (FAPAR) aggregated to NUTS2.</li> <li><strong>Soil Data</strong>: Data includes soil moisture information that can be used to calculate soil water holding capacity. The data comes from MCYFS (see Lecerf et al., 2019).</li> <li><strong>WOFOST data </strong>(NUTS2): The World Food Studies (WOFOST) crop model (van Diepen et al., 1989; Supit et al., 1994; de Wit et al.&nbsp; 2019) is a simulation model for the quantitative analysis of the growth and production of annual field crops. It is a mechanistic, dynamic model that explains daily crop growth on the basis of the underlying processes, such as photosynthesis, respiration and how these processes are influenced by environmental conditions.&nbsp;The crop simulation is fed by weather, soil and crop data. Observed meteorological data is interpolated on a regular 25 km grid using a method based on the distance, altitude and climatic region similarity between the center of grid cells and weather stations (see Van der Goot, 1998). WOFOST runs on the intersection between the 25 km meteorological grid and soil units based on the European soil map (http://esdac.jrc.ec.europa.eu/). In order to have the output data aggregated to administrative regions such as countries or provinces, simulation units are further intersected with the boundaries of these regions. The outputs at soil unit (STU) level are aggregated to grid level in an area weighted manner. Gridded simulations are aggregated to lowest NUTS level 3 considering the arable land area of each grid, derived from GLOBCOVER and CORINE Land Cover (Cerrani and Lopez Lozano, 2017). From NUTS3 to higher levels, crop area fractions for the current year, retrieved from Eurostat, are used to weight and aggregate the output (Cerrani and Lopez Lozano, 2017).</li> <li><strong>GAES data</strong>: GAES data includes agro-climatic features of regions, such as&nbsp;elevation and slope (from USGS-EROS, 2021), field size (from&nbsp;Lesiv et al., 2019), irrigated (crop) areas (from&nbsp;EC-JRC, 2020) and crop areas&nbsp;(from&nbsp;EC-JRC, 2020).</li> <li><strong>National yield statistics&nbsp;</strong>(NUTS0): These are the official Eurostat national yield statistics (Eurostat, 2020a).&nbsp;We used these yield statistics as reference to compare&nbsp;the machine learning predictions aggregated to NUTS0 and the actual MCYFS forecasts (see van der Velde and Nisini, 2019).</li> <li><strong>Regional yield statistics&nbsp;</strong>(NUTS2): We used NUTS2 yield statistics&nbsp;as labels to train and evaluate machine learning algorithms. We got NUTS2&nbsp;yield statistics from&nbsp;The Central Bureau of Statistics (CBS) of the Netherlands&nbsp;(NL-CBS, 2020).</li> <li><strong>Past MCYFS Yield Forecasts&nbsp;</strong>(NUTS0): These are actual forecasts made by MCYFS in the past (see van der Velde and Nisini, 2019). We used the official Eurostat national yield statistics (see point 7 above) as the reference to compare the machine learning predictions aggregated to NUTS0 and&nbsp;MCYFS forecasts.</li> </ol> <p><strong>Crop ID and name mapping</strong></p> <p>2 : grain maize</p> <p>6 : sugar beets</p> <p>7 : potatoes</p> <p>90 : soft wheat</p> <p>93 : sunflower</p> <p>95 : spring barley</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>We would like to thank S. Niemeyer from the European Commission&rsquo;s Joint Research Centre (JRC) for the permission to provide open&nbsp;access to the Netherlands data. Similarly, we would like to thank M. van der Velde, L. Nisini and I. Cerrani from JRC for sharing with us past MCYFS forecasts&nbsp;and Eurostat national yield statistics.</p>

opencc-by-4.0Dec 2020View details →
OpenNeuro52/100

Route Learning

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
OpenNeuro52/100

OLVSL_ Object-location visual statistical learning

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
OpenNeuro52/100

Learning, Inhibitory Control, and Perception

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo52/100

Detecting repeating earthquakes on the San Andreas Fault with unsupervised machine-learning of spectrograms (supplementary material)

<p>Supplementary material for Sawi et al., 2023, <i>Detecting repeating earthquakes on the San Andreas Fault with unsupervised machine-learning of spectrograms </i>(The Seismic Record). Catalog of repeating earthquakes in sequences on a 10-km long segment of the San Andreas Fault in California from 1984-2019.&nbsp;</p><p>&nbsp;</p><p><strong>Catalog Header</strong></p><p>YR/MO/DY...........Date of event</p><p>HR/MN/SC...........Time of event</p><p>LAT/LON/DEP........Location of event</p><p>EX/EY/EZ...........Relative location uncertainty (in m)</p><p>MAG................NCSN magnitude</p><p>evID.................NCSN event ID</p><p>seqID................Repeating earthquake sequence ID</p><p>isRESp............Is quasi-periodic RES (bool)</p><p>&nbsp;</p><p><strong>References:&nbsp;</strong></p><p>Sawi T., Waldhauser F., Holtzman B. K., Groebner, N. (2023) Detecting repeating earthquakes on the San Andreas Fault with unsupervised machine-learning of spectrograms. The Seismic Record.&nbsp;</p><p>Waldhauser, F., and Schaff, D. P. (2021). A Comprehensive Search for Repeating Earthquakes in Northern California: Implications for Fault Creep, Slip Rates, Slip Partitioning, and Transient Stress. J Geophys Res B Solid Earth, 126(11), 1–22.&nbsp;<a href="https://doi.org/10.1029/2021JB022495">https://doi.org/10.1029/2021JB022495</a></p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Gravity Spy Machine Learning Classifications of LIGO Glitches from Observing Runs O1, O2, O3a, and O3b

<p>This data set contains all classifications that the Gravity Spy Machine Learning model for LIGO glitches from the first three observing runs (<a href="https://doi.org/10.7935/K57P8W9D">O1</a>, <a href="https://doi.org/10.7935/CA75-FM95">O2</a> and O3, where O3 is split into <a href="https://doi.org/10.7935/nfnt-hm34">O3a</a> and <a href="https://doi.org/10.7935/pr1e-j706">O3b</a>). Gravity Spy classified all noise events identified by the <a href="https://doi.org/10.1016/j.softx.2020.100620">Omicron trigger pipeline</a> in which Omicron identified that the signal-to-noise ratio was above 7.5 and the peak frequency of the noise event was between 10 Hz and 2048 Hz. To classify noise events, Gravity Spy made <a href="https://en.wikipedia.org/wiki/Constant-Q_transform">Omega scans</a> of every glitch consisting of 4 different durations, which helps capture the morphology of noise events that are both short and long in duration.</p> <p>There are <a href="https://doi.org/10.1088/1361-6382/aa5cea">22 classes</a> used for O1 and O2 data (including No_Glitch and None_of_the_Above), while there are <a href="https://doi.org/10.1088/1361-6382/ac1ccb">two additional classes</a> used to classify O3 data (while None_of_the_Above was removed).</p> <p>For O1 and O2, the glitch classes were: 1080Lines, 1400Ripples, Air_Compressor, Blip, Chirp, Extremely_Loud, Helix, Koi_Fish, Light_Modulation, Low_Frequency_Burst, Low_Frequency_Lines, No_Glitch, None_of_the_Above, Paired_Doves, Power_Line, Repeating_Blips, Scattered_Light, Scratchy, Tomte, Violin_Mode, Wandering_Line, Whistle</p> <p>For O3, the glitch classes were: 1080Lines, 1400Ripples, Air_Compressor, Blip, <strong>Blip_Low_Frequency</strong>, Chirp, Extremely_Loud, <strong>Fast_Scattering</strong>, Helix, Koi_Fish, Light_Modulation, Low_Frequency_Burst, Low_Frequency_Lines, No_Glitch, None_of_the_Above, Paired_Doves, Power_Line, Repeating_Blips, Scattered_Light, Scratchy, Tomte, Violin_Mode, Wandering_Line, Whistle</p> <p>The data set is described in <a href="https://doi.org/10.1088/1361-6382/acb633"><strong>Glanzer </strong><em>et al</em><strong>. (2023)</strong></a>, which we ask to be cited in any publications using this data release. Example code using the data can be found in this <a href="https://colab.research.google.com/drive/19q_lItODPk7qw_sohlHyWPnAbY0FZyt8?usp=sharing"><strong>Colab notebook</strong></a>.</p> <p>If you would like to download the Omega scans associated with each glitch, then you can use the gravitational-wave data-analysis tool <a href="https://gwpy.github.io/docs/stable/">GWpy</a>. If you would like to use this tool, please install anaconda if you have not already and create a virtual environment using the following command</p> <pre><code class="language-bash">conda create --name gravityspy-py38 -c conda-forge python=3.8 gwpy pandas psycopg2 sqlalchemy</code></pre> <p>After downloading one of the CSV files for a specific era and interferometer, please run the following Python script if you would like to download the data associated with the metadata in the CSV file. We recommend not trying to download too many images at one time. For example, the script below will read data on Hanford glitches from O2 that were classified by Gravity Spy and filter for only glitches that were labelled as Blips with 90% confidence or higher, and then download the first 4 rows of the filtered table.</p> <pre><code class="language-python">from gwpy.table import GravitySpyTable H1_O2 = GravitySpyTable.read('H1_O2.csv') H1_O2[(H1_O2["ml_label"] == "Blip") &amp; (H1_O2["ml_confidence"] &gt; 0.9)] H1_O2[0:4].download(nproc=1)</code></pre> <p>Each of the columns in the CSV files are taken from various different inputs:&nbsp;</p> <p>[&lsquo;event_time&rsquo;, &lsquo;ifo&rsquo;, &lsquo;peak_time&rsquo;, &lsquo;peak_time_ns&rsquo;, &lsquo;start_time&rsquo;, &lsquo;start_time_ns&rsquo;, &lsquo;duration&rsquo;, &lsquo;peak_frequency&rsquo;, &lsquo;central_freq&rsquo;, &lsquo;bandwidth&rsquo;, &lsquo;channel&rsquo;, &lsquo;amplitude&rsquo;, &lsquo;snr&rsquo;, &lsquo;q_value&rsquo;] contain metadata about the signal from the <a href="https://virgo.docs.ligo.org/virgoapp/Omicron/">Omicron pipeline</a>.&nbsp;</p> <p>[&lsquo;gravityspy_id&rsquo;] is the unique identifier for each glitch in the dataset.&nbsp;</p> <p>[&lsquo;1400Ripples&rsquo;, &lsquo;1080Lines&rsquo;, &lsquo;Air_Compressor&rsquo;, &lsquo;Blip&rsquo;, &lsquo;Chirp&rsquo;, &lsquo;Extremely_Loud&rsquo;, &lsquo;Helix&rsquo;, &lsquo;Koi_Fish&rsquo;, &lsquo;Light_Modulation&rsquo;, &lsquo;Low_Frequency_Burst&rsquo;, &lsquo;Low_Frequency_Lines&rsquo;, &lsquo;No_Glitch&rsquo;, &lsquo;None_of_the_Above&rsquo;, &lsquo;Paired_Doves&rsquo;, &lsquo;Power_Line&rsquo;, &lsquo;Repeating_Blips&rsquo;, &lsquo;Scattered_Light&rsquo;, &lsquo;Scratchy&rsquo;, &lsquo;Tomte&rsquo;, &lsquo;Violin_Mode&rsquo;, &lsquo;Wandering_Line&rsquo;, &lsquo;Whistle&rsquo;] contain the machine learning confidence for a glitch being in a particular Gravity Spy class (the confidence in all these columns should sum to unity). These use the original 22 classes in all cases.</p> <p>[&lsquo;ml_label&rsquo;, &lsquo;ml_confidence&rsquo;] provide the machine-learning predicted label for each glitch, and the machine learning confidence in its classification.&nbsp;</p> <p>[&lsquo;url1&rsquo;, &lsquo;url2&rsquo;, &lsquo;url3&rsquo;, &lsquo;url4&rsquo;] are the links to the publicly-available <a href="https://gwdetchar.readthedocs.io/en/stable/omega/">Omega scans</a> for each glitch. &lsquo;url1&rsquo; shows the glitch for a duration of 0.5 seconds, &lsquo;url2&rsquo; for 1 seconds, &lsquo;url3&rsquo; for 2 seconds, and &lsquo;url4&rsquo; for 4 seconds.</p> <p>For the most recently uploaded training set used in Gravity Spy machine learning algorithms, please see <a href="https://zenodo.org/record/1486046#.YZfcar3MJqs">Gravity Spy Training Set</a> on Zenodo.&nbsp;</p> <p><br> For detailed information on the training set used for the original Gravity Spy machine learning paper, please see <a href="https://zenodo.org/record/1476156#.YZfchL3MJqs">Machine learning for Gravity Spy: Glitch classification and dataset</a> on Zenodo.</p>

opencc-by-4.0Nov 2021View details →
zenodo52/100

Experimental data for "Deep Learning Methods for Colloidal Silver Nanoparticle Concentration and Size Distribution Determination from UV-Vis Extinction Spectra"

<p>Testing data (experimental data) for neural networks published in preprint https://doi.org/10.48550/arXiv.2404.10891</p> <p>The UV-VIS-NIR spectral data was also used in the dissertation of Nadzeya Khinevch, titled "Two-dimensional structures of nanoparticles for elements of surface-enhanced Raman scattering substrates".</p> <p>Emails of the corresponding authors:</p> <p>Tomas Klinavičius tomas.klinavicius@ktu.lt</p> <p>Tomas Tamulevičius tomas.tamulevicius@ktu.lt</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Unravelling the physiological and psychosocial signatures of pain by machine learning

<p>These datasets include information from 118 subjects, with 81 chronic pain patients across three different cohorts, Complex Regional Pain Syndrome (CRPS), Low Back Pain (LBP) and Spinal Cord Injury with Neuropathic Pain (SCI NP) and 37 healthy subjects. Each participant underwent 40 repetitions of experimentally induced pain, resulting in a total of 4,697 pain trials. Physiological signals (EDA and EEG) and psychosocial information have been recorded and collected. Age, gender, height, weight, BMI, medications, fatigue, sleep quality, perceived health, quality of life, sleep quality, and sick leave and validated questionnaires: Hospital Anxiety and Depression Scale (HADS), Pain Catastrophizing Score (PCS), and Pain Self-Efficacy Questionnaire (PSEQ) and Multidimensional Assessment of Interoceptive Awareness (MAIA).</p> <p>This information has been used for the publication "Unravelling the physiological and psychosocial signatures of pain by machine learning".</p> <p><span><em><strong>Cite this dataset as:</strong></em></span><br>N. Gozzi, G. Preatoni, F. Ciotti, M. Hubli, P. Schweinhardt, A. Curt, S. Raspopovic, Unraveling the physiological and psychosocial signatures of pain by machine learning. Med 0 (2024). &nbsp;<a href="https://doi.org/10.1016/j.medj.2024.07.016" target="_blank" rel="noopener">10.1016/j.medj.2024.07.016</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Exploring AdaBoost and Random Forests machine learning approaches for infrared pathology on unbalanced data sets

<p>The use of infrared spectroscopy to augment decision-making in histopathology is a promising direction for the diagnosis of many disease types. Hyperspectral images of healthy and diseased tissue, generated by infrared spectroscopy, are used to build chemometric models that can provide objective metrics of disease state. It is important to build robust and stable models to provide confidence to the end user. The data used to develop such models can have a variety of characteristics which can pose problems to many model-building approaches. Here we have compared the performance of two machine learning algorithms &ndash; AdaBoost and Random Forests &ndash; on a variety of non-uniform data sets. Using samples of breast cancer tissue, we devised a range of training data capable of describing the problem space. Models were constructed from these training sets and their characteristics compared. In terms of separating infrared spectra of cancerous epithelium tissue from normal-associated tissue on the tissue microarray, both AdaBoost and Random Forests algorithms were shown to give excellent classification performance (over 95% accuracy) in this study. AdaBoost models were more robust when datasets with large imbalance were provided. The outcomes of this work are a measure of classification accuracy as a function of training data available, and a clear recommendation for choice of machine learning approach.</p>

opencc-by-4.0May 2021View details →
zenodo52/100

SQLite database to accompany the paper, "Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring"

<p>This dataset is a SQLite database that accompanies methods and analysis described in the paper, &quot;Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring&quot; (Balantic &amp; Donovan 2019, Bioacoustics, https://www.tandfonline.com/doi/full/10.1080/09524622.2019.1605309).&nbsp;</p> <p>A Github repository containing code for using the SQLite&nbsp;database also accompanies this paper at:&nbsp;<a href="https://github.com/cbalantic/false-positive-mitigation">http://github.com/cbalantic/false-positive-mitigation</a></p>

opencc-by-4.0May 2019View details →
zenodo52/100

Dataset for a machine learning tool to improve lymph node staging with FDG-PET/CT

<p>This upload provides Open Data associated with the publication&nbsp;&quot;A machine learning tool to improve prediction of mediastinal lymph node metastases in non-small cell lung cancer using routinely obtainable [<sup>18</sup>F]FDG-PET/CT parameters&quot; by Rogasch JMM <em>et al.</em> (2022).</p> <p>The upload contains the&nbsp;anonymized dataset&nbsp;with 10 features necessary for the final GBM model that was presented in the publication. However, the original full dataset with&nbsp;40 features was excluded from this Open Data repository because it may not comply with strict rules of data anonymization. The full dataset can be obtained from the corresponding author (julian.rogasch@charite.de) upon reasonable request.</p> <p>Besides the dataset, this upload provides the original python and R scripts that were used as well as&nbsp;their output.</p> <p>A description of all&nbsp;files&nbsp;can be found in &quot;content_description_2022_11_19.txt&quot;.</p> <p>A user-friendly web tool that implements the final machine learning model can be found here:&nbsp;<a href="https://baumgagl.github.io/PET_LN_calculator/">PET_LN_calculator</a>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo52/100

Datasets for evaluating scalable supervised learning for synthesize-on-demand chemical libraries

<p>This repository contains datasets for the manuscript &quot;Evaluating scalable supervised learning for synthesize-on-demand chemical libraries&quot;:</p> <ul> <li><strong>ams_all_preds.csv.gz</strong>: The AMS dataset predictions when using an RF or baseline model trained on the training dataset. Includes the predicted score and rank from each model for each compound. We started with 8,434,707 AMS compounds and detected that 247,025 were in the LC or MLPCN training data. These were removed from the AMS list, leaving 8,187,682 compounds to score. The compound matching was done on the SMILES that we canonicalized in rdkit.</li> <li><strong>ams_order_results.csv.gz</strong>: Information about the 1,024 compounds purchased from the AMS library. Excludes the 4 AMS compounds that were incompletely dissolved. Includes the chemical feature representation, information from the vendor, RF and baseline model predictions, screening results, and clustering results.</li> <li><strong>baseline_weight.npy</strong>: The saved Similarity Baseline model, which consists of the active compounds in the training data. This model was used to score the AMS library. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a>&nbsp;for code to load the model and make predictions on new compounds.</li> <li><strong>cdd_training_data.tar.gz</strong>: The LC1234 and MLPCN PriA-SSB screening data exported from CDD.</li> <li><strong>enamine_costs_clustered_v3_with_nneighbor.csv.gz</strong>: Contains 5,620 Enamine compounds that were selected based on the RF prediction score and availability. This file also contains the Taylor-Butina cluster ID when clustering the training compounds, 1,024 tested AMS compounds, and top-ranked Enamine compounds at a 0.4 threshold. The nearest neighbor compounds in the training and AMS sets are also included along with compound information from Enamine, RF model scores, and chemical feature representations.</li> <li><strong>enamine_dose_response_curve_plots.xlsx</strong>: Images of the dose response curves from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, multiple curves are shown in the same plot. The compound structure images and SMILES are exported from CDD, not generated with RDKit.</li> <li><strong>enamine_dose_response_curves.tsv</strong>: The dose response curve summaries from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, only the highest-quality dose response curve was used.</li> <li><strong>enamine_final_list.csv.gz</strong>: The final 100 filtered compounds from&nbsp;<code>enamine_top_10000.csv.gz</code>. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>enamine_PriA-SSB_dose_response_data.tar.gz</strong>: The dose response screening data from all three runs on the 68 Enamine compounds. The 2021-06-16 run was originally screened on 2020-08-24. 2021-06-16 is the date the compound identities were corrected. This run contains two 1,536 well plates.</li> <li><strong>enamine_top_10000.csv.gz</strong>: Top 10,000 predictions from the Enamine REAL dataset using the selected RF model. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>master_df.csv.gz</strong>: The output of preprocessing the files in&nbsp;<code>cdd_training_data.tar.gz</code>. Contains 441,900 rows.</li> <li><strong>random_forest_classification_139.pkl</strong>: The saved RF classification model with&nbsp;hyperparameter ID 139. This model was used to score the AMS and Enamine REAL libraries. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a> directory for code to load the model and make predictions on new compounds.</li> <li><strong>train_ams_real_cluster.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering at a 0.4 threshold applied to the training compounds, 1,024 tested AMS compounds, and top-ranked compounds from Enamine. Includes the chemical features, dataset to which the compound belongs, leader compound for each cluster, and whether the compound is a known hit.</li> <li><strong>training_df_single_fold.csv.gz</strong>: This is all ten folds in&nbsp;<code>training_folds.tar.gz</code>&nbsp;merged for convenience. Contains 427,300 compounds.</li> <li><strong>training_df_single_fold_with_ams_clustering.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering applied to the 427,300 training compounds and the 1,024 tested AMS compounds. Different clustering results are shown at the 0.2, 0.3, and 0.4 thresholds. Includes the leader compound for each cluster. Although the training and AMS compounds were clustered jointly, only the training compounds&#39; clusters are shown. The AMS compounds&#39; clusters are in&nbsp;<code>ams_order_results.csv.gz</code>.</li> <li><strong>training_folds.tar.gz</strong>: The LC1234 and MLPCN training data split into ten folds. This dataset with 427,300 compounds was used for cross validation and model selection. This dataset is derived from&nbsp;<code>master_df.csv.gz.</code></li> </ul> <p>If you use&nbsp;these&nbsp;datasets in a publication, please cite:</p> <p>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</p> <p>See&nbsp;PubChem AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>, AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1918986">1918986</a>,&nbsp;and the associated publications for details about the PriA-SSB screening data. The screening datasets were compiled from three separate sources that should all be cited if the training dataset is used in a publication:</p> <ul> <li>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</li> <li>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.8b00363">Practical model selection for prospective virtual screening</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2018.</li> <li>Andrew F. Voter<sup>+</sup>, Michael P. Killoran<sup>+</sup>, Gene E. Ananiev, Scott A. Wildman, F. Michael Hoffmann, James L. Keck.&nbsp;<a href="https://doi.org/10.1177/2472555217712001">A high-throughput screening strategy to identify inhibitors of SSB protein&ndash;protein interactions in an academic screening facility</a>.&nbsp;<em>SLAS Discovery</em>&nbsp;2018.</li> </ul> <ul> </ul>

opencc-by-4.0Oct 2021View details →
edi52/100

Bioacoustic Dataset of African and Florida Manatee Vocalizations for Machine Learning Applications, 2020-2022

This data package presents a comprehensive acoustic library of manatee vocalizations for machine learning (ML) and classifier development. It includes recordings from two species, African and Florida manatees, sampled across four locations. The species are combined due to the acoustic similarity of their vocalizations, providing a diverse and representative training set for ML algorithms. The dataset consists of 0.5-second WAV clips categorized as either containing manatee vocalizations (MV, n=18,129 clips) or not (Noise, n=23,444 clips). MV clips may include multiple vocalizations or truncated calls. All clips were manually verified by two researchers with expertise in manatee acoustics. Recordings were collected using stationary hydrophones deployed in natural habitats, with variable signal-to-noise ratios (SNR) resulting from changes in distance between the vocalizing manatees and the recorders. Background noise across sites is relatively low, with minimal anthropogenic noise; caution is advised when applying models to noisier environments. No dolphin species are believed to be present at the recording sites, and models trained on this dataset should be used cautiously in dolphin-inhabited regions to avoid false positives. If you use this dataset, please reach out to the listed contacts, we are interested in learning how it supports your work.

openCC (other)Sep 2025View details →
edi52/100

Estimation of Abundance and Distribution of Salt Marsh Plants from Images Using Deep Learning

Recent advances in computer vision and machine learning, most notably deep convolutional neural networks (CNNs), are exploited to identify and localize various plant species in salt marsh images. Three different approaches are explored that provide estimations of abundance and spatial distribution at varying levels of granularity in terms of spatial resolution. In the coarsest-grained approach, CNNs are tasked with identifying which of six plant species are present/absent in large patches within the salt marsh images. CNNs with diverse topological properties and attention mechanisms are shown capable of providing accurate estimations with > 90% precision and recall in the case of the more abundant plant species whereas the performance of the CNNs is observed to decline in the case of less common plant species. Estimation of percent cover of each plant species is performed at a finer spatial resolution, where smaller image patches are extracted and the CNNs tasked with identifying the plant species or substrate at the center of the image patch. In an ecological setting, several image patches (~100) are extracted and classified using this approach to estimate the percent cover of the various plant species in the image. For the percent cover estimation task, the CNNs are observed to exhibit a performance profile similar to that for the presence/absence estimation task, but with an ~ 5–10% reduction in precision and recall. Finally, estimation of the spatial distribution of the various plant species is performed via semantic segmentation of the input images at the finest level of granularity in terms of spatial resolution. The Deeplab-V3 semantic segmentation architecture is observed to provide very accurate estimations for abundant plant species; however, a significant degradation in performance is observed in the case of less abundant plant species and, in extreme cases, rare plant classes are seen to be ignored entirely. Overall, a clear trade-off is observed between

openCC (other)Jan 2020View details →
OpenNeuro48/100

Motor sequence learning

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
OpenNeuro48/100

EEG: Probabilistic Learning with Affective Feedback: Exp #2

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
OpenNeuro48/100

EEG: Probabilistic Learning with Affective Feedback: Exp #1

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo48/100

The DNNLikelihood: enhancing likelihood distribution with Deep Learning

<p>Datasets and trained models&nbsp;corresponding to version 2 of <a href="https://arxiv.org/abs/1911.03305">arXiv:1911.03305</a> and complementing&nbsp;the code on&nbsp;<a href="https://github.com/riccardotorre/DNNLikelihood/releases/tag/1911.03305v2">GitHhub</a>.</p> <p>Notice that the code on GitHub includes scripts to automatically download these data.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record