Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Seafloor Density Measurements, Prediction, and Associated Uncertainty for "Predicting global marine sediment density using the random forest regressor machine learning algorithm"
<p>Global seafloor density prediction results using the random forest regressor machine learning algorithm. </p> <p>Dataset S1. Seafloor density measurements. Columns are labeled with a header and include associated drilling project and measurement type for each sample. File format: CSV text file</p> <p>Dataset S2. Seafloor density prediction results from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p> <p>Dataset S3. Seafloor density prediction standard deviation from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p>
Data and code for training and evaluating machine learning models for thunderstorm prediction from reanalysis data
<p>FIXED Data and Python code for training and evaluating machine learning models for predicting thunderstorms, associated with the paper:</p> <p>"Evaluation of machine learning classifiers for predicting deep convection"</p> <p>by Peter Ukkonen and Antti Mäkelä (to appear in JAMES)</p> <p>The data (preprocessed inputs and outputs) is stored as netCDF files and .mat files which can be loaded with Python. </p>
Characterization of descriptors in machine learning for data-based sputtering yield prediction
<p>figures and a table</p>
Data and code for paper "A gray-box model for a probabilistic estimate of regional ground magnetic perturbations: Enhancing the NOAA operational Geospace model with machine learning"
<p>Simulation results from the NOAA/SWPC Geospace model used in the paper Camporeale et al. (2020) "A gray-box model for a probabilistic estimate of regional ground magnetic perturbations: Enhancing the NOAA operational Geospace model with machine learning" published in J. Geophys. Res. (2020)</p> <p>MATLAB code is provided to train process the data and train the machine learning model and plot results.</p> <p>Manuscript available on <a href="https://arxiv.org/abs/1912.01038">https://arxiv.org/abs/1912.01038</a></p>
Supporting user preferences in search-based product line architecture design using Machine Learning
<p>The Product Line Architecture (PLA) is one of the most important artifacts of a Software Product Line. PLA design requires intensive human effort as it involves several conflicting factors. In order to support this task, an interactive search-based approach, automated by a tool named OPLA-Tool, was proposed in a previous work. Through this tool the software architect evaluates the generated solutions during the optimization process. Considering that evaluating PLA is a complex task and search-based algorithms demand a high number of generations, the evaluation of all solutions in all generations cause human fatigue. In this work, we incorporated in OPLA-Tool a Machine Learning (ML) model to represent the architect in some moments during the optimization process aiming to decrease the architect's effort. Through the execution of a quanti-qualitative exploratory study it was possible to demonstrate the reduction of the fatigue problem and that the solutions produced at the end of the process, in most cases, met the architect’s needs.</p>
Machine learning protocol code
<p>README.md</p> <p>This project "A Machine Learning Protocol for Predicting Protein Infrared Spectra" was supported by Prof. Shaul Mukamel(the University of California, Irvine), Prof. Jonathan D. Hirst(University of Nottingham),and Prof.Jun Jiang(University of Science and Technology of China).</p> <p>Simulation data and code of ML protocl for IR spectra of proteins.</p> <p>Any researchers who interested in protein spectroscopy can use our ML protcol online service: <a href="http://dcaiku.com:12880/platform/first">http://dcaiku.com:12880/platform/first</a></p> <p>For the machine learning protocol source code written in Python and Bash language which including:</p> <p>1.1.py: Split the protein into individual peptide bonds and dipeptides.</p> <p>1.2.py: Calculate the center of mass for each each peptide bond and dipeptide.</p> <p>1.3.sh: Convert the pdb format file to xyz format</p> <p>2.py: Extracte the Coulomb Matrix (CM) descriptors of peptide bond and dipeptide..</p> <p>3.py: Predict the vibrational frequency and vibrational transition dipole moment of each peptide bond from trained NMA Neural Networks (NN) model.</p> <p>4.py: Predict the neighboring coupling of each dipeptide from trained GLDP NN model.</p> <p>5.sh: Generate the input file for SPECTRON program to calculate the IR spectra of proteins.</p> <p>6.py: Construct the model Hamiltonian for amide I vibrations in a protein based on vibration exciton model theory.</p> <p>IR.sh: Diagonalize the Hamilton matrix calculate the IR spectra by using the SPECTRON.</p> <p> </p>
NEMO, HIDRA and Tide Gauge Datasets for HIDRA Machine Learning Algorithm Verification
<p>Supporting sea level datasets for paper:</p> <p>"HIDRA 1.0: Deep-Learning-Based Ensemble Sea Level Forecastingin the Northern Adriatic"</p> <p>by Lojze Žust, Anja Fettich, Matej Kristan, and Matjaž Ličer</p>
Predicting amphibian intraspecific diversity with machine learning: Challenges and prospects for integrating traits, geography, and genetic data
<p>The growing availability of genetic datasets, in combination with machine learning frameworks, offer great potential to answer long-standing questions in ecology and evolution. One such question has intrigued population geneticists, biogeographers, and conservation biologists: What factors determine intraspecific genetic diversity? This question is challenging to answer because many factors may influence genetic variation, including life history traits, historical influences, and geography, and the relative importance of these factors varies across taxonomic and geographic scales. Furthermore, interpreting the influence of numerous, potentially correlated variables is difficult with traditional statistical approaches. To address these challenges, we analyzed repurposed data using machine learning and investigated predictors of genetic diversity, focusing on Nearctic amphibians as a case study. We aggregated species traits, range characteristics, and >42,000 genetic sequences for 299 species using open-access scripts and various databases. After identifying important predictors of nucleotide diversity with random forest regression, we conducted follow-up analyses to examine the roles of phylogenetic history, geography, and demographic processes on intraspecific diversity. Although life history traits were not important predictors for this dataset, we found significant phylogenetic signal in genetic diversity within amphibians. We also found that salamander species at northern latitudes contain lower genetic diversity. Data repurposing and machine learning provide valuable tools for detecting patterns with relevance for conservation, but concerted efforts are needed to compile meaningful datasets with greater utility for understanding global biodiversity.</p>
An Artificial Eye for Palaeography. Applying Deep Machine Learning for the Study of Medieval Latin Scripts
<p>The project “Digital Forensics for Historical Documents” (at Huygens ING, Amsterdam) attempts to create a digital tool, based on a deep learning system, in which the unique characteristics of one medieval script sample will be matched with similar script samples by making use of digitized manuscript collections available in the world wide web.</p> <p>Project website and contact: <a href="https://www.youtube.com/redirect?q=https%3A%2F%2Fen.huygens.knaw.nl%2Fprojecten%2Fdigital-forensics-for-historical-documents%2F&v=WYtseNK-1Dc&event=video_description&redir_token=QUFFLUhqbDM0WDRJRjN0V3QzOXV6d0ZNbDB2TVYzV1hUQXxBQ3Jtc0ttLVJHMEE1RlRkZjJVV3poYnpQOHZseTFzZHVqWHRGR2k2eWpVWTJSSldtV2p4aFhKbFRIQTFjVHQ5YnM0Mkd4ajBzOTNtdE1yOVJLVktqWlhqWUgzZXc3YmNQMS1nNGFpZ3p2amNOTFNqTDVUSTdtZw%3D%3D">https://en.huygens.knaw.nl/projecten/...</a> </p> <p>Presented as a Lightning Talk for the Schoenberg Symposium 2020</p>
FORCE 2020 Well well log and lithofacies dataset for machine learning competition
<p>This well log dataset from 118 wells in the Norwegian Sea that has been used in the FORCE 2020 machine learning competition with seismic and wells to predict the lithofacies using machine learning models. </p> <p>The well logs have been slightly cleaned up and partially despiked.</p> <p>The lithofacies and lithology interpretation has been hand crafted using skilled geoscientists (Thanks to Explocrowd for excellent work). For citation in addition to the DOI please also refer to the github repository where the documentation and trained models reside</p> <p><a href="https://github.com/bolgebrygg/Force-2020-Machine-Learning-competition">https://github.com/bolgebrygg/Force-2020-Machine-Learning-competition </a></p> <p> </p> <p>The original well log data comes form the Norwegian government and is provided by a NOLD 2.0 license</p> <p> </p>
The Camouflage Machine: Optimising protective colouration using deep learning with genetic algorithms
Evolutionary biologists frequently wish to measure the fitness of alternative phenotypes using behavioural experiments. However, many phenotypes are complex. For example colouration: camouflage aims to make detection harder, while conspicuous signals (e.g. for warning or mate attraction) require the opposite. Identifying the hardest and easiest to find patterns is essential for understanding the evolutionary forces that shape protective colouration, but the parameter space of potential patterns (coloured visual textures) is vast, limiting previous empirical studies to a narrow range of phenotypes. Here we demonstrate how deep learning combined with genetic algorithms can be used to augment behavioural experiments, identifying both the best camouflage and the most conspicuous signal(s) from an arbitrarily vast array of patterns. To show the generality of our approach, we do so for both trichromatic (e.g. human) and dichromat (e.g. typical mammalian) visual systems, in two different habitats. The patterns identified were validated using human participants; those identified as the best for camouflage were significantly harder to find than a tried-and-tested military design, while those identified as most conspicuous were significantly easier than other patterns. More generally, our method, dubbed the 'Camouflage Machine', will be a useful tool for identifying the optimal phenotype in high dimensional state-spaces.
Machine learning reactive mixing dataset-0
<p>Reactive-mixing dataset-0 for machine learning analyses</p>
Machine learning reactive mixing dataset-2
<p>Reactive-mixing dataset-1.1 for machine learning analyses</p>
Machine learning reactive mixing dataset-1
<p>Reactive-mixing dataset-1 for machine learning analyses</p>
Machine learning reactive mixing dataset-9
<p>Reactive-mixing dataset-3 for machine learning analyses</p>
Machine learning reactive mixing dataset-4
<p>Reactive-mixing dataset-1.3 for machine learning analyses</p>
Machine learning reactive mixing dataset-3
<p>Reactive-mixing dataset-1.2 for machine learning analyses</p>
Machine learning reactive mixing dataset-8
<p>Reactive-mixing dataset-2.3 for machine learning analyses</p>
Machine learning reactive mixing dataset-7
<p>Reactive-mixing dataset-2.2 for machine learning analyses</p>
Machine learning reactive mixing dataset-6
<p>Reactive-mixing dataset-2.1 for machine learning analyses</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.