Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
A Machine Learning Approach for Real-time Cortical State Estimation
<p>Data and code accompanying the following publication: Weiss, D. A., Borsa, A. M., Pala, A., Sederberg, A. J., & Stanley, G. B. (2024). A machine learning approach for real-time cortical state estimation. <em>Journal of neural engineering</em>, <em>21</em>(1), 10.1088/1741-2552/ad1f7b. https://doi.org/10.1088/1741-2552/ad1f7b</p>
Landslide susceptibility maps using ensemble machine learning models on basin and regional level in Lombardy, Italy
<p>A selection of landslide susceptibility maps computed through ensemble machine learning models for the basins of Val Tartano, Upper Valtellina and Valchiavenna, and on a regional level for the Lombardy region in Italy.</p> <p>A list of the used base machine learning methods:</p> <ul> <li>Random Forest,</li> <li>AdaBoost,</li> <li>Neural Networks.</li> </ul> <p>A list of the used ensemble models:</p> <ul> <li>Stacking,</li> <li>Blending,</li> <li>Soft Voting.</li> </ul> <p>A full list of the model combinations can be found in the "Case Studies" document.</p> <p>The maps are in WGS 84/ UTM zone 32N (EPSG:32632).</p> <p>The map production process details are discussed in Xu et al. 2024. If you use the dataset, please, cite also the paper:</p> <p><em>Qiongjie Xu, Vasil Yordanov, Lorenzo Amici & Maria Antonia Brovelli (2024) Landslide susceptibility mapping using ensemble machine learning methods: a case</em><br><em>study in Lombardy, Northern Italy, International Journal of Digital Earth, 17:1, 2346263, DOI:10.1080/17538947.2024.2346263</em></p> <p>The maps are produced as part of the "Geoinformatics and Earth Observation for Landslide Monitoring" Italy-Vietnam.</p> <p>The work is partially funded by the Italian Ministry of Foreign Affairs and International Cooperation within the project “Geoinformatics and Earth Observation for Landslide Monitoring” CUP D19C21000480001.</p> <p> </p>
[Supplementary material] Machine Learning-driven Testing of Web APIs
<p>This is the supplementary material of the paper entitled "Machine Learning-driven Testing of Web APIs".</p>
Using knowledge-guided machine learning to assess patterns of areal change in waterbodies across the contiguous U.S.: Data
<p>Data used for generating figures in the knowledge-guided machine learning manuscript and supplemental info by Wander et al.</p><p>This repository contains three folders:</p><ul><li><strong>Results: </strong>csv file with the final KGML groups for each waterbody. Latitude, longitude, RealSAT (Khandelwal et al., 2022) waterbody id, and HydroLAKES (Messager et al., 2016) waterbody id are provided.</li><li><strong>Code_data</strong>: data used for generating figures in the knowledge-guided machine learning manuscript and supplemental info by Wander et al.</li><li><strong>prism</strong>: data from PRISM dataset (Matsuura and Willmott) used for preliminary driver analysis in Wander et al.</li></ul><p>Khandelwal, A.; Karpatne, A.; Ravirathinam, P.; Ghosh, R.; Wei, Z.; Dugan, H.; Hanson, P. C.; Kumar, V. ReaLSAT, a global dataset of reservoir and lake surface area variations. <i>Sci. Data</i> <strong>2022</strong>, <i>9</i>, 356. <a href="https://doi.org/10.1038/s41597-022-01449-5">https://doi.org/10.1038/s41597-022-01449-5</a></p><p>Messager, M. L.; Lehner, B.; Grill, G.; Nedeva, I.; Schmitt, O. Estimating the volume and age of water stored in global lakes using a geo-statistical approach. <i>Nat. Commun.</i> <strong>2016</strong>, <i>7</i>, 1–11. <a href="https://doi.org/10.1038/ncomms13603">https://doi.org/10.1038/ncomms13603</a><br><br>Willmott, C. J.; Matsuura, K. Terrestrial Air Temperature and Precipitation: 1900-2014 Gridded Monthly Time Series: NOAA Physical Sciences Laboratory Terrestrial Air Temperature and Precipitation: 1900-2014 Gridded Monthly Time Series, <strong>2015</strong>. <a href="https://psl.noaa.gov/data/gridded/data.UDel_AirT_Precip.html">https://psl.noaa.gov/data/gridded/data.UDel_AirT_Precip.html</a>. (Accessed Jan 2022).</p><p>Danielson, J. J.; Gesch, D. B. TEMIS -- GMTED2010 Elevation Data at Different Resolutions, <strong>2011</strong>. <a href="https://www.usgs.gov/centers/eros/science/usgs-eros-archive-digital-elevation-global-multi-resolution-terrain-elevation">https://www.usgs.gov/centers/eros/science/usgs-eros-archive-digital-elevation-global-multi-resolution-terrain-elevation</a>. (Accessed Jan 2022).</p>
Exploring the Global Reaction Coordinate for Retinal Photoisomerization: A Graph Theory-Based Machine Learning Approach
<p>This repository contains i. optimized geometry of the cis and trans retinal, ii. figure labelling the internal coordinates of retinal. </p>
Supplementary Data: Uncovering DFG-out sequence propensity determinants of kinases with machine learning
<h2>General description</h2><p>This submission accompanies the paper "Uncovering DFG-out sequence propensity determinants of kinases with machine learning" and covers:</p><ul><li><strong>Trained models (train/models_latest)</strong>. A subset of these is also published as a part of the <a href="https://github.com/edikedik/kinactive">KinActive tool</a>; each model is of `KinactiveClassifier` type and can be loaded using this tool as well via `kinactive.io.load()` function).</li><li><strong>Patched sequences (patched_seqs)</strong>. PDB sequences, with missing regions patched by UniProt sequences. This is a collection of <a href="https://github.com/edikedik/lXtractor">lXtractor</a> ChainSequence objects.</li><li><strong>All labels, variables, and datasets</strong> (datasets and labels).</li><li><strong>Initial chains and predictions for SwissProt proteins</strong> (SP_predictions).</li></ul><h2>Note on model abbreviations</h2><p>The main text emphasized datasets that yielded more interpretable results. These were constructed from domain sequences, labeled as apo, inactive, DFG-in, or DFG-out, and further divided into TK and STK subsets. We refer to these datasets with the abbreviation <strong>AAIO</strong> (Apo All In or Out) to distinguish them from additional datasets.</p><p>In addition to <strong>AAIO</strong>, we explored two alternative labeling strategies:</p><ul><li><strong>AHAO</strong> (Apo/Holo Any Out): This dataset includes sequences from ligand-bound entries. All sequences within 95\% identity clusters are labeled as DFG-out if the cluster contains at least one sequence in this state. All others are labeled as DFG-in.</li><li><strong>AAO</strong> (Apo Any Out): This dataset excludes sequences corresponding to ligand-bound entries but includes those with conflicting conformational tendencies within 95\% identity clusters.</li></ul><p>Each of these datasets had two versions:</p><ol><li>A seed version, denoted by a "*" symbol.</li><li>A version enriched with orthologous sequences (no special designation).</li></ol><p>Together with the <strong>TkST</strong> datasets (encompassing TK and STK labels) used for testing the methodology, a total of 14 datasets were used, and both <i>RF</i> and <i>XGB</i> models were applied to each, using the same initial settings, resulting in 28 different models. This additional information is provided for completeness.</p>
Graph Machine Learning Dataset LPWC
<p><strong>LPWC</strong> is a <strong>heterogenous graph machine learning dataset</strong> based on the RDF knowledge graph <a href="https://arxiv.org/pdf/2310.20475.pdf">Linked Papers With Code</a> (version from 2023-06-24).</p><p>LPWC contains four node types - papers (376,557 nodes), datasets (8,322 nodes), tasks (4,267 nodes) and methods (2,101), and six edge types.</p><p>Each node has rich semantic node features for node representation (content-based and topology-based node features are available).</p><p> </p><p>More information can be found in README.txt and on <a href="https://github.com/davidlamprecht/AutoRDF2GML">https://github.com/davidlamprecht/AutoRDF2GML</a></p>
Graph Machine Learning Dataset SOA-SW
<p><strong>SOA-SW</strong> is a heterogeneous graph machine learning dataset based on the RDF knowledge graph <a href="doi.org/10.5281/zenodo.10299132">SemOpenAlex-SemanticWeb</a>.</p><p>SOA-SW contains six node types - works (95,575 nodes), authors (19,970 nodes), concepts (38,050 nodes), sources (10,739 nodes), institutions (5,846 nodes), and publishers (786 nodes), and seven edge types. </p><p>Each node has rich semantic node features as node representation (content-based and topology-based node features are available).</p><p>More information can be found in the README.txt and on <a href="https://github.com/davidlamprecht/AutoRDF2GML">https://github.com/davidlamprecht/AutoRDF2GML.</a></p><p> </p><p><strong>soa-sw-homogeneous-author</strong> only models the co-author network of SOA-SW.</p><p>It is a homogeneous graph containing the author node type (19,970 nodes) and the edge type author- author.</p><p>The authors' content-based features (nodes-nld) are based on the titles and abstracts of the authors' works (128-dimensional SciBERT embeddings).</p>
Machine Learning Augmented DA Code and Data
Open the record for dataset details and reuse information.
Supplemental Data for the Journal Article Entitled Accelerating FEM-based Corrosion Predictions using Machine Learning submitted for publication to the Journal of the Electrochemical Society.
<p>This repository contains the Supplemental Data for the Journal Article Entitled Accelerating FEM-based Corrosion Predictions using Machine Learning submitted for publication to the Journal of the Electrochemical Society.</p> <p>Authors:</p> <p>David Montes de Oca Zapiain 1, Demitri Maestas 1, Matthew Roop 1,2, Philip Noel 1, Michael Melia 1, Ryan Katona 1</p> <p>1 Sandia National Laboratories, Albuquerque, NM 87185, USA <br>2 University of New Mexico, Albuquerque, NM 87131, USA</p> <p>Each folder contains a ReadMe.txt describing the files and their organization within each folder. </p> <p><br>Acknowledgements:<br>Sandia National Laboratories is a multi-mission laboratory managed and operated by National Technology and Engineering Solutions of Sandia, LLC., a wholly owned subsidiary of Honeywell International, Inc.,<br>for the U.S. Department of Energy National Nuclear Security Administration under contract DE-NA0003525. The views expressed in the article<br>do not necessarily represent the views of the U.S. Department of Energy or the United States Government. SAND No: SAND2023-14419O</p>
Code and data for "Machine-learning-boosted ab-initio study of the thermal conductivity of Janus PtSTe van der Waals heterostructures"
<h1>Code and data for <em>Machine-learning-boosted ab-initio study of the thermal conductivity of Janus PtSTe van der Waals heterostructures</em></h1> <p> </p> <h2>Contents:</h2> <ul> <li><strong>neuralil.tar.xz</strong>: version used in the manuscript of the force-field code described in the articles <a href="https://doi.org/10.1021/acs.jcim.1c01380">A Differentiable Neural-Network Force Field for Ionic Liquids</a> and <a href="https://doi.org/10.1063/5.0146905">Deep ensembles vs committees for uncertainty estimation in neural-network force fields: Comparison and application to active learning</a>. General-purpose releases can be found <a href="https://github.com/Madsen-s-research-group/neuralil-public-releases">here</a>.</li> <li><strong>DFT_data.tar.xz</strong>: first-principles data created for training and validating the force field, stored as <a href="https://wiki.fysik.dtu.dk/ase/ase/db/db.html">ASE databases</a> in JSON format.</li> <li><strong>model_params_plain_ensemble_DEEP_413E12A9.pkl</strong>: saved parameters of the fully trained force field.</li> <li><strong>0001-Use-equipartition-occupancies.patch</strong>: patch for <a href="https://phonopy.github.io/phono3py">Phono3py</a> to use classical (equipartition) occupations instead of Bose-Einstein values.</li> </ul>
Machine learning simulation protocol for protein amide II spectra
<p>Machine learning simulation data for protein amide II spectra.</p>
Supplementary Information: CHAPTER 3 - Classification of genomic features of plant-associated bacteria using machine learning
<p>Appendix A- List of all bacterial genomes used in orthologous genes clustering in the feature extraction step and in the further steps to build and test classifiers’ models. The list includes the isolation source information and the related category for the genome classification and features selection purposes.</p> <p>Appendix B - Distribution of genomes by phylum, family, and genus among the categories defined according to bacteria lifestyle association.</p> <p>Appendix C - Enriched orthogroups by genus according to each enrichment test (Material and Methods). Values for each test are "Y" (enriched), "N" (not enriched), or "Untested" (clusters were untested when there was insufficient phylogenetic signal, they were too small or were found in all genomes).</p> <p>Appendix D - Classification performance of random forest and logistic regression techniques applied to genus-specific datasets of genomic features (orthogroups) using both matrices from gene count number and presence/absence values. Sensitivity is a measure of how well a test identifies true positives; Specificity: is a measure how well a test or model avoids false positives; Positive Predictive Value (Pos. Pred. Value): The probability that a positive prediction is correct; Negative Predictive Value (Neg. Pred. Value): The probability that a negative prediction is correct; Precision: The accuracy of positive predictions; Recall (Sensitivity): The ability to find all relevant cases; F1 Score: A combined measure of precision and recall; Prevalence: The proportion of positive cases in the total; Detection Rate: The proportion of true positive cases identified; Detection Prevalence: The proportion of positive predictions; Balanced Accuracy: An average of sensitivity and specificity; Area Under the Curve (AUC): The overall performance of the model in distinguishing between positive and negative cases.</p> <p>Appendix E - Orthogroups assigned with predicted COGs as an important feature for classifying plant-associated genomes. COG categories: A - RNA processing and modification; B - Chromatin structure and dynamics; C - Energy production and conversion; D - Cell cycle control, cell division, chromosome partitioning; E - Amino acid transport and metabolism; F - Nucleotide transport and metabolism; G - Carbohydrate transport and metabolism; H - Coenzyme transport and metabolism; I - Lipid transport and metabolism; J - Translation, ribosomal structure and biogenesis; K - Transcription; L - Replication, recombination and repair; M - Cell wall/membrane/envelope biogenesis; N - Cell motility; O - Posttranslational modification, protein turnover, chaperones; P - Inorganic ion transport and metabolism; Q - Secondary metabolites biosynthesis, transport and catabolism; R - General function prediction only; S - Function unknown; T - Signal transduction mechanisms; U - Intracellular trafficking, secretion, and vesicular transport; V - Defense mechanisms; W - Extracellular structures; X - Mobilome: prophages, transposons; Y - Nuclear structure; Z - Cytoskeleton.</p>
Potential and training data for 'Structure-property relations of silicon oxycarbides studied using a machine learning interatomic potential'
<p>Fitted potential, training and testing data.</p>
dataset for machine learning example in WarpX / ImpactX
<p>This contains data for a machine learning workflow using codes from the BLAST suite supported at Lawrence Berkeley National Laboratory. There are 2 directories: </p> <p>lab_particle_diags: contains particle diagnostics in openPMD format for use in training neural networks<br>models: parameters of pre-trained neural networks for use in inference</p>
Innovative Strategies for Blood-Brain Barrier (BBB) Permeability Modeling: Harnessing the Power of Machine Learning-based q-RASAR Approach
<p>In the current research, we have unveiled an advanced technique termed the quantitative Read-Across Structure-Activity Relationship (q-RASAR) framework to harnesses the power of machine learning (ML) for significantly enhancing the precision of predictions related to blood-brain barrier (BBB) permeability. It is important to emphasize that the central objective of this study is not to introduce another model for predicting BBB permeability. Instead, our focus is on highlighting the improvement in predicting the BBB permeability of organic compounds by introducing the q-RASAR approach. This innovative methodology strives to enhance the precision of evaluating neuropharmacological implications and streamline the drug development process. In this investigation, we developed an ML-based q-RASAR PLS model using a dataset comprising 1012 compounds of diverse classes of heterocyclic and aromatic hydrocarbons, obtained from the freely accessible B3DB database (accessible at <a href="https://github.com/theochem/B3DB">https://github.com/theochem/B3DB</a>) to predict BBB permeability during the lead discovery phase for central nervous system (CNS) drugs. The model's predictive capability underwent validation using two external sets, encompassing a total of 1,130,315 compounds, including synthetic compounds and natural products (NPs) for data gap filling and other two external sets comprising 116 drug-like/drug compounds from FDA and ChEMBL databases to assess the model's reliability. This study aimed to bridge the data gap by employing a predictive model to estimate the impact of brain-plasma concentration ratios on BBB permeability for both synthetic compounds and natural products (NPs). To further enhance predictability, we have developed various other ML-based q-RASAR models. The insights from the developed model highlight the pivotal roles played by hydrophobicity, electronic effects, degree of ionization and steric factors as essential features facilitating the traversal of the blood-brain barrier. This research not only advances our understanding of the molecular determinants influencing the permeability of central nervous system drugs but also establishes a versatile computational platform for the rapid assessment of diverse compounds, facilitating informed decision-making in the realms of drug development and design.</p>
Machine Learning and Radiomics analysis by Computed Tomography in colorectal liver metastases patients for RAS mutational status prediction
<p>We uploaded the Radiomics features raw data of the manuscript "<span>Machine Learning and Radiomics analysis by Computed Tomography in colorectal liver metastases patients for RAS mutational status prediction</span>"</p>
Machine Learning Potential for Modelling H2 Adsorption/Diffusion in MOFs with Open Metal Sites
<p>Data for reproducibility of the paper "Machine Learning Potential for Modelling H2 Adsorption/Diffusion in MOFs with Open Metal Sites"</p> <p>Corresponding paper:</p> <p>ShanPing Liu, Romain Dupuis, Dong Fan, Salma Benzaria, Mickaele Bonneau, Prashant Bhatt, Mohamed Eddaoudi, and Guillaume Maurin. "Machine Learning Potential for Modelling H2 Adsorption/Diffusion in MOFs with Open Metal Sites." <em>Chemical Science</em> (2024). https://doi.org/10.1039/D3SC05612K</p>
Research data for "Understanding Defects in Amorphous Silicon with Million-Atom Simulations and Machine Learning"
<p>This dataset contains structural data and ML local energies for the publication "Understanding Defects in Amorphous Silicon with Million-Atom Simulations and Machine Learning". Details of the contents can be found in README.txt. Code to analyse structures and reproduce figures from the publication are available at https://github.com/MorrowChem/understanding_defects</p>
To enhance CO2 saturation prediction from seismic data by joint use of attenuation and elastic properties: A machine learning approach
<p>Wang and Zhao submitted for GJI</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.