Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

A Building Imagery Database for the Calibration of Machine Learning Algorithms

<p>An open database of 5276 building images from a parish in Lisbon (Alvalade), whose buildings have been classified according to a uniform taxonomy. This open database can be used for the testing and calibration of machine learning algorithms, as well as for the direct assessment of earthquake risk in Alvalade.</p> <p><strong>Full Changelog</strong>: <a href="https://github.com/vsilva028/ML/commits/v1.0.0">https://github.com/vsilva028/ML/commits/v1.0.0</a></p>

openother-openFeb 2023View details →
zenodo36/100

Raw Data for the publication of 'Machine-Learning of Piezoelectric Coefficients for Wurtzite Crystals'

<p>The dataset used for the ML model described in&nbsp;<strong>Machine-Learning of Piezoelectric Coefficients for Wurtzite Crystals.&nbsp;</strong></p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
dryad36/100

Data for: Advances and critical assessment of machine learning techniques for prediction of docking scores

<p class="MsoNormal"><span>Semi-flexible docking was performed using AutoDock Vina 1.2.2 software on the SARS-CoV-2 main protease M<sup>pro</sup> (PDB ID: 6WQF). <br></span></p> <p class="MsoNormal"><span>Two data sets are provided in the xyz format containing the AutoDock Vina docking scores. These files were used as input and/or reference in the machine learning models using TensorFlow, XGBoost, and SchNetPack to study their docking scores prediction capability. The first data set originally contained 60,411 in-vivo labeled compounds selected for the training of ML models. The second data set,denoted as in-vitro-only, originally contained 175,696 compounds active or assumed to be active at 10 μM or less in a direct binding assay. These sets were downloaded on the 10th of December 2021 from the ZINC15 database. Four compounds in the in-vivo set and 12 in the in-vitro-only set were left out of consideration due to presence of Si atoms. Compounds with no charges assigned in mol2 files were excluded as well (523 compounds in the in-vivo and 1,666 in the in-vitro-only set). Gasteiger charges were reassigned to the remaining compounds using OpenBabel. In addition, four in-vitro-only compounds with docking scores greater than 1 kcal/mol have been rejected.</span></p> <p class="MsoNormal"><span>The provided in-vivo and the in-vitro-only sets contain 59,884 (in-vivo.xyz) and 174,014 (in-vitro-only.xyz) compounds, respectively. Compounds in both sets contain the following elements: H, C, N, O, F, P, S, Cl, Br, and I. The in-vivo compound set was used as the primary data set for the training of the ML models in the referencing study. <br></span></p> <p class="MsoNormal"><span>The file in-vivo-splits-data.csv contains the exact composition of all (random) 80-5-15 train-validation-test splits used in the study, labeled I, II, III, IV, and V. Eight additional random subsets in each of the in-vivo 80-5-15 splits were created to monitor the training process convergence. These subsets were constructed in such a manner, that each subset contains all compounds from the previous subset (starting with the 10-5-15 subset) and was enlarged by one eighth of the entire (80-5-15) train set of a given split. These subsets are further referred to as in_vivo_10_(I, II, ..., V), in_vivo_20_(I, II, ..., V),..., in_vivo_80_(I, II, ... V). </span></p>

opencc-zeroMar 2023View details →
zenodo36/100

Machine learning based lightning parameterizations for CONUS

<p>The dataset used for training&nbsp;Machine Learning models for lightning parameterization for CONUS, as well as the code to reproduce the analysis. The trained models with network weights are also available.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Machine Learning Techniques Application for Installation Torque Prediction of Helical Piles

<p>This database contains torque observations in the installation of helical piles used as a foundation in an infrastructure construction for power transmission towers. In this way, this database trains ML models for torque prediction.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)

<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder &quot;pattern&quot; - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder &quot;shapes&quot; - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File &quot;swemls-ontology.ttl&quot; - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File &quot;swemls-instances.ttl&quot; - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File &quot;swemls-kg.ttl&quot; - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using &quot;swemls-toolkit&quot; [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Estimating global GPP from the plant functional type perspective using a machine learning approach

<p><span>The long-term monitoring of gross primary production (GPP) is crucial to the assessment of the carbon cycle of terrestrial ecosystems. In this study, a well-known machine learning model (Random Forest, RF) is established to reconstruct the global GPP dataset named ECGC_GPP. The model distinguished nine functional plant types, including C3 and C4 crops, using eddy fluxes, meteorological variables, and leaf area index as training data of the RF model. Based on ERA5_Land and the corrected GEOV2 data, the global monthly GPP dataset at a 0.05-degree resolution from 1999 to 2019 was estimated. The results showed that the RF model could explain 74.81% of the monthly variation of GPP in the testing dataset, of which the average contribution of Leaf Area Index (LAI) reached 41.73%. The average annual and standard deviation of GPP during 1999–2019 were 117.14 ± 1.51 Pg C yr<sup>-1</sup>, with an upward trend of 0.21 Pg C yr<sup>-2</sup> (<em>p</em> &lt; 0.01). By using the plant functional type classification, the underestimation of cropland is improved. Therefore, ECGC_GPP provides reasonable global spatial patterns and long-term trends of annual GPP.</span></p>

opencc-zeroMar 2023View details →
dryad36/100

ECG and EEG stress features for: ECG and EEG based detection and multilevel classification of stress using machine learning for specified genders: A preliminary study

<p>Mental health, especially stress, plays a crucial role in the quality of life. During different phases (luteal and follicular phases) of the menstrual cycle, women may exhibit different responses to stress from men. This, therefore, may have an impact on stress detection and classification accuracy of machine learning models that genders are not taken into account. However, this has never been investigated before. In addition, only a handful of stress detection devices are scientifically validated. To this end, this work proposes stress detection and multilevel stress classification models for unspecified and specified genders through ECG and EEG signals. Models for stress detection are achieved through developing and evaluating multiple individual classifiers. On the other hand, stacking technique is employed to obtain models for multilevel stress classification. ECG and EEG features extracted from 40 subjects (21 females and 19 males) were used to train and validate the models. In the low&amp;high combined stress condition, RBF-SVM and kNN yielded the highest average classification accuracy for females (79.81%) and males (73.77%), respectively. Combining ECG and EEG, the average classification accuracy increased to at least 87.58% (male, high stress) and up to 92.70% (female, high stress). For multilevel stress classification from ECG and EEG, the accuracy for females was 62.60% and for males was 71.57%. This study shows that the difference in genders influences the classification performance for both the detection and multilevel classification of stress. The developed models can be used for both personal (through ECG) and clinical (through ECG and EEG) stress monitoring with and without taking genders into account.</p>

opencc-zeroMar 2023View details →
zenodo36/100

Replication Package for Identifying Self-Admitted Technical Debt in Issue Tracking Systems using Machine Learning

<p>This dataset includes pre-trained word embeddings and a weighted file that can be used to identify self-admitted technical debt (SATD) from issue tracking systems.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data supporting "Machine Learning-based Modeling of Olfactory Receptors: Human OR51E2 as a Case Study"

<p>Simulation data and input files in support of the Manuscript:&quot;Machine Learning-based Modeling of Olfactory Receptors: Human OR51E2 as a Case Study&quot;.<br> <br> The archive is organized in 6&nbsp;different folders:</p> <p>1. <strong>7x7_rmsd</strong>, which contains a tcl script (to be run in VMD) to compute the 7x7 RMSD matrix (see Wang et al. J Struct. Bio, (2017)).<br> 2. <strong>a100_plumed</strong>, which contains PLUMED input files to compute A<sup>100&nbsp;</sup>index on OR51E2 trajectories.<br> 3. <strong>initial_structures</strong>, which contains the 6 different conformation for hOR51E2 obtained by the different predictors in the pdb format.<br> 4.&nbsp;<strong>mdp_files</strong>, which contains the GROMACS mdp files to perform all the protocol described in the paper.<br> 5. <strong>topologies</strong>, which contains the 6 different topologies (in .top format) and the initial conformation (in .gro format)&nbsp;for the hOR51E2 embedded in the membrane and solvated.<br> 6. <strong>trajectories</strong>, which contains the 18 (6 systems, 3 replicas per system)&nbsp;different trajectories with the sodium in place close to D69<sup>2.50</sup> without solvent and ions&nbsp;(in .xtc format with a frame every 100 ps) and a reference conformation (in .gro format). Here we have also added the trajectory for the SwissModel-derived simulation without sodium ion in place, where we observed the ion binding. This last trajectory contains also solvent and ions, but with a lower printing frequency (1 ns).</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Supporting Data for "Does a Machine-Learned Potential Perform Better Than an Optimally Tuned Traditional Force Field? A Case Study on Fluorohydrins"

<p>Supporting Data for &quot;Does a Machine-Learned Potential Perform Better Than an Optimally Tuned Traditional Force Field? A Case Study on Fluorohydrins&quot;</p>

opencc-by-4.0Nov 2022View details →
dryad36/100

Data from: Machine learning confirms new records of maniraptoran theropods in Middle Jurassic UK microvertebrate faunas

<p>Current research suggests that the initial radiation of maniraptoran theropods occurred in the Middle Jurassic, although their fossil record is known almost exclusively from the Cretaceous. However, fossils of Jurassic maniraptorans are scarce, usually consisting solely of isolated teeth, and their identifications are often disputed. Here, we apply different machine learning models, in conjunction with morphological comparisons, to a suite of isolated theropod teeth from Bathonian microvertebrate sites in the UK in order to determine if any of these can be confidently assigned to Maniraptora. We generated three independent models developed on a training dataset with a wide range of theropod taxa and broad geographical and temporal coverage. Classifying the Middle Jurassic teeth in our sample against these models indicates the presence of at least three distinct dromaeosaur morphotypes, plus a therizinosaur and troodontid, in these assemblages, a conclusion supported by morphological comparison. These new referrals significantly extend the ranges of Therizinosauroidea and Troodontidae, by some 27 million years. These results indicate that not only were maniraptorans present in the Middle Jurassic, as predicted by previous phylogenetic analyses, but had already radiated into a diverse fauna that pre-dated the break-up of Pangaea. This study also demonstrates the power of machine learning to provide quantitative assessments of isolated teeth in providing a robust, testable framework for taxonomic identifications, and highlights the importance of assessing and including evidence from microvertebrate sites in faunal and evolutionary analyses.</p>

opencc-zeroApr 2023View details →
zenodo36/100

Flnc: Machine Learning Improves the Identification of Novel Long Noncoding RNAs from Stand-Alone RNA-Seq Data

<p>Flnc is software that can accurately identify full-length long noncoding RNAs (lncRNAs) from human RNA-seq data. lncRNAs are linear transcripts of more than 200 nucleotides that do not encode proteins. The most common approach for identifying lncRNAs from RNA-seq data which examines the coding abilities of assembled transcripts will result in a very high false-positive rate (30%-75%) of lncRNA identification. The falsely discovered lncRNAs lack transcriptional start sites and most of them are RNA fragments or result from transcriptional noise. Unlike the false-positive lncRNAs, true lncRNAs are full-length lncRNA transcripts that include transcriptional start sites (TSSs). To exclude these false lncRNAs, H3K4me3 chromatin immunoprecipitation sequencing (ChIP-seq) data had been used to examine transcriptional start sites of putative lncRNAs, which are transcripts without coding abilities. However, because of cost, time, and the limited availability of sample materials for generating H3K4me3 ChIP-seq data, most samples (especially clinical biospecimens) may have available RNA-seq data but lack matched H3K4me3 ChIP-seq data. This Flnc method solves the problem of lacking transcriptional initiation profiles when identifying lncRNAs.</p> <p>Flnc integrates seven machine-learning algorithms built with four genomic features. Flnc achieves state-of-the-art prediction power with a AUROC score over 0.92. Flnc significantly improves the prediction accuracy from less than 50% using the common approach to over 85% on five independent datasets without requiring matched H3K4me3 ChIP-seq data. In addition to the stranded polyA-selected RNA-seq data, Flnc can also be applied to identify lncRNAs from stranded RNA-seq data of ribosomal RNA depleted samples or unstranded RNA-seq data of polyA-selected samples.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Machine learning enabled identification of sheet metal localization

<p>Data and Machine Learning codes for:</p> <ul> <li><strong>Machine learning enabled identification of sheet metal localization</strong></li> <li>Journal:&nbsp;<strong>International Journal of Solids and Structures</strong></li> </ul> <p><strong>Abstract</strong>: The Forming Limit Curve (FLC), which describes the maximum applicable strain before localization, depends on the particular material, but also on the applied load and history of the load. Recent investigations have shown that the non-<br> proportional loading effect on the FLC can be predicted with data-driven or machine-learning-based methods. Here<br> we compare different ML methods to their applicability in predicting localization point under multi-segmented non-<br> proportional loading. Therefore, a FE-based metamodel is developed that allows imposing an arbitrary loading history<br> on a sheet metal to predict the point of localization. A series of virtual experiments are conducted with this metamodel<br> to generate a database of bi-linear loading paths that are used for training. Different ML-based methods were used<br> to predict the localization point based on the strain history data. The 1D-Convolutional Neural Network (1D-CNN),<br> with an ability to learn dependency between input features, has the best accuracy in predicting the localization point.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Development of a Machine Learning-Based Model to Determine the Optimum and Safe Restriping Timing of Thermoplastic Pavement Markings in Hot and Humid Climates

<p>Due to limited budget, most transportation agencies restripe their thermoplastic pavement markings based on a fixed schedule or based on visual inspection instead of monitoring the retroreflectivity and restriping when the retroreflectivity drops below a pre-determined threshold. These strategies are questionable in terms of efficiency and economy. Therefore, previous studies proposed degradation models to predict the retroreflectivity of thermoplastic markings based on key variables. Yet, most of these studies reported low R<sup>2</sup> (as low as 0.1), which placed little confidence in these models.&nbsp; Therefore, the objective of this study was to evaluate and predict the field performance of thermoplastics and to propose cost-effective restriping strategies for thermoplastics used in hot and humid climate service conditions. To achieve this objective, National Transportation Product Evaluation Program (NTPEP) data were mined and analyzed. Results indicated that the service life (SL) of thermoplastics ranged between 0.4 and 12.1 years (according to the initial retroreflectivity, traffic, and surface type) with an average value of 3.4 &plusmn; 0.2 years. Four regression models with relatively high accuracy were developed to predict the SL of thermoplastics based on key variables. In addition, the genetic algorithm was used to develop a model that predicts the future retroreflectivty of these pavement markings. The predicted values were compared against actual retroreflectivity measurements collected from a field experiment at Louisiana State University. The results of this study could be used to make effective decisions related to restriping scheduling. Using the proposed models in restriping scheduling can result in considerable cost savings (up to $8,212 per lane-mile), as compared to the conventional restriping strategy.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Global runoff partitioning based on Budyko-constrained machine learning

<p>Global 0.25&deg;&nbsp;datasets&nbsp;of runoff partitioning are developed by Budyko-constrained machine learning. The datasets include runoff (<em>Q</em>, TFlow_0.25), runoff coefficient (<em>Q</em>/<em>P</em>, TFC_0.25), baseflow (<em>Q<sub>b</sub></em>, BFlow_0.25) and baseflow coefficient (<em>Q<sub>b</sub></em>/<em>P</em>, BFC_0.25).</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Development of Distress Index Prediction Models for Rehabilitation Treatments in Louisiana Using Advanced Machine Learning Techniques

<p>Performance prediction models are used by state agencies to predict future trends in distress indices, hence, determining the required maintenance and/or rehabilitation treatment as well as the deterioration rate and remaining pavement service life. However, most of these models are based on a limited number of parameters and cannot predict the performance distress indices reliably. Such limitation resulted in having, most of the time, a maximum prediction period of five years. As a solution and coping with the ever-increasing size of pavement data, machine learning techniques have become a promising alternative. The objective of this study was to develop a machine-learning-based framework for states with a hot and humid climate that can predict the long-term field performance (for 11 years) of their asphalt (AC) overlays based on their key project conditions. Two machine learning algorithms were examined, namely Random Forest (RF) and CatBoost, and the one yielding a higher accuracy was considered. In this study, the well-known pavement condition index (PCI) was used as the pavement performance indicator. A total of 892 log miles of AC overlay data were obtained from the Louisiana Department of Transportation and Development (LaDOTD) Pavement Management System (PMS) database.&nbsp; Based on the collected data, six models were trained (for each algorithm) and validated to predict the future PCI of AC overlays for up to 11 years. Results indicated that the RF algorithm yielded higher accuracy than the CatBoost Algorithm and thus the RF-based models were considered in the proposed decision-making framework.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Tailored machine learning for evaluating the long-term diabetes risk in older individuals: Findings from The Irish Longitudinal Study on Ageing (TILDA)

<p>The original data for &quot;Tailored machine learning for evaluating the long-term diabetes risk in older individuals: Findings from The Irish Longitudinal Study on Ageing (TILDA)&quot; that published in BMJ Open.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

A large expert-curated cryo-EM image dataset for machine learning protein particle picking

<p>Cryo-electron microscopy (cryo-EM) is a powerful technique for determining the structures of biological macromolecular complexes. Picking single-protein particles from cryo-EM micrographs is a crucial step in reconstructing protein structures. However, the widely used template-based particle picking process is labor-intensive and time-consuming. Though machine learning and artificial intelligence (AI) based particle picking can potentially automate the process, its development is hindered by lack of large, high-quality labelled training data. To address this bottleneck, we present CryoPPP, a large, diverse, expert-curated cryo-EM image dataset for protein particle picking and analysis. It consists of labelled cryo-EM micrographs (images) of 34 representative protein datasets selected from the Electron Microscopy Public Image Archive (EMPIAR). The dataset is 2.6 terabytes and includes 9,893 high-resolution micrographs with labelled protein particle coordinates. The labelling process was rigorously validated through 2D particle class validation and 3D density map validation with the gold standard. The dataset is expected to greatly facilitate the development of both AI and classical methods for automated cryo-EM protein particle picking.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Supplemetary Data for the article: Machine-learning identified molecular fragments responsible for infrared emission features of polycyclic aromatic hydrocarbons

<p>This is a set of&nbsp;Supplementary materials for&nbsp;the article &#39;Machine-learning identified molecular fragments responsible for infrared emission features of polycyclic aromatic hydrocarbons&#39;, by Meng et al.</p> <p>Supplementary_Data_I.pdf&nbsp;contains an extensive table spanning 36 pages that lists the top-10 molecular fragments accountable for the spectral bands between 2.761 and 1172.745 &mu;m. To access this table, hyperlinks within the document can be used for navigation.</p> <p>Supplementary_Data_II.pdf comprises a large table that encompasses 10,691 pages, including the top-100 molecular fragments responsible for the spectral bands between 2.761 and 1172.745 &mu;m. Navigation through the hyperlinks enables access to this table.</p> <p>Supplementary_Data_III.csv&nbsp;encompasses the chemical formulas, number of unpaired valence electrons, spin multiplicities, xyz data, and SMILES strings of the PAHs carrying the additional spectra.</p> <p>Supplementary_data_IV.zip&nbsp;includes the input and output datasets along with the ML code. The code script is written in Python 3.7, and is supported by the following libraries: sklearn, json, numpy, and pandas.</p> <p>Supplementary_Information.pdf contains the evidence supporting the choice of the cutoff radius, as well as the figures of the count of the molecules in the dataset, the FI with changing datasets and hyperparameters, of cross-validation, and of UIE bands and emission features of four SH PAHs. Importance of three&nbsp;carbon skeleton fragments for&nbsp;bands in different&nbsp;intervals is also demonstrated.</p>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record