Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Mapping Phyllosilicates on the Asteroid Bennu Using Thermal Emission Spectra and Machine Learning Model Applications
<p>We provide the laboratory spectra and metadata that was used to construct the PLS model coefficients. Through the application of the model, we provide the prediction values in volume% for Mg-rich serpentine, cronstedtite, and saponite for the BBD1, EQ3, and TAG datasets.</p>
Comparing Storm Resolving Models and Climates via Unsupervised Machine Learning
<p>Storm-resolving climate models (SRMs) have gained international interest for their unprecedented detail with which they globally resolve convection. However, this high resolution also makes it difficult to quantify the emergent differences or similarities among complex atmospheric formations induced by different parameterizations of sub-grid information. This paper uses modern unsupervised machine learning methods to analyze and intercompare SRMs based on their high-dimensional simulation data, learning low-dimensional latent representations and assessing model similarities based on these representations. To quantify such inter-SRM ``distribution shifts'', we use variational autoencoders in conjunction with vector quantization. Our analysis involving nine different global SRMs reveals that only six of them are aligned in their representation of atmospheric dynamics. Our analysis furthermore reveals regional and planetary signatures of the convective response to global warming in a fully unsupervised, data-driven way. In particular, this approach can help elucidate the effects of climate change on rare convection types, such as ``Green Cumuli''. Our study provides a path toward evaluating future high-resolution global climate simulation data more objectively and with less human intervention than has historically been needed.</p>
Raw NGS Data for "Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins"
<p>This directory contains relevant fastq files used for deep sequencing analysis in the publication “Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins”. </p> <p>Fastq files are provided for presorted, uninduced and induced populations from DMS experiments of four homologs (TtgR, TetR, RolR, and MphR). Three replicates were performed for each sample.</p> <p>Data analysis of this deep sequencing data was performed using custom scripts, which are described in the methods section of the publication.</p>
Artifact for ESEC/FSE Paper: "23 Shades of Self-Admitted Technical Debt: An Empirical Study on Machine Learning Software"
<p>Artifact for ESEC/FSE paper entitled "23 Shades of Self-Admitted Technical Debt: An Empirical Study on Machine Learning Software"</p>
Observing flow of He II with unsupervised machine learning
<p>Data repository for observing flow in He II with unsupervised machine learning.</p>
A fast machine-learning-guided primer design pipeline for selective whole genome amplification
<p>Addressing many of the major outstanding questions in the fields of microbial evolution and pathogenesis will require analyses of populations of microbial genomes. Although population genomic studies provide the analytical resolution to investigate evolutionary and mechanistic processes at fine spatial and temporal scales – precisely the scales at which these processes occur – microbial population genomic research is currently hindered by the practicalities of obtaining sufficient quantities of the relatively pure microbial genomic DNA necessary for next-generation sequencing. Here we present swga2.0, an optimized and parallelized pipeline to design selective whole genome amplification (SWGA) primer sets. Unlike previous methods, swga2.0 incorporates active and machine learning methods to evaluate the amplification efficacy of individual primers and primer sets. Additionally, swga2.0 optimizes primer set search and evaluates strategies, including parallelization at each stage of the pipeline, to dramatically decrease program runtime from weeks to minutes. Here we describe the swga2.0 pipeline, including the empirical data used to identify primer and primer set characteristics, that improve amplification performance. Additionally, we evaluated the novel swga2.0 pipeline by designing primers sets that successfully amplify <em>Prevotella melaninogenica</em>, an important component of the lung microbiome in cystic fibrosis patients, from samples dominated by human DNA.</p>
Machine Learning the Hohenberg-Kohn Map to Molecular Excited States
<p>The dataset and code used in paper”Machine Learning the Hohenberg-Kohn Map to Molecular Excited States” For detailed information of each file, see Readme.txt</p>
Exploring Supernova Gravitational Waves with Machine Learning
<p>The gravitational wave strain from сore-collapse supernova simulations used in our analysis. The file contains 402 signals labeled as <em>s00A0O00</em> or s00A0O00.0, where:</p> <p><em>s00</em> -- corresponds to (zero-age) progenitor mass, e.g. s27 means 27 solar mass</p> <p><em>A0</em> -- corresponds for a degree of differential rotation</p> <p><em>O00 </em>or <em>O00.0 </em>-- corresponds to central angular velocity, e.g. O07 or O07.5 means that our model has a central angular velocity of 7 or 7.5 rad/s, respectively</p> <p>Our waveforms are represented as a quadrupole wave amplitude. One can get a strain <em>h</em> multiplied by the distance <em>D </em>(= 10 kpc) by the following formula: <em>hD</em> = <strong><em>our_data</em></strong>/3.66 cm; see Eq (20) of Dimmelmeier et al 2008 [<a href="https://journals.aps.org/prd/abstract/10.1103/PhysRevD.78.064056">link</a>] for more information. All waveforms are represented in the time range from -15 to 20 ms with a 0.001 ms step size. The time of zero corresponds to the time of bounce. See [<a href="https://arxiv.org/abs/2209.14542">https://arxiv.org/abs/2209.14542</a>] for more information.</p>
Thermal conductivity of hydrous wadsleyite determined by non-equilibrium molecular dynamics based on machine learning
<p>This repository contains data used in "Thermal conductivity of hydrous wadsleyite determined by non-equilibrium molecular dynamics based on machine learning" submitted by Dong Wang, Zhongqing Wu and Xin Deng.</p> <p>Figure S4 : "MLP test-Energy" in <strong><a href="https://zenodo.org/api/files/353c5b70-3e53-4192-92af-7bf7f12328fd/Data%20for%20Figures.xlsx?versionId=96ac296c-13f9-4871-976a-d7ad087e0b62">Data for Figures.xlsx</a></strong>、<strong><a href="https://zenodo.org/api/files/353c5b70-3e53-4192-92af-7bf7f12328fd/MLP%20test-force.txt?versionId=acfeccd1-ecbf-42c8-a6a3-0d6daf72b078">MLP test-force.txt</a></strong></p> <p>Figure 1 : "MLP test-NEMD" in <strong><a href="https://zenodo.org/api/files/353c5b70-3e53-4192-92af-7bf7f12328fd/Data%20for%20Figures.xlsx?versionId=96ac296c-13f9-4871-976a-d7ad087e0b62">Data for Figures.xlsx</a></strong></p> <p>Figure 3 : "Modeing" in <strong><a href="https://zenodo.org/api/files/353c5b70-3e53-4192-92af-7bf7f12328fd/Data%20for%20Figures.xlsx?versionId=96ac296c-13f9-4871-976a-d7ad087e0b62">Data for Figures.xlsx</a></strong></p>
EXPLORE Machine Learning Lunar Data Challenges 2022 - QGIS project
<p>This dataset contains the the EXPLORE Machine Learning Data Challenge 2022 QGIS project.</p> <p>The project embed the following Archytas Dome layers:</p> <p><strong>Raster</strong></p> <ul> <li>Narrow Angle Camera (NAC)</li> <li>DEM derived from NAC</li> <li>Slope computer on DEM</li> </ul> <p><strong>Vectorial</strong></p> <ul> <li>POIs - Points Of Interest to be used in STEP 3 </li> </ul> <p> </p> <p> </p> <p> </p> <p>More information at: https://exploredatachallenges.space/</p> <p> </p> <p>Images were processed from NASA PDS raw data using USGS ISIS and NASA ASP tools.</p>
Geospatial data used in "Estimation of river water surface elevation using UAV photogrammetry and machine learning"
<p>Geospatial data used in article "Estimation of river water surface elevation using UAV photogrammetry and machine learning" by Radosław Szostak, Marcin Pietroń, Przemysław Wachniew, Mirosław Zimnoch and Paweł Ćwiąkała (AGH UST).</p> <p>Each zip archive contains the following files:</p> <ul> <li>dsm.tif - raster of digital surface model,</li> <li>ortho.tif - raster of orthophoto,</li> <li>gnss_wse.json - geojson multipoint shape containing RTN GNSS measurements of water surface elevation,</li> <li>grid.json - geojson multipolygon shape containing square areas of samples used in deep learning solution.</li> <li>centerline.json - geojson multipoint shape containing values sampled from DSM along centerline,</li> <li>wateredge.json - geojson multipoint shape containing values sampled from DSM along "water-edge".</li> </ul> <p>Data in AMO18.zip archive was collected by Bandini et. al (https://doi.org/10.5281/zenodo.3519888).</p> <p>Preprocessed machine learning dataset and source codes are available in github repository at: https://github.com/radekszostak/river-wse-uav-ml</p>
Decoding diabetes biomarkers and related molecular mechanisms using machine learning, text mining, and gene expression analysis
<p>The molecular basis of diabetes mellitus is yet to be fully elucidated. We aimed to identify the most frequently reported and differential expressed genes (DEGs) in diabetes using bioinformatics approaches. Text mining was used to screen 40,225 article abstracts from diabetes literature. These studies highlighted 5939 diabetes-related genes spread across 22 human chromosomes, with 112 genes mentioned in more than 50 studies. Among these genes, HNF4A, PPARA, VEGFA, TCF7L2, HLA- DRB1, PPARG, NOS3, KCNJ11, PRKAA2, and HNF1A were mentioned in more than 200 articles. These genes are correlated with the regulation of glycogen and polysaccharide, adipogenesis, AGE/RAGE, and macrophage differentiation. Three datasets (44 patients and 57 controls) were subjected to gene expression analysis. The analysis revealed 135 significant DEGs, of which CEACAM6, ENPP4, HDAC5, HPCAL1, PARVG, STYXL1, VPS28, ZBTB33, ZFP37 and CCDC58 were the top ten DEGs. These genes were enriched in aerobic respiration, T-Cell antigen receptor pathway, Tricarboxylic acid metabolic process, vitamin D receptor pathway, Toll-like receptor signaling, and endoplasmic reticulum (ER) unfolded protein response. The results of text mining and gene expression analyses used as attribute values for ML analysis . The "Decision tree", "Extra-tree regressor" and "Random forest" algorithms were used in ML analysis to identify unique markers that could be used as diabetes diagnosis tools. These algorithms produced prediction models with accuracy ranges from 0.6364 to 0.88 and overall confidence interval (CI) of 95%. There were 39 biomarkers that could distinguish diabetic and non-diabetic patients, 12 of which were repeated multiple times. The majority of these genes are associated with stress response, signalling regulation, locomotion, cell motility, growth, and muscle adaptation. ML algorithms highlighted the use of the HLA-DQB1 gene as a biomarker for diabetes early detection. Our data mining and gene expression analysis have provided useful information about potential biomarkers in diabetes.</p>
Preprocessed Data for "Comparing Storm Resolving Models and Climates via Unsupervised Machine Learning"
<p>Preprocessed Data (training and test) for 3 SRMs used in "Comparing Storm Resolving Models and Climates via Unsupervised Machine Learning". Here we included ICON, SPCAM, and SPCAM with sea surface tem[eratures warmed by +4K. Additionally we include lat/lon information for the test data.</p>
Two-dimensional Energy Histograms as Features for Machine Learning to Predict Adsorption in Diverse Nanoporous Materials
<p>This repo contains the supplementary data sets for the to-be-published paper entitled "Two-dimensional Energy Histograms as Features for Machine Learning to Predict Adsorption in Diverse Nanoporous Materials".</p> <p> </p> <p>This repo contains the following data sets:</p> <p>1. CIF files for amorphous porous materials (activated carbon, hyper-cross-linked polymers, Kerogen, PIMs).</p> <p>2. Grand canonical Monte Carlo (GCMC) simulation results for single-component adsorption isotherms in ToBaCCo1.0 MOFs and in amorphous porous materials. Gas molecules include Kr, Xe, ethane, propane, butane, n-hexane, and 2,2-dimethylbutane.</p> <p>3. Textural properties of ToBaCCo1.0 MOFs and amorphous porous materials.</p> <p>4. Trained machine learning models. R code that can work with these ML models is hosted on <a href="https://github.com/snurr-group/2D-energy-histogram">GitHub</a>. </p>
Dataset: Quantifying cell densities and biovolumes of phytoplankton communities and functional groups using scanning flow cytometry, machine learning and unsupervised clustering
<p>This dataset contains all relevant data for the manuscript (in submission) "<em>Quantifying cell densities and biovolumes of phytoplankton communities and functional groups using scanning flow cytometry, machine learning and unsupervised clustering</em>".</p> <p>Code written to analyse this dataset (which may be adapted for other flow cytometry datasets) is found at https://zenodo.org/record/999747</p> <p>--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>Naming convention for raw flow cytometry data files (located in /Script 3. Generating raw data subset/input/):</p> <p>[Allparameters] _ [Year] - [Month] - [Date] [Hour] [u] [Minute] _ [Depth]</p> <p>e.g: Allparameters_2014-07-31 08u08_1.0m</p> <p>The date, time and depth indicate the location and time at which the measurement was taken.</p>
Data for Developing Machine Learning Models to Predict Base Resistance of Pile Foundation
<p>This data was collected from 86 static pile load tests across 37 different high-rise buildings in Vietnam, especially soft soil region in Mekong Delta (Ho Chi Minh City). The data was used to develop machine learning models to predict base resistance of piles. Further details can be found in publication: "<strong>Influence of Settlement on Base Resistance of Long Piles in Soft Soil—Field and Machine Learning Assessments</strong>", Link: https://www.mdpi.com/2673-7094/4/2/25.</p> <p>Recommended citation: Nguyen, Thanh T., Viet D. Le, Thien Q. Huynh, and Nhu H.T. Nguyen. 2024. "Influence of Settlement on Base Resistance of Long Piles in Soft Soil—Field and Machine Learning Assessments" <em>Geotechnics</em> 4, no. 2: 447-469. https://doi.org/10.3390/geotechnics4020025</p> <p> </p>
SDO 2H Machine Learning Dataset
<p>This dataset provides a compact Machine Learning ready dataset of SDO EUV and HMI medium-resolution (1024x1024 pixels) images, for a total of 56,664 samples from May 14, 2010, to April 18, 2023, with a temporal cadence of 2 hours.</p> <p>EUV images are provided at the following wavelength : 1600A, 304A, 211A, 193A, 171A and 94A.<br>They are processed from the level 1.5 AIA-synoptic dataset (http://jsoc.stanford.edu/data/aia/synoptic/) and are successively:</p> <ul> <li>corrected for instrument degradation</li> <li>normalised by exposure time</li> <li>log-transformed (x->log(1+x)), symetrically on positive and negative values</li> <li>saturated to the 99.9 percentile maximum pixel value of the dataset, up to 2020*, for each channel</li> <li>linearly scaled between 0 and 255, converted to 8bit integers and compressed as jpegs</li> </ul> <p>The HMI's line-of-sight magnetograms (blos.zip) are retrieved from JSOC from the level 1.5 45-second line-of-sight serie and are successively :</p> <ul> <li>downscaled to 1024x1024 pixels</li> <li>standardized to a 2.4 arcec-to-pixel resolution (equal to the EUV images)</li> <li>aligned with the EUV images</li> <li>log-transformed (x->log(1+x))</li> <li>saturated to the 99.9 percentile maximum pixel value of the dataset, up to 2020* </li> <li>linearly scaled between 0 and 255, with 127 representig original null values, 0 and 255 respecivelly the negative and positive saturation value (approximately 4644G before log-transformation)</li> <li>converted to 8bit integers and compressed as jpegs</li> </ul> <p>Downscaled and cropped images (224x448 pixels) used in <a title="Francisco et al., 2023" href="https://doi.org/10.22541/essoar.170688972.24631782/v1" target="_blank" rel="noopener">Francisco et al., 2023</a> are aso provided in pcnn_images.zip</p> <p>An outlier study is also provided in anomalies.zip, from which '{wavelength}_anomalies_grades.csv' files can be used to exclude the dates where abnormal samples of a given type ('anomalies_grade_scale.txt') are identified.</p> <p>*The percentile values are computed on the pixels joint distribution using all sample from 2010 to 2019-12 included, so that the period starting from 2020-01 can be used as a completely independant test set.<br>Original exposure and instrument degradation corrected values can be retrieved using the saturation values provided belows. Althought the JPEG encoding results in the loss of small scale information, the dataset processing preserve the physical intensity of the original inputs so that the provided compressed images can efficiently be used to estimate large and medium-scale Active Regions physical features.</p> <table> <tbody> <tr> <td>HMI / BLOS</td> <td>±4644 G</td> </tr> <tr> <td>1600</td> <td>9,360 DN</td> </tr> <tr> <td>304</td> <td>44,488 DN</td> </tr> <tr> <td>171</td> <td>29,599 DN</td> </tr> <tr> <td>193</td> <td>81,139 DN</td> </tr> <tr> <td>211</td> <td>8,179 DN</td> </tr> <tr> <td>94</td> <td>6,099 DN</td> </tr> </tbody> </table> <p> </p>
Exome sequence analysis identifies rare coding variants associated with a machine learning-based marker for coronary artery disease.
<p>*.sh and *.R are codes to test rare coding variants for association with ISCAD.</p> <p>Petrazzini_etal_2024_*_level_meta_analysis.txt.gz are summary statistics of variant- and gene-level associations of rare coding variants in the exome sequences of 604,914 individuals with an in-silico score for coronary artery disease (ISCAD).</p> <p>Chromosomal positions are mapped to the GRCh38 (hg38) human genome reference.</p> <p>Directions of effect correspond to associations in the UK Biobank, the All of Us Research Program, the BioMe Biobank sample 1 and the BioMe Biobank sample 2, in that order.</p>
Dataset de entrenamiento para un Modelo tipológico de juguete para clasificar fragmentos cerámicos basado en Machine Learning
<p>Este dataset contiene información sobre 577 fragmentos de cerámica arqueológica de Colombia usada para entrenar un modelo tipológico de juguete basado en aprendizaje de máquinas</p>
Machine Learning for O2 Project (ML4O2)
<p>Gridded maps of dissolved oxygen based on machine learning algorithms. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.