Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Big data-driven target identification by machine learning: DRD2 as a therapeutic target for psoriasis.
GEO Series GSE276511. Mus musculus. 10 samples. Type: Expression profiling by high throughput sequencing.
A robust machine learning approach for missing persons cases with high genotyping errors
GEO Series GSE209804. Homo sapiens. 24 samples. Type: Genome variation profiling by SNP array; SNP genotyping by SNP array.
Multi-omics and machine learning reveal context-specific gene regulatory activities of PML-RARA in Acute Promyelocytic Leukemia [Cut&Tag]
GEO Series GSE209833. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Application of machine learning (ML) / deep learning (DL) using multiple epigenetic features reveals H3K27Ac as driver of gene expression prediction across patients with glioblastoma [ChIP-Seq]
GEO Series GSE296944. Homo sapiens. 7 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
DNA methylation-based machine learning classification distinguishes pleural mesothelioma from chronic pleuritis, pleural carcinosis, and pleomorphic lung carcinomas
GEO Series GSE203061. Homo sapiens. 76 samples. Type: Methylation profiling by genome tiling array.
A Machine-learning approach to define and predict the therapeutic landscape of pan-cancer Hippo pathway dependency (RNA-seq)
GEO Series GSE161010. Homo sapiens. 42 samples. Type: Expression profiling by high throughput sequencing.
A Novel Piperine Derivative Inhibits Colorectal Cancer Progression by Modulating EMT Signaling Pathways: An Integrated Transcriptomic and Machine Learning Analysis
GEO Series GSE275190. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
Identifying transcription factor-DNA interactions using machine learning
GEO Series GSE193400. Glycine max. 13 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Revealing the Grammar of Small RNA Secretion Using Interpretable Machine Learning
GEO Series GSE230012. Homo sapiens. 31 samples. Type: Other; Non-coding RNA profiling by high throughput sequencing.
A Machine-learning approach to define and predict the therapeutic landscape of pan-cancer Hippo pathway dependency
GEO Series GSE161019. Homo sapiens. 50 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Development of an exosomal gene signature to detect residual disease in dogs with osteosarcoma using a novel xenograft platform and machine learning
GEO Series GSE183191. Canis lupus familiaris; Mus musculus. 28 samples. Type: Expression profiling by high throughput sequencing.
Metabolomics and Machine learning approaches for more accurate Paracoccidioidomycosis diagnosis.
<p>High resolution mass spectrometry data from serum samples of Paracoccidioidomycosis patients and healthy volunteers. </p>
Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions"
<p>This is the replication package for the paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package in the following.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py </em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirements.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirements.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx </em>contains 848 labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>
Database for Automatic Risk Tuning in Short-Term Electricity Market Models Using a Machine Learning Proxy
<p>Database (2014-2018) used in the paper entitled "Automatic Risk Tuning in Short-Term Electricity Market Models Using a Machine Learning Proxy".</p> <p>This database includes :</p> <ul> <li>the inputs and outputs of the probabilistic forecaster which predicts the Belgian system imbalance.</li> <li>the market data related to the construction of the balancing market.</li> </ul> <p>These data are obtained from the Belgian transmission system operator (Elia) and the European Network of Transmission system Operators (ENTSO-E).</p> <p>If you use these data, please refer to the following paper:</p> <p>J. Bottieau, K. Bruninx, A. Sanjab, Z. De Grève, F. Vallée and J-F. Toubeau, “Automatic Risk Tuning in Short-Term Electricity Market Models Using a Machine Learning Proxy,”.</p>
Replication Package for the Paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List"
<p>This is the replication package for the paper: "A Machine Learning Based Ensemble Method for Automatic Classification of Decisions: A Study of the Hibernate Developer Mailing List". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. code folder</strong></p> <ul> <li><em>experiment.py </em>contains the source code for our experiment, which is conducted on Windows 10 and Python 3.7.0. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>decisions.xlsx </em>contains 844 labelled sentence-level decisions from the Hibernate developer mailing list.</li> </ul>
Investigating the Combined Impact of Climate Change and Land Use/Land Cover on Flood Vulnerability Using a Machine Learning Algorithm
<p>Using the uploaded code in preparation of the Flood vulnerability maps.</p><p> </p>
Ovarian cancer is detectable from peripheral blood using machine learning over T cell repertoires
<p>To see if TCR repertoire obtained from peripheral blood is associated with tumor status, we collected blood samples from 85 women with or without ovarian cancer and produced T cell receptor information. Using machine learning, we stratified, with high success rates, the two groups, thereby associating peripheral blood T cell repertoire with the existence of ovarian cancer tumors.</p>
Prediction and interpretation microglia cytotoxicity by machine learning
<p>Data sets for microglia cytotoxicity. </p>
Data-Driven Discovery of Carbonyl Organic Electrode Molecules: Machine Learning and Experiment
Open the record for dataset details and reuse information.
Data for "Battery cycle life study through relaxation and forecasting the lifetime via machine learning"
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.