Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
A Multi-Omics Interpretable Machine Learning Model Reveals Modes of Action of Small Molecules (ChIP-Seq)
GEO Series GSE129141. Mus musculus. 6 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Machine Learning Identifies Candidates for Drug Repurposing in Alzheimer's Disease
GEO Series GSE164788. Homo sapiens. 764 samples. Type: Expression profiling by high throughput sequencing.
Quantifying Biopolymer Sequence Recognition using Biophysically Informed Machine Learning
GEO Series GSE175942. Drosophila melanogaster; Homo sapiens. 43 samples. Type: Other.
Machine learning unveils an immune-related DNA methylation profile in germline DNA from breast cancer patients
GEO Series GSE243529. Homo sapiens. 524 samples. Type: Methylation profiling by genome tiling array.
Using prognostic signatures and machine learning to identify core features associated with response to CDK4/6 inhibitor-based therapy in metastatic breast cancer
GEO Series GSE285861. Homo sapiens. 168 samples. Type: Expression profiling by high throughput sequencing.
Machine-learning Directed Conversion of Glioblastoma to Dendritic Cell-like Antigen Presenting Cells as Cancer Immunotherapy
GEO Series GSE270855. Mus musculus. 3 samples. Type: Expression profiling by high throughput sequencing.
Machine-learning Classification Identifies Early Systemic Sclerosis Patients that Improve with Abatacept treatment by Modulating CD28-Pathways
GEO Series GSE217067. Homo sapiens. 233 samples. Type: Expression profiling by high throughput sequencing.
Machine Learning Reveals Common Transcriptomic Signature Across Brain and Placenta Following Developmental Organophosphate Ester Exposure
GEO Series GSE230516. Rattus norvegicus. 24 samples. Type: Expression profiling by high throughput sequencing.
Machine learning sequence prioritization for cell type-specific enhancer design
GEO Series GSE171549. Mus musculus. 44 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Establishment of interpretable cytotoxicity prediction models using machine learning analysis of transcriptome features
GEO Series GSE252529. Homo sapiens. 15 samples. Type: Expression profiling by high throughput sequencing.
Distinguishing reproductive phasiRNAs in grasses from other identically-sized small RNAs using machine learning methods
GEO Series GSE108105. Zea mays. 4 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Accurate age prediction from blood using a small set of DNA methylation sites and a cohort-based machine learning algorithm
GEO Series GSE207605. Homo sapiens. 0 samples. Type: Methylation profiling by array; Third-party reanalysis.
Design of an unbiased machine learning workflow to predict Multiple Sclerosis staging from blood transcriptome
GEO Series GSE136411. Homo sapiens. 336 samples. Type: Expression profiling by array.
Machine learning for discovery: deciphering RNA splicing logic
GEO Series GSE200096. Homo sapiens. 4 samples. Type: Other.
Wikidata Dump Machine Learning
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/110">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Characterization of descriptors in machine learning for data-based sputtering yield prediction
<p>descriptor and target variables</p>
Artifact for Counterfactual Explanations for Machine Learning on Multivariate HPC Time Series Data
<p>This includes the data sets used in the SC'20 submission "Counterfactual Explanations for Machine Learning on Multivariate HPC Time Series Data".</p>
Machine learning for predicting the average length of vertically aligned TiO2 nanotubes. Dataset Description
<p>Machine learning for predicting the average length of vertically aligned TiO2 nanotubes. Dataset Description. Dataset is found in two versions: MATLAB mat file and CSV format. Rows contains samples and Columns Features. Last Column represents the Response variable.</p>
Results of experiments in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks
<p>Results of the experiments in paper Simulated Hyperparameter Optimization for Statistical Tests in Machine Learning Benchmarks</p>
Beware of the generic machine learning-based scoring functions in structure-based virtual screening
<p>Data sets and the rescoing scores utilized in the paper "Beware of the generic machine learning-based scoring functions in structure-based virtual screening" (DOI: 10.1093/bib/bbaa070)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.