Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,666

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,666 results for “human genome”

Learn how ShareScore rates datasets ↗
dryad36/100

Neandertal ancestry through time: Insights from genomes of ancient and present-day humans

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad36/100

Genomic analysis reveals close genetic similarity between ESBL-producing E. coli isolates from humans and dogs, suggesting potential for inter-species transmission

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

De novo genome assembly of human cell line CHM13 nanopore ultra-long reads using Shasta

Open the record for dataset details and reuse information.

publicMay 2022View details →
dryad36/100

Supporting data for: Three-dimensional genome re-wiring in loci with Human Accelerated Regions

Open the record for dataset details and reuse information.

publicJan 2023View details →
zenodo32/100

Eigen scores for human genome assembly GRCh38 Part 1 (Chr12 - Chr22)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Dec 2018View details →
zenodo32/100

Eigen scores for human genome assembly GRCh38 Part 2 (Chr6 - Chr11)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Jun 2019View details →
zenodo32/100

Systematic shotgun of human genome [fastq.gz]

<p>Systematic shotgun of hg38 genome for testing the ability to identify human reads.</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

Synthetic Noisy Reads Generated from Human Genomes

<p>This repository contains datasets with size 200K, 400K and 1 Million noisy reads generated from 5000 transcripts of GTcenters.fasta file. The assigned task is to recover the ground-truth transcripts (GTcenters) based on the given noisy reads.</p>

opencc-by-4.0Sep 2020View details →
zenodo32/100

Human genomics of the humoral immune response against polyomaviruses

<p>Meta-analyses&nbsp;results&nbsp;of the manuscript &quot;Human genomics of the humoral immune response against polyomaviruses&quot;.</p>

opencc-by-4.0Nov 2020View details →
zenodo32/100

Data for "Analysis of metagenome-assembled viral genomes from the human gut reveals diverse putative CrAss-like phages with unique genomic features"

<p>Data for &quot;Analysis of metagenome-assembled viral genomes from the human gut reveals diverse putative CrAss-like phages with unique genomic features&quot; (submitted to Nature Communications)</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

Human genome ancestral state files

<p>Files containing ancestral states at all positions of the human genome, inferred based on consensus support among three ape species; Gorilla, Chimpanzee and Orangutan.</p>

opencc-by-4.0Jan 2021View details →
dryad32/100

Data from: Differential requirements for the RAD51 paralogs in genome repair and maintenance in human cells

Deficiency in several of the classical human RAD51 paralogs [RAD51B, RAD51C, RAD51D, XRCC2 and XRCC3] is associated with cancer predisposition and Fanconi anemia. To investigate their functions, isogenic disruption mutants for each were generated in non-transformed MCF10A mammary epithelial cells and in transformed U2OS and HEK293 cells. In U2OS and HEK293 cells, viable ablated clones were readily isolated for each RAD51 paralog; in contrast, with the exception of RAD51B, RAD51 paralogs are cell-essential in MCF10A cells. Underlining their importance for genomic stability, mutant cell lines display variable growth defects, impaired sister chromatid recombination, reduced levels of stable RAD51 nuclear foci, and hyper-sensitivity to mitomycin C and olaparib. Altogether these observations underscore the contributions of RAD51 paralogs in diverse DNA repair processes, and demonstrate essential differences in different cell types. Finally, this study will provide useful reagents to analyze patient-derived mutations and to investigate mechanisms of chemotherapeutic resistance deployed by cancers.

opencc-zeroSep 2019View details →
zenodo32/100

"interactive" version of data associated with the eLife paper "Integrative genomic analysis of the human immune response to influenza vaccination"

This archive contains an "interactive" version of data associated with the eLife paper "Integrative genomic analysis of the human immune response to influenza vaccination" by Luis M Franco, Kristine L Bucasas, Janet M Wells, Diane Niño, Xueqing Wang, Gladys E Zapata, Nancy Arden, Alexander Renwick, Peng Yu, John M Quarles, Molly S Bray, Robert B Couch, John W Belmont, Chad A Shaw http://dx.doi.org/10.7554/eLife.00299 (doi:10.7554/eLife.00299) Installation: (1) Download the rar file, link available from the paper. (2) Unrar the file: Mac OS X: Use UnRarX - http://www.unrarx.com Linux : unrar command - http://en.wikipedia.org/wiki/Unrar (3) Open the file "vaxgenomics.htm" in your favorite browser use File>Open : Once opened in your browser, you should see a "Circos"-style circular plot of the data, and links to the tables and figures used in this application. Each table includes a link-out from GeneID to the Entrez entry for that Gene as well as additional links to the data. The link to Entrez requires access to NCBI, so will only work if you have Internet availability for this to work. All other links point to content contained within the application/archive. (4) Figures: High-resolution copies of the figures included in the paper. Fig 1: eQTL profile of flu vaccination response. Markers associated with cis gene expression identified in the discovery cohort were replicated in a vaidation cohort and -log10 p-values for both data sets are shown in the genome wide circularized graphic. Fig 2: Effect of the treatment on eQTL assocation. We observed that the pattern of association between gene expression and SNP changed over time after vaccination. Panel A of this figure shows this phenomena for a single gene NECAB2. Panel B shows the aggregate character of this phenomena across large numbers of markers. The change in R2 compared to the initial time point is depicted; this change in R squared appears to correspond to an increase the magnitude of the slope (additive association with genotype). Fig 3: Pathways and processes identified as enriched in our candidates. Both Ingenuity IPA Analysis and GO and KEGG pathway databases were used. Heavily implicated immunologic response classes are repsresented. Fig 4: Human immune response as measured by our Antibody Response scores are correlated with gene expression changes, and these patterns are recapituoated in discovery and validation (male/female) cohorts. Fig 5: Immune cellular context of the genes identifeies in our eQTL and immune response analysis. A striking number of our validated genes occuue in the antigen processing and presentatio pnthway. Fig 6: Q-Q plot depicting that the strength of association between genotype and phenotype (titer response) is stronger for markers that have a SNP association with expression and where expression is associated with titer response than would be expected for random SNPs. Fig 7: Causal and Reactive Model Analyses. Three-way association between genotype, expression and trait. Our data are more consistent with a causal relationship compared to reactive, but the results are not definitive. We explored the sample size necessary to investigate this in the supplement. Fig 8: Diagram demonstrating the experimental design of this study. Fig 9: Population structure analysis performed on our cohort using the genome wide SNP data confirms the European ancestry and ethnic homogeoneity of our ty sample Fig 10: Schematic of the eQTL analysis. The time course of gene expression change is integrated with a single model considering effects of Day, Genotype, Day-Genotype interaction and random effects for each individual to account for the longitudinal nature of the design. email: cashaw@bcm.edu with questions or comments.

opencc-zeroJul 2013View details →
zenodo32/100

ChromBERT: Uncovering Chromatin State Motifs in the Human Genome using a BERT-based Approach

<ol> <li>Pretrain data and results for:&nbsp; <ul> <li>Promoter regions for all genes in 127 different cell lines in ROADMAP&nbsp;</li> <li>CRM (<em>cis</em>-regulatory module) regions longer than 2k bps in 127 different cell lines in ROADMAP</li> <li>Whole genome regions without continuous low signal state ("O") in 4-mer</li> </ul> </li> <li>Fine-tuning data and result for: <ul> <li>[Classification] Promoter regions of high-expressed genes (from RPKM&gt;10 to RPKM&gt;50) compared to not expressed genes (RPKM=0) or low-expressed genes (RPKM&gt;0) with the directory names:<br> <ul> <li>not_n_rpkm0 : RPKM=0 vs. RPKM&gt;0</li> <li>not_n_rpkm10 : RPKM=0 vs. RPKM&gt;10</li> <li>not_n_rpkm20 : RPKM=0 vs. RPKM&gt;20&nbsp;</li> <li>not_n_rpkm30 : RPKM=0 vs. RPKM&gt;30&nbsp;</li> <li>not_n_rpkm50 : RPKM=0 vs. RPKM&gt;50&nbsp;</li> <li>rpkm0_n_rpkm10 : RPKM&gt;0 vs. RPKM&gt;10&nbsp;</li> <li>rpkm0_n_rpkm20 : RPKM&gt;0 vs. RPKM&gt;20&nbsp;</li> <li>rpkm0_n_rpkm30 : RPKM&gt;0 vs. RPKM&gt;30&nbsp;</li> <li>rpkm0_n_rpkm50 : RPKM&gt;0 vs. RPKM&gt;50&nbsp;</li> <li>rpkm10_n_rpkm20 : RPKM&gt;10 vs. RPKM&gt;20&nbsp;</li> <li>rpkm10_n_rpkm30 : RPKM&gt;10 vs. RPKM&gt;30&nbsp;</li> <li>rpkm10_n_rpkm50 : RPKM&gt;10 vs. RPKM&gt;50&nbsp;</li> <li>rpkm20_n_rpkm30 : RPKM&gt;20 vs. RPKM&gt;30&nbsp;</li> <li>rpkm20_n_rpkm50 : RPKM&gt;20 vs. RPKM&gt;50&nbsp;</li> <li>rpkm30_n_rpkm50 : RPKM&gt;30 vs. RPKM&gt;50&nbsp;</li> </ul> </li> <li>[Regression] Quantitative gene expression prediction for promoter regions</li> <li>CRM&nbsp;regions longer than 2k bps compared to non-CRM regions</li> </ul> </li> </ol>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Updated Metagenomic Species Pan-genomes (MSPs) of the human gastrointestinal microbiota

<p></p><h1>Gene catalog construction</h1><br>The methodology for creating the IGC2 catalog is described in the original papers: Li et al., 2014 and Wen et al., 2017<br><h1>MSP creation</h1><br>Reads from publicly available human gut metagenomes were aligned against the IGC2 catalog with the Meteor to produce a raw gene abundance table (10.4M genes quantified in &gt;2000 samples). Then, co-abundant genes were binned in 1,989 Metagenomic Species Pan-genomes (MSPs, i.e. clusters of co-abundant genes that likely belong to the same microbial species) using MSPminer.<br><h1>MSPs taxonomic annotation</h1><br>MSPs taxonomic annotation was performed by aligning MSP core and accessory genes against representative genomes of the Genome Taxonomy Database (GTDB r207) using blastn (task = megablast, word_size = 16). The 20 best hits for each gene were kept (--max-target-seq 20). Using an in-house pipeline, a species-level assignment was given if &gt; 50% of the genes matched the representative genome of a given species, with a mean identity ≥ 95% and mean gene length coverage ≥ 90%. The remaining MSPs were assigned to a higher taxonomic level (genus to superkingdom), if more than 50% of their genes had the same annotation.<br><h1>Construction of the phylogenetic tree</h1><br>39 universal phylogenetic markers genes were extracted from the MSPs with fetchMGs. Then, the markers were separately aligned with MUSCLE. The alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni). <h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB1786, PRJEB5224, PRJEB6337, PRJNA422434 (cohort used in catalogue assembly) and PRJEB11532, PRJEB33500, PRJEB37249, PRNJNA834801 (independent cohort not used in assembly).<p></p>

opencc-zeroDec 2020View details →
zenodo32/100

CRISPR-Cas12a-integrated transgenes in genomic safe harbors retain high expression in human hematopoietic iPSC-derived lineages and primary cells

<p>We identified and characterized&nbsp;potential integration safe harbor sites (SHS) in human cells. Using the CRISPR-MAD7 system, we integrated transgenes at these genomic sites in iPSC, primary T and NK cells, and Jurkat cell line, and demonstrated efficient and stable expression at these loci. Subsequently, we validated the differentiation capabilities of engineered iPSC towards CD34+&nbsp;hematopoietic stem and progenitor cells (HSPC), lymphoid progenitors (LPC), and natural killer (NK) cells, and showed that transgene expression was retained in these lineages.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

A genomic compendium of cultivated human gut fungi characterizes the gut mycobiome and its relevance to common diseases

<p><span>Morphological and scanning electron microscopy images of 206 fungal species cultivated from human feces.&nbsp;</span></p>

opencc-by-4.0Dec 2023View details →
dryad32/100

Data from: Demographic inference from whole-genome and RAD sequencing data suggests alternating human impacts on goose populations since the last ice age

We investigated how population changes and fluctuations in the pink-footed goose might have been affected by climatic and anthropogenic factors. First, genomic data confirmed the existence of two separate populations: western (Iceland) and eastern (Svalbard/Denmark). Second, emographic inference suggests that the species survived the last glacial period as a single ancestral population with a low population size (100-1,000 individuals) that split into the current populations at the end of the Last Glacial Maximum with Iceland being the most plausible glacial refuge. While population changes during the last glaciation were clearly environmental, we hypothesize that more recent demographic changes are human-related: (1) the inferred population increase in the Neolithic is due to deforestation to establish new lands for agriculture, increasing available habitat for pink-footed geese (2) the decline inferred during the Middle Ages is due to human persecution and (3) improved protection explains the increasing demographic trends during the 20th century. Our results suggest both environmental (during glacial cycles) and anthropogenic effects (more recent) can be a threat to species survival.

opencc-zeroDec 2016View details →
zenodo32/100

Human genome compressibility

<p>Data for&nbsp;https://bioinformatics.stackexchange.com/questions/12824/theoretical-limit-of-human-genome-compression and&nbsp;https://github.com/karel-brinda/human-genome-compression</p>

opencc-by-4.0May 2021View details →
zenodo32/100

Genome-wide survey of D/E repeats in human proteins uncovers their instability and aids in identification of their role in the chromatin regulator ATAD2

<p>Raw digital images of Western blots that comprise Figure 8E of the manuscript</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record