Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,666

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,666 results for “human genome”

Learn how ShareScore rates datasets ↗
ClinicalTrials.gov32/100

Human Papillomavirus Association and Genomic Exploration in Head-Neck Squamous Cell Carcinomas

ClinicalTrials.gov study NCT06642324. IPD Sharing: NO. Countries: 1. Publications: 4.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Human Genomic Population Structure and Phenotype-genotype Variation in ADME Genes in Four Populations

ClinicalTrials.gov study NCT02789527. IPD Sharing: UNDECIDED. Countries: 4. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Data from: Distinctive microbial community and genome structure in coastal seawater from a human-made port and nearby offshore island in northern Taiwan facing the Northwestern Pacific Ocean

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad32/100

Data from: Comparative landscape genomics reveals species-specific spatial patterns and suggests human-aided dispersal in a global hotspot for biological invasions

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad32/100

Data from: Differential requirements for the RAD51 paralogs in genome repair and maintenance in human cells

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad32/100

Data from: Demographic inference from whole-genome and RAD sequencing data suggests alternating human impacts on goose populations since the last ice age

Open the record for dataset details and reuse information.

publicSep 2017View details →
dryad32/100

Data from: Signatures of human-commensalism in the house sparrow genome

Open the record for dataset details and reuse information.

publicJul 2018View details →
dryad32/100

Most damaging CADD scores for hg19 human genome build (CADD scores generated with bStatistic removed)

Open the record for dataset details and reuse information.

publicAug 2023View details →
zenodo28/100

YAMP Resources with Human Genome indexed bbmap 38.76

<p>Resource dataset for <strong>YAMP</strong> (https://github.com/alesssia/YAMP) with Human Genome indexed using <strong>&nbsp;bbmap 38.76</strong></p> <p><em>Original dataset:</em></p> <p>Alessia. (2017). Data for YAMP (https://github.com/alesssia/YAMP) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1068229</p> <p><em>From the original dataset :</em></p> <p>This dataset includes all the databases/datasets queried by YAMP (https://github.com/alesssia/YAMP). This&nbsp;dataset has been created to help the&nbsp;users to get started with YAMP, and&nbsp;to save them from the hassle of collecting and downloading the data from different sources.</p> <p>More in details, this dataset contains:</p> <ul> <li>a FASTA file listing the adapter sequences to remove in the trimming step. This file is usually&nbsp;provided within the&nbsp;BBmap installation (https://sourceforge.net/projects/bbmap, version 37.68).&nbsp;</li> <li>two FASTA files describing synthetic contaminants (sequencing_artifacts.fa.gz&nbsp;and&nbsp;phix174_ill.ref.fa.gz). These files are usually provided within the&nbsp;BBmap installation (https://sourceforge.net/projects/bbmap, version 37.68).&nbsp;</li> <li>a FASTA file provided by Brian Bushnell for removing human contamination (described here:&nbsp;http://seqanswers.com/forums/showthread.php?t=42552).&nbsp;&nbsp;Please note that this file should be indexed beforehand. This can be done using BBMap, using the following command: `bbmap.sh -Xmx24G ref=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz`.&nbsp;</li> <li>the BowTie2 database file for MetaPhlAn2. This file is usually provided within the MetaPhlAn2 installation (version 2.6.0)</li> <li>the ChocoPhlAn and UniRef (Uniref50, Uniref90)&nbsp;databases, downloaded directly by HUMAnN2 (version 0.9.9), as explained here:&nbsp;https://bitbucket.org/biobakery/humann2/wiki/Home#markdown-header-5-download-the-databases</li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo28/100

Eigen scores for human genome assembly GRCh37 Part 3 (Chr9 - Chr16)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

Eigen scores for human genome assembly GRCh37 Part 2 (Chr4 - Chr8)

<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

Genomic insights into the formation of human populations in East Asia

<p>The genotypes of&nbsp;383 present-day individuals from 46 populations indigenous to China (n=337) and Nepal (n=46) using the Affymetrix Human Origins array.</p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

QDNAseq.hg38: QDNAseq bin annotation for the human genome build hg38

<p><strong>QDNAseq</strong> bin annotations of size 1, 5, 10, 15, 30, 50, 100, 500, and 1000 kbp for the human genome build hg38.</p>

openother-openNov 2020View details →
zenodo28/100

Human genome for contaminant removal of metagenome reads

<p>See&nbsp;http://seqanswers.com/forums/archive/index.php/t-42552.html</p>

opencc-by-4.0Jan 2021View details →
dryad28/100

Data from: Comparative genomics, infectivity and cytopathogenicity of Zika viruses produced by acutely and persistently Zika virus-infected human hematopoietic cell lines

Zika virus (ZIKV), an arthropod-borne virus, has emerged as a major human pathogen. Prolonged or persistent ZIKV infection of human cells and tissues may serve as a reservoir for the virus and present serious challenges to the safety of public health. Human hematopoietic cell lines with different developmental properties revealed differences in susceptibility and outcomes to ZIKV infection. In 3 separate studies involving the prototypic MR 766 ZIKV strain and the human monocytic leukemia U937 cell line, ZIKV initially developed only a low-grade infection at a slow rate. After continuous culture for several months, persistently ZIKV-infected cell lines were observed with most, if not all, cells testing positive for ZIKV antigen. The infected cultures produced ZIKV RNA (v-RNA) and infectious ZIKVs persistently ("persistent ZIKVs") with distinct infectivity and pathogenicity when tested using various kinds of host cells. When the genomes of ZIKVs from the three persistently infected cell lines were compared with the genome of the prototypic MR 766 ZIKV strain, distinct sets of mutations specific to each cell line were found. Significantly, all three "persistent ZIKVs" were capable of infecting fresh U937 cells with high efficiency at rapid rates, resulting in the development of a new set of persistently ZIKV-infected U937 cell lines. The genomes of ZIKVs from the new set of persistently ZIKV-infected U937 cell lines were further analyzed for their different mutations. The 2nd generation of persistent ZIKVs continued to possess most of the distinct sets of mutations specific to the respective 1st generation of persistent ZIKVs. We anticipate that the study will contribute to the understanding of the fundamental biology of adaptive mutations and selection during viral persistence. The persistently ZIKV-infected human cell lines that we developed will also be useful to investigate critical molecular pathways of ZIKV persistence and to study drugs or countermeasures against ZIKV infections and transmission.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Rates of genomic divergence in humans, chimpanzees and their lice

The rate of DNA mutation and divergence is highly variable across the tree of life. However, the reasons underlying this variation are not well understood. Comparing the rates of genetic changes between hosts and parasite lineages that diverged at the same time is one way to begin to understand differences in genetic mutation and substitution rates. Such studies have indicated that the rate of genetic divergence in parasites is often faster than that of their hosts when comparing single genes. However, the variation in this relative rate of molecular evolution across different genes in the genome is unknown. We compared the rate of DNA sequence divergence between humans, chimpanzees and their ectoparasitic lice for 1534 protein-coding genes across their genomes. The rate of DNA substitution in these orthologous genes was on average 14 times faster for lice than for humans and chimpanzees. In addition, these rates were positively correlated across genes. Because this correlation only occurred for substitutions that changed the amino acid, this pattern is probably produced by similar functional constraints across the same genes in humans, chimpanzees and their ectoparasites.

opencc-zeroDec 2013View details →
dryad28/100

MtDNA genomes from Ranis individuals aligned with previously published ancient & modern humans

<p>The Middle to Upper Palaeolithic transition in Europe is associated with the regional disappearance of Neanderthals and the spread of Homo sapiens. Archaeological evidence indicates the presence of technocomplexes at the interface of this transition, complicating our understanding of the period and the association of those with specific hominin groups. One such technocomplex where the maker is unknown is the Lincombian-Ranisian-Jerzmanowician (LRJ), which covers an area in northwestern and central Europe from the UK to Poland. This paper presents the morphological and proteomic species identification, mitochondrial DNA analysis, and direct radiocarbon dating of human remains directly associated to an LRJ assemblage at the cave site of Ilsenhöhle in Ranis (Germany).  </p>

opencc-zeroOct 2023View details →
zenodo28/100

10X Genomics Human Visium Spatial Transcriptomics Demo Dataset for Cellxgene VIP

<p>4 Visium Spatial Transcriptomics datasets downloaded 10X Genomics data site ,and organized in the way to be used for Cellxgene VIP input.</p> <p>10X_demo_data_Breast_Cancer_Block_A_Section_1<br> 10X_demo_data_Breast_Cancer_Block_A_Section_2<br> 10X_demo_data_Human_Heart<br> 10X_demo_data_Human_Lymph_Node<br> &nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo28/100

MACIE scores for human genome assembly GRCh37 Part 2 (Chr4 - Chr7)

<p>MACIE (Multi-dimensional Annotation Class Integrative Estimation) is an unsupervised multivariate mixed model framework to assess multi-dimensional functional impacts for both coding and non-coding variants in the human genome. MACIE integrates a variety of functional annotations, including protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics, and estimates the joint posterior probabilities of each genetic variant being functional.</p> <p>For each non-coding and synonymous coding variant, the MACIE score is a vector of length 4, representing the estimated joint posterior probabilities of &ldquo;not evolutionarily conserved and regulatory functional&rdquo; (MACIE01); &ldquo;evolutionarily conserved and not regulatory functional&rdquo; (MACIE10); &ldquo;not evolutionarily conserved and not regulatory functional&rdquo; (MACIE00); &ldquo;both evolutionarily conserved and regulatory functional (MACIE11). MACIE_conserved is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo;, which is the sum of MACIE10 and MACIE11; MACIE_regulatory is the estimated posterior probability of &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01 and MACIE11; MACIE_anyclass is the estimated posterior probability of &ldquo;evolutionarily conserved&rdquo; or &ldquo;regulatory functional&rdquo;, which is the sum of MACIE01, MACIE10, and MACIE11.</p>

opencc-by-4.0Dec 2021View details →
dryad28/100

Data from: Genome sequences reveal cryptic speciation in the human pathogen Histoplasma capsulatum

Histoplasma capsulatum is a pathogenic fungus that causes life-threatening lung infections. About 500,000 people are exposed to H. capsulatum each year in the United States, and over 60% of the U.S. population has been exposed to the fungus at some point in their life. We performed genome-wide population genetics and phylogenetic analyses with 30 Histoplasma isolates representing four recognized areas where histoplasmosis is endemic and show that the Histoplasma genus is composed of at least four species that are genetically isolated and rarely interbreed. Therefore, we propose a taxonomic rearrangement of the genus. IMPORTANCE: The evolutionary processes that give rise to new pathogen lineages are critical to our understanding of how they adapt to new environments and how frequently they exchange genes with each other. The fungal pathogen Histoplasma capsulatum provides opportunities to precisely test hypotheses about the origin of new genetic variation. We find that H. capsulatum is composed of at least four different cryptic species that differ genetically and also in virulence. These results have implications for the epidemiology of histoplasmosis because not all Histoplasma species are equivalent in their geographic range and ability to cause disease.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record