Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,666

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,666 results for “human genome”

Learn how ShareScore rates datasets ↗
zenodo32/100

Pre-computed MGCs from human microbiome reference genomes

<p>This dataset contains non-redundant metabolic gene clusters (MGCs) collected by running gutSMASH and BiG-MAP on a collection of unique high-quality reference genomes. This collection consist of MGCs predicted by gutSMASH using&nbsp;1,520 genomes from the Culturable Genome Reference (CGR), 2,308 genomes from the&nbsp;Human Microbiome Project (HMP)&nbsp;and 414 Clostridia&nbsp;genomes as input and&nbsp;then filtered for redundancy using the family module of BiG-MAP. For more information:&nbsp;<a href="http://doi.org/10.1101/2021.02.25.432841">https://doi.org/10.1101/2021.02.25.432841</a></p> <p><strong>BiG-MAP_mg.pickle</strong> -&gt; suitable for <strong>metagenome</strong> analyses</p> <p><strong>BiG-MAP_mt.pickle </strong>-&gt; suitable for <strong>metatranscriptome </strong>analyses</p> <p>The files can be used as direct input in the third module of&nbsp;BiG-MAP&nbsp;(BiG-MAP.map.py:&nbsp;<a href="https://github.com/medema-group/BiG-MAP">https://github.com/medema-group/BiG-MAP</a>).</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Annotation for human genomic variation during the BMP4-induced conversion from embryonic stem cells to trophoblast by bone

<p>Whole genomic data of three cell lines were sequenced.&nbsp;These three cell lines are two hESC lines (H1 &amp; H9) by invasion assays and iPSC cell line MRucR. &nbsp;These three cell lines are two hESC lines (H1 &amp; H9) by invasion assays and iPSC cell line MRucR. Paired-end DNA libraries were prepared according to manufacturer&rsquo;s instructions (Illumina Truseq Library Construction). And then, valid sequencing data is mapped to the reference genome (UCSC hg19) by Burrows-Wheeler Aligner (BWA) software. Reads that aligned to genomic regions were collected for mutation identification and subsequent analysis. Samtools mpileup and bcftools are used to do variant calling and identify SNP, indels. Control-free (Boeva V et al.2012) is utilized to do CNV detection. And BreakDancer (Chen K et al.2009) is applied to detect SV information. And, Single nucleotide variant (SNV) and small somatic insertions and deletions (indels) were identified using Strelka2.</p>

opencc-by-4.0Aug 2019View details →
zenodo32/100

Whole genome RNA-sequencing reveals modulation of genes related to brain disorders by Withania somnifera in human neuroblastoma SK-N-SH cells

<p>Table S1: Human reference genome based differential gene expression; Figure S1: Reactome Pathway (50 &mu;g/mL_3h vs C_3h); Figure S2: Reactome Pathway (50 &mu;g/mL_9h vs C_9h); Figure S3: Reactome pathway (100 &mu;g/mL_3h vs C_3h); Figure S4: GO Dose comparison; Reactome pathway (100 &mu;g/mL_3h vs 50 &mu;g/mL_3h); Figure S5: GO Dose comparison; Reactome pathway (100 &mu;g/mL_9h vs 50 &mu;g/mL_9h); Figure S6: GO Time comparison; Reactome pathway (50 &mu;g/mL_9h vs 50 &mu;g/mL_3h); Figure S7: GO Time comparison; Reactome pathway (100 &mu;g/mL_9h vs 100 &mu;g/mL_3h); Table S2: Disease ontology analysis of 100 &mu;g/mL_3h vs 50 &mu;g/mL_3h WS-treated SK-N-SH cells; Table S3: Disease ontology analysis of 100 &mu;g/mL_9h vs 50 &mu;g/mL_9h WS-treated SK-N-SH cells; Table S4: Disease ontology analysis of 100 &mu;g/mL_9h vs 100 &mu;g/mL_3h WS-treated SK-N-SH cells.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

CAMI2 Challenge - Human Microbiome Project Toy Database - sample 19 - regenerated using recent RefSeq representative genomes

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Human genome dataset (TargetCall-D1.2)

<p>Continuation of dataset&nbsp;Human genome dataset (TargetCall-D1.1) with DOI accession&nbsp;10.5281/zenodo.7334648&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Human genome dataset (TargetCall-D1.1)

<p>Detailed description:</p> <p>The fast5 files in this dataset is generated from ONT machine.</p> <p>This dataset includes the 196K fast5 files from the sequencing run #2 (sampled&nbsp;from&nbsp;<strong>GIAB_GM24385_UB_Run2</strong>)&nbsp;</p> <p>This dataset contains only 98% of the dataset used in the paper. Remaining fast5 files can be downloaded from&nbsp;10.5281/zenodo.7402342</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Genome graphs detect human polymorphisms in active epigenomic states during influenza infection: validation

<p>Sanger sequencing and qPCR validation data for &quot;Genome graphs detect human polymorphisms in active epigenomic states during influenza infection&quot; manuscript.</p>

opencc-by-4.0Dec 2022View details →
dryad32/100

Data from: Distinctive microbial community and genome structure in coastal seawater from a human-made port and nearby offshore island in northern Taiwan facing the Northwestern Pacific Ocean

<p><span>Pollution in human-made fishing ports caused by petroleum </span><span>from</span><span> boats, dead fish, toxic </span><span>chemicals</span><span>, and effluent </span><span>poses</span><span> a challenge to the organisms in seawater. To decipher the impact of pollution on the microbiome, we collected surface water </span><span>from</span><span> a fishing port and a nearby offshore island in northern Taiwan facing the </span><span>Northwestern Pacific Ocean. By employing 16S </span><span>rRNA gene</span><span> amplicon sequencing and whole-genome shotgun sequencing, we discovered that </span><span>Rhodobacteraceae, Vibrionaceae, and Oceanospirillaceae emerged as the dominant species in the fishing port</span><span>,</span><span> where we found many genes harboring the functions of </span><span>antibiotic</span><span> resistance (</span><span>ansamycin, nitroimidazole, and aminocoumarin), metal tolerance (copper, chromium, iron and multimetal), virulence factors (chemotaxis, flagella, T3SS1), carbohydrate metabolism (biofilm formation and remodeling of bacterial cell </span><span>walls</span><span>), nitrogen metabolism (denitrification, N<sub>2</sub> fixation, and ammonium assimilation), and ABC transporters (phosphate, lipopolysaccharide, and branched-chain amino </span><span>acids</span><span>). The dominant bacteria at the nearby offshore island (</span><span>Alteromonadaceae, Cryomorphaceae, Flavobacteriaceae, Litoricolaceae, and Rhodobacteraceae) were partly similar to those in the South China Sea and the East China Sea. Furthermore, we inferred</span><span> that</span><span> the microbial community network of </span><span>the cooccurrence</span><span> of dominant bacteria </span><span>on the</span><span> offshore island was connected to dominant bacteria in </span><span>the </span><span>fishing port by mutual</span> <span>exclusion. By examining the assembled microbial genomes collected from the coastal seawater of the fishing port, we revealed four genomic islands containing large gene-containing sequences</span><span>,</span><span> including phage integrase, DNA</span> <span>invertase, restriction enzyme, DNA gyrase inhibitor, and antitoxin HigA-1.</span><span> In this study, </span><span>we provided </span><span>clues </span><span>for the possibility of genomic islands as the units of horizontal transfer and as the tools of microbes for facilitating adaptation in a human-made port environment.</span></p>

opencc-zeroJan 2023View details →
zenodo32/100

AGC archives of human and SARS-CoV-2 genomes

<p><a href="https://github.com/refresh-bio/agc">AGC</a> is a tool to compress a collection of similar genomes. This Zenodo record provides pre-built AGC-3.0 archives of several datasets:</p> <ul> <li>File &quot;HPRC-yr1.agc&quot; contains <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_009914755.3/">CHM13</a> and 94 haploid human assemblies <a href="https://github.com/human-pangenomics/HPP_Year1_Data_Freeze_v1.0">released by HPRC</a>&nbsp;in 2021. The telomere-to-telomere&nbsp;CHM13 v2&nbsp;plus&nbsp;chrY from GRCh38 is used as the reference genome.</li> <li>File &quot;sars-cov-2_ncbi-620k.agc&quot; contains 619,750 complete SARS-CoV-2 genomes with&nbsp;<a href="https://www.ncbi.nlm.nih.gov/nuccore/NC_045512.2">NC_045512.2</a> as the reference. It was&nbsp;created with AGC command line &quot;agc create -cb10000 -s3000&quot;.&nbsp;SARS-CoV-2 genomes were downloaded <a href="https://www.ncbi.nlm.nih.gov/datasets/coronavirus/genomes/">from NCBI</a> at the end of year 2021. The original FASTA is provided as &quot;sars-cov-2_ncbi-620k.fa.xz&quot;.</li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Inversion polymorphism in a complete human genome assembly

<p>Supplementary data including code for an journal article titled &#39;Inversion polymorphism in a complete human genome assembly&#39;.</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

A genome-wide genetic screen uncovers novel determinants of human pigmentation

<p>1717.gwas.imputed_v3.both_sexes.variants.nolowconf.maf01.clumped.txt - GWAS results for LD-clumped genome-wide significant variants from White British skin color GWAS<br> 1717.gwas.imputed_v3.both_sexes.variants.nolowconf.maf01.pbsfilt.100kb.tsv - GWAS results for LD-pruned variants within 100kb of pro-melanin genes<br> eqtl.gwas.minpergene.txt - melanocyte eQTL and White British skin color GWAS results for SNPs near pro-melanin gene. Columns 1-12 are eQTL info, 13 onwards are GWAS info. Select rows with &quot;closest&quot; = TRUE and both p-values &lt; 0.01 to recreate figure in the paper.<br> csv_and_r_script_files.tar- csv and R script files to recreate the figures in the paper.&nbsp;<br> PBS_Pigmentation.tar- txt and R script files to recreate PBS figures in the Figure&nbsp;<br> &nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Human genome fasta file from ensembl (GRCh38 v109)

<p>Human genome fasta file from ensembl&nbsp; (GRCh38 v109), <span>corresponds to GenBank Assembly ID </span><span>GCA_000001405.28</span></p>

opencc-by-4.0Jun 2023View details →
dryad32/100

Most damaging CADD scores for hg19 human genome build (CADD scores generated with bStatistic removed)

<p>Analyses of genetic variation in many taxa have established that neutral genetic diversity is shaped by natural selection at linked sites. Whether the mode of selection is primarily the fixation of strongly beneficial alleles (selective sweeps) or purifying selection on deleterious mutations (background selection) remains unknown, however. We address this question in humans by fitting a model of the joint effects of selective sweeps and background selection to autosomal polymorphism data from the 1000 Genomes Project. After controlling for variation in mutation rates along the genome, a model of background selection alone explains ~60% of the variance in diversity levels at the megabase scale. Adding the effects of selective sweeps driven by adaptive substitutions to the model does not improve the fit, and when both modes of selection are considered jointly, selective sweeps are estimated to have had little or no effect on linked neutral diversity. The regions under purifying selection are best predicted by phylogenetic conservation, with ~80% of the deleterious mutations affecting neutral diversity occurring in non-exonic regions. Thus, background selection is the dominant mode of linked selection in humans, with marked effects on diversity levels throughout autosomes.</p>

opencc-zeroAug 2023View details →
zenodo32/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p><p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

A compendium of 3562 human and animal papillomavirus genomes

<p>Papillomaviruses (PVs)&nbsp;are a ubiquitous group of DNA viruses that infect a wide range of vertebrate hosts, including human. It is well established that the infection of some PV types (e.g., HPV16 and HPV18) are strongly associated with the occurrence and progression of cancers in human, but it remains elusive how such pathogenicity got evolved at the genomic level.&nbsp;Here we curated and annotated a compendium of 3,562 PV genomes (including 3,329&nbsp;human&nbsp;PVs and 233&nbsp;animal PVs) and performed the corresponding comparative genomics analyses.</p>

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov32/100

Evaluating Genomic Testing in Human Cancer & Outcomes of Targeted Therapies

ClinicalTrials.gov study NCT03089554. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

2000 HIV Human Functional Genomics Partnership Program

ClinicalTrials.gov study NCT03994835. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Associations, Outcomes and Genomics of GB Virus C, Hepatitis C Virus and Human Immunodeficiency Virus Infection

ClinicalTrials.gov study NCT00164060. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

NICUSeq: A Trial to Evaluate the Clinical Utility of Human Whole Genome Sequencing (WGS) Compared to Standard of Care in Acute Care Neonates and Infants

ClinicalTrials.gov study NCT03290469. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Beneficial Effects of Oral Premarin Estrogen Replacement Therapy Assessed by Human Genome Array

ClinicalTrials.gov study NCT00318318. IPD Sharing: Not stated. Countries: 1. Publications: 16.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record