Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,666

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,666 results for “human genome”

Learn how ShareScore rates datasets ↗
zenodo36/100

Genome graphs detect human polymorphisms in active epigenomic states during influenza infection: code and processed data

<p>Manuscript, figure, and analysis code and processed data for the Groza et al (2022) preprint.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Quantitative PCR from human genomic DNA: the determination of gene copy numbers for congenital adrenal hyperplasia and RCCX copy number variation

<p>The dataset is related a study in which we aimed to simultaneously assess the performance of 7 quantitative polymerase chain reaction (qPCR) assays for the gene copy number (GCN) determination of the genetic elements of RCCX copy number variation (CNV). A single laboratory method validations of duplex qPCR assays with hydrolysis probes on <em>CYP21A1P</em> and <em>CYP21A2</em> genes, which are responsible for congenital adrenal hyperplasia, were performed using 46 human genomic DNA samples. We also performed the verifications on 5 qPCR assays for the genetic elements of RCCX CNV such as <em>C4A</em> gene, <em>C4B</em>, gene, RCCX CNV breakpoint, HERV-K(C4) CNV deletion and insertion alleles. The dataset contains the data of genomic DNA samples, the raw quantification cycle values of all qPCR experiments, the peak heights and dosage quotient of multiplex ligation-dependent probe amplification (MLPA) experiments, and the detailed GCN results based on qPCR and MLPA. All other analyses are available in our publication under the same title.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Early-life human gut metagenome-assembled genomes and proteins catalogs

<p>The description of the files:</p> <p>(1) The 32,277 genomes include&nbsp;six parts:&nbsp;ELGG_part_1.zip,&nbsp;ELGG_part_2.zip,&nbsp;ELGG_part_3.zip,&nbsp;ELGG_part_4.zip,&nbsp;ELGG_part_5.zip,&nbsp;ELGG_part_6.zip.</p> <p>(2) The 2,172 representative&nbsp;species: ELGG_representatives_2172.zip.</p> <p>(3) The&nbsp;ELGP&nbsp;catalog&nbsp;clustered at 95% amino acid identity:&nbsp;ELGP_95.faa.gz.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Genomes and associated scripts for paper: Potential millennial-scale avian declines by humans in southern China

<p><span>Mounting observational records demonstrate human-caused faunal decline in recent decades, while accumulating archaeological evidence suggests an early biodiversity impact of human activities during the Holocene. A fundamental question arises concerning whether modern wildlife population declines began during early human disturbance. Here, we performed population genomic analysis of six common forest birds in East Asia to address this question. For five of them, demographic history inference based on 25-33 genomes of each species revealed dramatic population declines by 4-48-fold over millennia (two to five thousand years ago). Nevertheless, <a name="_Hlk39081165"></a>ecological niche models predicted extensive range persistence during the Holocene and imply limited demographic impact of historical climate change. Summary statistics further suggest high negative correlations between these population declines and human disturbance intensities and indicate a potential driver of human activities. These findings provide deep-time and large-scale insight into the recently recognized avifaunal decline and support an early origin hypothesis of human effects on biodiversity. Overall, our study sheds light on the current biodiversity crisis in the context of long-term human-environment interactions and offers a multievidential framework for quantitatively assessing the ecological consequences of human disturbance.</span></p>

opencc-zeroAug 2022View details →
zenodo36/100

30-mer mappable regions in the human hg19 genome

<p>Knowing where reads can uniquely map in the genome is useful for nascent RNA assays, both in statistical calculations and to make predictions.</p> <p>The dataset was created using the bowtie 1 aligner.&nbsp; The genome was windows at 30 basepair genomic intervals and mapped back to the genome.&nbsp; If the read maps to more than one place, the read is thrown away.&nbsp; Therefore the regions captured in the dataset are regions that any read at least 30 basepairs long will map to uniquely.&nbsp; The shell script originally used to create this dataset has been lost.</p>

opencc-by-4.0May 2019View details →
zenodo36/100

Virulence and antibiotic resistance plasticity of Arcobacter butzleri: insights on the genomic diversity of an emerging human pathogen (genome assembly, annotation dataset, core- and pan-genome loci)

<p>This dataset refers to the analysis of 49 <em>Arcobacter butzleri</em> genomes and includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the predicted&nbsp;transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files), the respective amino acid sequences of the translated CDS sequences (.faa files), the nucleotide alignments of all the 1165 core-genome loci,&nbsp;the nucleotide alignments of the genes <em>hecA</em>, <em>tetR </em>and <em>porA</em>, the categorized amino acid sequences of the six hypervariable regions of PorA, and the nucleotide sequences of the first allele of each of the 7474 pan-genome loci with the respective complete allelic profile matrix.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB34441).</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Comparative whole genome phylogeny of animal, environmental and human strains confirms the genogroups organization and the diversity of Stenotrophomonas maltophilia

<p>Reannotation of Smc genomes from Refseq (Prokka&nbsp;v1.13) and&nbsp;gene presence and absence spreadsheet from Roary.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on Human Genome Diversity Project Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the HGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>2b388c1fa446ecec70e33ea0471e06f8</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>50b80ed32b1ae542c8967cc31986dd19</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>995f30b74c4bb094a674b1a994853246</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>e79e61efab4c491fa2825b7d1853df58</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Machine learning reveals the diversity of human 3D chromatin contact patterns (example predictions genome wide)

<p>Example data for the paper: Machine learning reveals the diversity of human 3D chromatin contact patterns</p> <p>GitHub: https://github.com/erin-n-gilbertson/3DGenome-diversity/tree/main</p> <p>biorXiv: https://www.biorxiv.org/content/10.1101/2023.12.22.573104v1.full</p> <p>Manuscript accepted at Molecular Biology and Evolution</p> <p>Of primary interest will be the example predictions genome wide for hg38 reference, human-archaic hominin ancestor and most divergent 1KG individual per genome along with the Jupyter notebook tutorial for making your own Akita predictions given any input 1MB sequence.</p> <div> <ul> <li>bin: contains python script for and qsub array shell script for generating example predictions. These scripts can be modified to take in any fasta files as input.</li> <li>akita_predictions: contains both Akita prediction output arrays and SVG files with predicted contact maps for the hg38 reference, human-archaic hominin ancestor and most divergent 1KG individual in each of 4,873 1MB windows</li> <li>anc_window_spearman.csv: spearman correlation between each 1KG individual and the ancestor for each 1MB window. To calculate 3D divergence subtract these values from 1.</li> <li>basenji: basenji dir from their github, necessary in the directory to run predictions - https://github.com/calico/basenji/tree/master</li> <li>genomes: fasta genomes for hg38 reference and human-archaic hominin ancestor used to make akita predictions</li> <li>divergent_windows: variants and expected divergence distributions for 392 more divergent than expected windows. Defined in the manuscript as windows where 3D divergence between 1KG indiivudals and the ancestor is greater than what would be expected based on sequence divergence. See manuscript Fig. S9 for more details.&nbsp;</li> <li>windows.txt: 4,873 1MB genomic windows with 100% coverage in hg38 used for Akita predictions</li> <li>making_examples.ipynb: jupyter notebook with tutorial instructions for making Akita predictions on any human genome sequence.</li> </ul> <br><br></div>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Mode-of-inheritance predictions for all possible missense variants in the human genome (hg38)

<p><span>Ensemble and consensus approaches to prediction of recessive inheritance for missense variants in human disease.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Metagenome-assembled genomes(MAGs) generated from CRC human gut (PRJEB27928).

<p>MAGs generated from&nbsp;CRC human gut (PRJEB27928) with Maxbin2, VAMB, Metabat2, SemiBin(single-sample binning) and VAMB, SemiBin(multi-sample binning).</p> <p>Single-sample binning: Maxbin2.tar.gz, Metabat2.tar.gz, VAMB.tar.gz and SemiBin(_pretrain).tar.gz.&nbsp;</p> <p>Multi-sample binning: VAMB_multi.tar.gz and SemiBin_multi.tar.gz.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Human genomes for GGCAT benchmarks - part 1

<p>The first 50 genomes of the 100 Human genomes used for GGCAT benchmarks</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Convergent genomics of longevity in rockfishes highlights the genetics of human lifespan variation

<p>Longevity is a defining, heritable trait that varies dramatically between species. To resolve the genetic regulation of this trait, we have mined genomic variation in rockfishes, which range in longevity from 11 to over 205 years, along with the sea robin as an outgroup. We used the phylomapping methodology to target sequencing and assembly to conserved, functional loci, defined using available genomic resources from well-characterized species. These annotations are in the gtf. For analysis of coding elements, assembled contigs were trimmed to just the exons. Exons were concatenated and analyzed to generate gene trees with IQTree, fixed to match the species tree. Additional details of these assembly methods are available at&nbsp;DOI:&nbsp;<a href="https://doi.org/10.1038/s41559-019-0914-2">10.1038/s41559-019-0914-2</a> and <a href="https://doi.org/10.1126/sciadv.add2743">DOI: 10.1126/sciadv.add2743</a></p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

FUN-LDA scores for human genome assembly GRCh37

<p>FUN-LDA is based on a Latent Dirichlet Allocation (LDA) model for predicting functional effects of non-coding genetic variants in a cell type and tissue-specific way by integrating diverse epigenetic annotations for specific cell types and tissues from large-scale genomics projects such as ENCODE and Roadmap Epigenomics. Using this unsupervised approach, we predict tissue-specific functional effects for every position in the human genome for 127 tissues and cell types in ENCODE and Roadmap Epigenomics.&nbsp;This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&amp;usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Format</strong></p> <p>The FUN-LDA scores are stored in the UCSC Genome Browser bigWig Track Format.&nbsp;</p> <p>To extract FUN-LDA scores, the bigWigAverageOverBed utility is required. It can be downloaded from the Genome Browser website at&nbsp;<a href="https://hgdownload.cse.ucsc.edu/admin/exe">https://hgdownload.cse.ucsc.edu/admin/exe</a>.</p> <p>User should prepare a UCSC Genome Browser bed file, a tab-separated four-column file. The first column is the chromosome; the second is the zero-based coordinate of the position&nbsp;of interest; the third is that zero-based coordinate plus one; and the fourth is a unique identifier for the position. The command below is an example&nbsp;of extracting FUN-LDA scores in tissue E007 for the positions defined in a bed file, &quot;input.bed&quot;.</p> <pre><code class="language-bash">bigWigAverageOverBed E007.valley9.c89.bigwig input.bed output.tab</code></pre> <p>It produces a six-column file, output.tab.&nbsp;The first column includes the unique position id defined in the input bed file.&nbsp;The last column, average over the covered bases, is the score for this position.</p> <p><strong>Reference</strong></p> <p>Daniel Backenroth, Zihuai He, Krzysztof Kiryluk, Valentina Boeva, Lynn Pethukova, Ekta Khurana, Angela Christiano, Joseph Buxbaum, Iuliana Ionita-Laza. FUN-LDA: A latent Dirichlet allocation model for predicting tissue-specific functional effects of noncoding variation: Methods and applications. American Journal of Human Genetics, 2018.<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Impressive pan-genomic diversity of E. coli from a wild animal community near urban development reflects human impacts

<p>Data provided here support the findings of this&nbsp;study and code will allow the&nbsp;replication of analyses therein. To do so, first&nbsp;download and unzip the associated data files.</p> <p>For each analysis, the following files are required:</p> <p>PCAs: PCA.R, prokka_annotations.zip (and VirulenceFinder_results.csv&nbsp;for virulence factor PCA)</p> <p>Sankey diagram: Plasmid_AMR_Sankey.R, mlplasmids_results.zip, MOBsuite_results.zip, ResFinder_results.csv</p> <p>Pathotype assignment: VF_data_analysis.R,&nbsp;VirulenceFinder_results.csv</p> <p>TableS1_Ecoli_metadata.csv contains the isolate metadata (n=143) for all above analyses.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record