Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,666
datasets available to search
ShareScore release 0.9.0
Dataset results
1,666 results for “human genome”
Neandertal ancestry through time: Insights from genomes of ancient and present-day humans
Open the record for dataset details and reuse information.
Genomic analysis reveals close genetic similarity between ESBL-producing E. coli isolates from humans and dogs, suggesting potential for inter-species transmission
Open the record for dataset details and reuse information.
De novo genome assembly of human cell line CHM13 nanopore ultra-long reads using Shasta
Open the record for dataset details and reuse information.
Supporting data for: Three-dimensional genome re-wiring in loci with Human Accelerated Regions
Open the record for dataset details and reuse information.
Eigen scores for human genome assembly GRCh38 Part 1 (Chr12 - Chr22)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Eigen scores for human genome assembly GRCh38 Part 2 (Chr6 - Chr11)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Systematic shotgun of human genome [fastq.gz]
<p>Systematic shotgun of hg38 genome for testing the ability to identify human reads.</p>
Synthetic Noisy Reads Generated from Human Genomes
<p>This repository contains datasets with size 200K, 400K and 1 Million noisy reads generated from 5000 transcripts of GTcenters.fasta file. The assigned task is to recover the ground-truth transcripts (GTcenters) based on the given noisy reads.</p>
Human genomics of the humoral immune response against polyomaviruses
<p>Meta-analyses results of the manuscript "Human genomics of the humoral immune response against polyomaviruses".</p>
Data for "Analysis of metagenome-assembled viral genomes from the human gut reveals diverse putative CrAss-like phages with unique genomic features"
<p>Data for "Analysis of metagenome-assembled viral genomes from the human gut reveals diverse putative CrAss-like phages with unique genomic features" (submitted to Nature Communications)</p>
Human genome ancestral state files
<p>Files containing ancestral states at all positions of the human genome, inferred based on consensus support among three ape species; Gorilla, Chimpanzee and Orangutan.</p>
Data from: Differential requirements for the RAD51 paralogs in genome repair and maintenance in human cells
Deficiency in several of the classical human RAD51 paralogs [RAD51B, RAD51C, RAD51D, XRCC2 and XRCC3] is associated with cancer predisposition and Fanconi anemia. To investigate their functions, isogenic disruption mutants for each were generated in non-transformed MCF10A mammary epithelial cells and in transformed U2OS and HEK293 cells. In U2OS and HEK293 cells, viable ablated clones were readily isolated for each RAD51 paralog; in contrast, with the exception of RAD51B, RAD51 paralogs are cell-essential in MCF10A cells. Underlining their importance for genomic stability, mutant cell lines display variable growth defects, impaired sister chromatid recombination, reduced levels of stable RAD51 nuclear foci, and hyper-sensitivity to mitomycin C and olaparib. Altogether these observations underscore the contributions of RAD51 paralogs in diverse DNA repair processes, and demonstrate essential differences in different cell types. Finally, this study will provide useful reagents to analyze patient-derived mutations and to investigate mechanisms of chemotherapeutic resistance deployed by cancers.
"interactive" version of data associated with the eLife paper "Integrative genomic analysis of the human immune response to influenza vaccination"
This archive contains an "interactive" version of data associated with the eLife paper "Integrative genomic analysis of the human immune response to influenza vaccination" by Luis M Franco, Kristine L Bucasas, Janet M Wells, Diane Niño, Xueqing Wang, Gladys E Zapata, Nancy Arden, Alexander Renwick, Peng Yu, John M Quarles, Molly S Bray, Robert B Couch, John W Belmont, Chad A Shaw http://dx.doi.org/10.7554/eLife.00299 (doi:10.7554/eLife.00299) Installation: (1) Download the rar file, link available from the paper. (2) Unrar the file: Mac OS X: Use UnRarX - http://www.unrarx.com Linux : unrar command - http://en.wikipedia.org/wiki/Unrar (3) Open the file "vaxgenomics.htm" in your favorite browser use File>Open : Once opened in your browser, you should see a "Circos"-style circular plot of the data, and links to the tables and figures used in this application. Each table includes a link-out from GeneID to the Entrez entry for that Gene as well as additional links to the data. The link to Entrez requires access to NCBI, so will only work if you have Internet availability for this to work. All other links point to content contained within the application/archive. (4) Figures: High-resolution copies of the figures included in the paper. Fig 1: eQTL profile of flu vaccination response. Markers associated with cis gene expression identified in the discovery cohort were replicated in a vaidation cohort and -log10 p-values for both data sets are shown in the genome wide circularized graphic. Fig 2: Effect of the treatment on eQTL assocation. We observed that the pattern of association between gene expression and SNP changed over time after vaccination. Panel A of this figure shows this phenomena for a single gene NECAB2. Panel B shows the aggregate character of this phenomena across large numbers of markers. The change in R2 compared to the initial time point is depicted; this change in R squared appears to correspond to an increase the magnitude of the slope (additive association with genotype). Fig 3: Pathways and processes identified as enriched in our candidates. Both Ingenuity IPA Analysis and GO and KEGG pathway databases were used. Heavily implicated immunologic response classes are repsresented. Fig 4: Human immune response as measured by our Antibody Response scores are correlated with gene expression changes, and these patterns are recapituoated in discovery and validation (male/female) cohorts. Fig 5: Immune cellular context of the genes identifeies in our eQTL and immune response analysis. A striking number of our validated genes occuue in the antigen processing and presentatio pnthway. Fig 6: Q-Q plot depicting that the strength of association between genotype and phenotype (titer response) is stronger for markers that have a SNP association with expression and where expression is associated with titer response than would be expected for random SNPs. Fig 7: Causal and Reactive Model Analyses. Three-way association between genotype, expression and trait. Our data are more consistent with a causal relationship compared to reactive, but the results are not definitive. We explored the sample size necessary to investigate this in the supplement. Fig 8: Diagram demonstrating the experimental design of this study. Fig 9: Population structure analysis performed on our cohort using the genome wide SNP data confirms the European ancestry and ethnic homogeoneity of our ty sample Fig 10: Schematic of the eQTL analysis. The time course of gene expression change is integrated with a single model considering effects of Day, Genotype, Day-Genotype interaction and random effects for each individual to account for the longitudinal nature of the design. email: cashaw@bcm.edu with questions or comments.
ChromBERT: Uncovering Chromatin State Motifs in the Human Genome using a BERT-based Approach
<ol> <li>Pretrain data and results for: <ul> <li>Promoter regions for all genes in 127 different cell lines in ROADMAP </li> <li>CRM (<em>cis</em>-regulatory module) regions longer than 2k bps in 127 different cell lines in ROADMAP</li> <li>Whole genome regions without continuous low signal state ("O") in 4-mer</li> </ul> </li> <li>Fine-tuning data and result for: <ul> <li>[Classification] Promoter regions of high-expressed genes (from RPKM>10 to RPKM>50) compared to not expressed genes (RPKM=0) or low-expressed genes (RPKM>0) with the directory names:<br> <ul> <li>not_n_rpkm0 : RPKM=0 vs. RPKM>0</li> <li>not_n_rpkm10 : RPKM=0 vs. RPKM>10</li> <li>not_n_rpkm20 : RPKM=0 vs. RPKM>20 </li> <li>not_n_rpkm30 : RPKM=0 vs. RPKM>30 </li> <li>not_n_rpkm50 : RPKM=0 vs. RPKM>50 </li> <li>rpkm0_n_rpkm10 : RPKM>0 vs. RPKM>10 </li> <li>rpkm0_n_rpkm20 : RPKM>0 vs. RPKM>20 </li> <li>rpkm0_n_rpkm30 : RPKM>0 vs. RPKM>30 </li> <li>rpkm0_n_rpkm50 : RPKM>0 vs. RPKM>50 </li> <li>rpkm10_n_rpkm20 : RPKM>10 vs. RPKM>20 </li> <li>rpkm10_n_rpkm30 : RPKM>10 vs. RPKM>30 </li> <li>rpkm10_n_rpkm50 : RPKM>10 vs. RPKM>50 </li> <li>rpkm20_n_rpkm30 : RPKM>20 vs. RPKM>30 </li> <li>rpkm20_n_rpkm50 : RPKM>20 vs. RPKM>50 </li> <li>rpkm30_n_rpkm50 : RPKM>30 vs. RPKM>50 </li> </ul> </li> <li>[Regression] Quantitative gene expression prediction for promoter regions</li> <li>CRM regions longer than 2k bps compared to non-CRM regions</li> </ul> </li> </ol>
Updated Metagenomic Species Pan-genomes (MSPs) of the human gastrointestinal microbiota
<p></p><h1>Gene catalog construction</h1><br>The methodology for creating the IGC2 catalog is described in the original papers: Li et al., 2014 and Wen et al., 2017<br><h1>MSP creation</h1><br>Reads from publicly available human gut metagenomes were aligned against the IGC2 catalog with the Meteor to produce a raw gene abundance table (10.4M genes quantified in >2000 samples). Then, co-abundant genes were binned in 1,989 Metagenomic Species Pan-genomes (MSPs, i.e. clusters of co-abundant genes that likely belong to the same microbial species) using MSPminer.<br><h1>MSPs taxonomic annotation</h1><br>MSPs taxonomic annotation was performed by aligning MSP core and accessory genes against representative genomes of the Genome Taxonomy Database (GTDB r207) using blastn (task = megablast, word_size = 16). The 20 best hits for each gene were kept (--max-target-seq 20). Using an in-house pipeline, a species-level assignment was given if > 50% of the genes matched the representative genome of a given species, with a mean identity ≥ 95% and mean gene length coverage ≥ 90%. The remaining MSPs were assigned to a higher taxonomic level (genus to superkingdom), if more than 50% of their genes had the same annotation.<br><h1>Construction of the phylogenetic tree</h1><br>39 universal phylogenetic markers genes were extracted from the MSPs with fetchMGs. Then, the markers were separately aligned with MUSCLE. The alignments were merged and trimmed with trimAl (parameters: -automated1). Finally, the phylogenetic tree was computed with FastTreeMP (parameters: -gamma -pseudo -spr -mlacc 3 -slownni). <h1>Mapping rate distribution across public cohorts</h1>We generated mapping rate distribution plots using Meteor2 (default parameters), comparing performance between: PRJEB1786, PRJEB5224, PRJEB6337, PRJNA422434 (cohort used in catalogue assembly) and PRJEB11532, PRJEB33500, PRJEB37249, PRNJNA834801 (independent cohort not used in assembly).<p></p>
CRISPR-Cas12a-integrated transgenes in genomic safe harbors retain high expression in human hematopoietic iPSC-derived lineages and primary cells
<p>We identified and characterized potential integration safe harbor sites (SHS) in human cells. Using the CRISPR-MAD7 system, we integrated transgenes at these genomic sites in iPSC, primary T and NK cells, and Jurkat cell line, and demonstrated efficient and stable expression at these loci. Subsequently, we validated the differentiation capabilities of engineered iPSC towards CD34+ hematopoietic stem and progenitor cells (HSPC), lymphoid progenitors (LPC), and natural killer (NK) cells, and showed that transgene expression was retained in these lineages. </p>
A genomic compendium of cultivated human gut fungi characterizes the gut mycobiome and its relevance to common diseases
<p><span>Morphological and scanning electron microscopy images of 206 fungal species cultivated from human feces. </span></p>
Data from: Demographic inference from whole-genome and RAD sequencing data suggests alternating human impacts on goose populations since the last ice age
We investigated how population changes and fluctuations in the pink-footed goose might have been affected by climatic and anthropogenic factors. First, genomic data confirmed the existence of two separate populations: western (Iceland) and eastern (Svalbard/Denmark). Second, emographic inference suggests that the species survived the last glacial period as a single ancestral population with a low population size (100-1,000 individuals) that split into the current populations at the end of the Last Glacial Maximum with Iceland being the most plausible glacial refuge. While population changes during the last glaciation were clearly environmental, we hypothesize that more recent demographic changes are human-related: (1) the inferred population increase in the Neolithic is due to deforestation to establish new lands for agriculture, increasing available habitat for pink-footed geese (2) the decline inferred during the Middle Ages is due to human persecution and (3) improved protection explains the increasing demographic trends during the 20th century. Our results suggest both environmental (during glacial cycles) and anthropogenic effects (more recent) can be a threat to species survival.
Human genome compressibility
<p>Data for https://bioinformatics.stackexchange.com/questions/12824/theoretical-limit-of-human-genome-compression and https://github.com/karel-brinda/human-genome-compression</p>
Genome-wide survey of D/E repeats in human proteins uncovers their instability and aids in identification of their role in the chromatin regulator ATAD2
<p>Raw digital images of Western blots that comprise Figure 8E of the manuscript</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.