Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,666
datasets available to search
ShareScore release 0.9.0
Dataset results
1,666 results for “human genome”
Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)
<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs "170630_M00528_0292_000000000-B9JY8" (aka "NC_LIMMS") and "180221_M00528_0334_000000000-B6PJM" (aka "NC_LIMMS2"). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics 2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p> </p> <p><em><strong>"170630_M00528_0292_000000000-B9JY8" ("NC_LIMMS") :</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>1 iPSC_control_rep1 4 ACAGAT NNNNNNNN</p> <p>2 iPSC_control_rep2 24 ATCGTG NNNNNNNN</p> <p>3 iPSC_control_rep3 31 CACGAT NNNNNNNN</p> <p>4 S3P1_OK_rep1 36 CACTGA NNNNNNNN</p> <p>5 S3P1_OK_rep2 46 CTGACG NNNNNNNN</p> <p>6 S3P1_OK_rep3 63 GAGTGA NNNNNNNN</p> <p>7 S4P1_OK_rep1 79 GTATAC NNNNNNNN</p> <p>8 S4P1_OK_rep2 92 TCGAGC NNNNNNNN</p> <p>9 S4P1_OK_rep3 9 ACATGA NNNNNNNN</p> <p>10 S4P2_OK_rep1 21 ATCATA NNNNNNNN</p> <p>11 S4P2_OK_rep2 33 CACGTG NNNNNNNN</p> <p>12 S4P2_OK_rep3 45 CGATGA NNNNNNNN</p> <p>13 S1P1_rep1 57 GAGATA NNNNNNNN</p> <p>14 S1P1_rep2 69 GCTCTC NNNNNNNN</p> <p>15 S1P1_rep3 81 GTATGA NNNNNNNN</p> <p>16 S3P1_FAILED_rep1 93 TCGATA NNNNNNNN</p> <p>17 S3P1_FAILED_rep2 11 AGTAGC NNNNNNNN</p> <p>18 S3P1_FAILED_rep3 23 ATCGCA NNNNNNNN</p> <p>19 S4P1_FAILED_rep1 35 CACTCT NNNNNNNN</p> <p>20 S4P1_FAILED_rep2 47 CTGAGC NNNNNNNN</p> <p>21 S4P1_FAILED_rep3 59 GAGCGT NNNNNNNN</p> <p>22 S4P2_FAILED_rep1 71 GCTGCA NNNNNNNN</p> <p>23 S4P2_FAILED_rep2 83 TATAGC NNNNNNNN</p> <p>24 S4P2_FAILED_rep3 95 TCGCGT NNNNNNNN</p> <p> </p> <p><em><strong>"180221_M00528_0334_000000000-B6PJM" ("NC_LIMMS2"):</strong></em></p> <p><strong>ID Sample_name Barcode_number Barcode_sequence Index_sequence</strong></p> <p>25 PETRI_rep1 04 ACAGAT NNNNNNNN</p> <p>26 PETRI_rep2 24 ATCGTG NNNNNNNN</p> <p>27 PETRI_rep3 31 CACGAT NNNNNNNN</p> <p>28 BIOCHIP_E_rep1 6 CACTGA NNNNNNNN</p> <p>29 BIOCHIP_M_rep1 46 CTGACG NNNNNNNN</p> <p>30 BIOCHIP_S_rep1 63 GAGTGA NNNNNNNN</p> <p>31 BIOCHIP_E_rep2 79 GTATAC NNNNNNNN</p> <p>32 BIOCHIP_M_rep2 92 TCGAGC NNNNNNNN</p> <p>33 BIOCHIP_S_rep2 09 ACATGA NNNNNNNN</p> <p>34 BIOCHIP_E_rep3 21 ATCATA NNNNNNNN</p> <p>35 BIOCHIP_M_rep3 33 CACGTG NNNNNNNN</p> <p>36 BIOCHIP_S_rep3 45 CGATGA NNNNNNNN</p> <p>37 HEPATOCYTES_rep1 57 GAGATA NNNNNNNN</p> <p>38 HEPATOCYTES_rep2 69 GCTCTC NNNNNNNN</p> <p>39 iPSC_control_rep1-2 81 GTATGA NNNNNNNN</p> <p>40 BIOCHIP_E_rep2-2 93 TCGATA NNNNNNNN</p> <p>41 BIOCHIP_M_rep1-2 11 AGTAGC NNNNNNNN</p> <p>42 BIOCHIP_S_rep2-2 23 ATCGCA NNNNNNNN</p> <p> </p>
Human genomes for GGCAT benchmarks - part 2
<p>The last 50 genomes of the 100 Human genomes used for GGCAT benchmarks<br> </p>
Human genome annotation file hg38
<p>Human genome annotation file hg38 from ensembl including snoRNA from snoDB database</p>
Human reference genome analysis sets
<p>For details, see the companion <a href="https://github.com/lh3/ref-gen">GitHub repo</a>.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics
<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>
Data from: Genome sequences reveal cryptic speciation in the human pathogen Histoplasma capsulatum
Open the record for dataset details and reuse information.
MtDNA genomes from Ranis individuals aligned with previously published ancient & modern humans
Open the record for dataset details and reuse information.
Data from: Comparative genomics, infectivity and cytopathogenicity of Zika viruses produced by acutely and persistently Zika virus-infected human hematopoietic cell lines
Open the record for dataset details and reuse information.
Data from: Rates of genomic divergence in humans, chimpanzees and their lice
Open the record for dataset details and reuse information.
Data from: Genomic DNA transposition induced by human PGBD5
Open the record for dataset details and reuse information.
Data from: Modeling human population separation history using physically phased genomes
Open the record for dataset details and reuse information.
Data from: Forward genetic screen of human transposase genomic rearrangements
Open the record for dataset details and reuse information.
Data from: Genomic characterization of large heterochromatic gaps in the human genome assembly
Open the record for dataset details and reuse information.
Genome wide identification of p63 binding sites in human neonatal foreskin keratinocytes
GEO Series GSE32061. Homo sapiens. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Elimination of mitochondrial DNA variants by nuclear genome transfer in human oocytes
GEO Series GSE42077. Homo sapiens. 11 samples. Type: Expression profiling by array.
Whole-genome microarray analysis of human skin fibroblasts
GEO Series GSE69447. Homo sapiens. 9 samples. Type: Expression profiling by array.
Genomic and transcriptomic profiles underlying the effect of SETD1A disruption on human neurodevelopmental trajectories
GEO Series GSE237165. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.
Genome-wide CRISPR screen for comprehensive identification of human factors involved in alternative polyadenylation based on differential localization of CD47 protein
GEO Series GSE288919. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
Genome-wide integration of microRNA and the transcriptome during human alveolar epithelial cell transdifferentiation identifies SGK1 as novel target of miR- 424/503
GEO Series GSE140073. Homo sapiens. 27 samples. Type: Non-coding RNA profiling by array; Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.