Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,666

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,666 results for “human genome”

Learn how ShareScore rates datasets ↗
zenodo28/100

Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)

<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs &quot;170630_M00528_0292_000000000-B9JY8&quot; (aka &quot;NC_LIMMS&quot;) and &quot;180221_M00528_0334_000000000-B6PJM&quot; (aka &quot;NC_LIMMS2&quot;). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics&nbsp;2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under&nbsp;the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p>&nbsp;</p> <p><em><strong>&quot;170630_M00528_0292_000000000-B9JY8&quot; (&quot;NC_LIMMS&quot;) :</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>1&nbsp;&nbsp; iPSC_control_rep1&nbsp;&nbsp; 4&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>2&nbsp;&nbsp; iPSC_control_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>3&nbsp;&nbsp; iPSC_control_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>4&nbsp;&nbsp; S3P1_OK_rep1&nbsp;&nbsp; 36&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>5&nbsp;&nbsp; S3P1_OK_rep2&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>6&nbsp;&nbsp; S3P1_OK_rep3&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>7&nbsp;&nbsp; S4P1_OK_rep1&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>8&nbsp;&nbsp; S4P1_OK_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>9&nbsp;&nbsp; S4P1_OK_rep3&nbsp;&nbsp; 9&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>10&nbsp;&nbsp; S4P2_OK_rep1&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>11&nbsp;&nbsp; S4P2_OK_rep2&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>12&nbsp;&nbsp; S4P2_OK_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>13&nbsp;&nbsp; S1P1_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>14&nbsp;&nbsp; S1P1_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>15&nbsp;&nbsp; S1P1_rep3&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>16&nbsp;&nbsp; S3P1_FAILED_rep1&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>17&nbsp;&nbsp; S3P1_FAILED_rep2&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>18&nbsp;&nbsp; S3P1_FAILED_rep3&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>19&nbsp;&nbsp; S4P1_FAILED_rep1&nbsp;&nbsp; 35&nbsp;&nbsp; CACTCT&nbsp;&nbsp; NNNNNNNN</p> <p>20&nbsp;&nbsp; S4P1_FAILED_rep2&nbsp;&nbsp; 47&nbsp;&nbsp; CTGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>21&nbsp;&nbsp; S4P1_FAILED_rep3&nbsp;&nbsp; 59&nbsp;&nbsp; GAGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>22&nbsp;&nbsp; S4P2_FAILED_rep1&nbsp;&nbsp; 71&nbsp;&nbsp; GCTGCA&nbsp;&nbsp; NNNNNNNN</p> <p>23&nbsp;&nbsp; S4P2_FAILED_rep2&nbsp;&nbsp; 83&nbsp;&nbsp; TATAGC&nbsp;&nbsp; NNNNNNNN</p> <p>24&nbsp;&nbsp; S4P2_FAILED_rep3&nbsp;&nbsp; 95&nbsp;&nbsp; TCGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p> <p><em><strong>&quot;180221_M00528_0334_000000000-B6PJM&quot; (&quot;NC_LIMMS2&quot;):</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>25&nbsp;&nbsp; PETRI_rep1&nbsp;&nbsp; 04&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>26&nbsp;&nbsp; PETRI_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>27&nbsp;&nbsp; PETRI_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>28&nbsp;&nbsp; BIOCHIP_E_rep1&nbsp;&nbsp; 6&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>29&nbsp;&nbsp; BIOCHIP_M_rep1&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>30&nbsp;&nbsp; BIOCHIP_S_rep1&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>31&nbsp;&nbsp; BIOCHIP_E_rep2&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>32&nbsp;&nbsp; BIOCHIP_M_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>33&nbsp;&nbsp; BIOCHIP_S_rep2&nbsp;&nbsp; 09&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>34&nbsp;&nbsp; BIOCHIP_E_rep3&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>35&nbsp;&nbsp; BIOCHIP_M_rep3&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>36&nbsp;&nbsp; BIOCHIP_S_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>37&nbsp;&nbsp; HEPATOCYTES_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>38&nbsp;&nbsp; HEPATOCYTES_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>39&nbsp;&nbsp; iPSC_control_rep1-2&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>40&nbsp;&nbsp; BIOCHIP_E_rep2-2&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>41&nbsp;&nbsp; BIOCHIP_M_rep1-2&nbsp;&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>42&nbsp;&nbsp; BIOCHIP_S_rep2-2&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p>

openOct 2017View details →
zenodo28/100

Human genomes for GGCAT benchmarks - part 2

<p>The last 50 genomes of the 100 Human genomes used for GGCAT benchmarks<br> &nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo28/100

Human genome annotation file hg38

<p>Human genome annotation file hg38 from ensembl including snoRNA from snoDB database</p>

opencc-by-4.0May 2023View details →
zenodo28/100

Human reference genome analysis sets

<p>For details, see the companion <a href="https://github.com/lh3/ref-gen">GitHub repo</a>.</p>

opencc-by-4.0Jun 2023View details →
zenodo28/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →
zenodo28/100

Genomic atlas of the human proteome from brain, CSF and plasma: Improvement with TOPMed imputed genomics

<p>Abstract</p> <p>Comprehensive expression quantitative trait loci (eQTL) studies have been instrumental for understanding tissue-specific gene regulation and pinpointing functional genes for disease-associated GWAS loci in a tissue-specific manner. Compared to gene expressions, proteins more directly affect various biological processes, often dysregulated in disease, and are important drug targets. We previously performed and identified tissue-specific protein QTL (pQTL) in neurologically relevant tissues. We now enhance this work by analyzing more proteins (1,300 versus 1,079) and an almost twofold increase in high-quality imputed genetic variants (8.4 million versus 4.4 million) by using TOPMed reference panel. We identified 38 genomic regions associated with 43 proteins in brain, 150 regions associated with 247 proteins in CSF, and 95 regions associated with 145 proteins in plasma. Compared to our previous study, this study newly identified 12 pQTL in brain, 30 pQTL in CSF, and 22 pQTL in plasma. Our improved genomic atlas uncovers the genetic control of protein regulation across multiple tissues. These pQTL findings are assessable through the Online Neurodegenerative Trait Integrative Multi-Omics Explorer (ONTIME) for use by the scientific community.</p>

opencc-by-4.0Jul 2023View details →
dryad28/100

Data from: Genome sequences reveal cryptic speciation in the human pathogen Histoplasma capsulatum

Open the record for dataset details and reuse information.

publicNov 2018View details →
dryad28/100

MtDNA genomes from Ranis individuals aligned with previously published ancient & modern humans

Open the record for dataset details and reuse information.

publicOct 2023View details →
dryad28/100

Data from: Comparative genomics, infectivity and cytopathogenicity of Zika viruses produced by acutely and persistently Zika virus-infected human hematopoietic cell lines

Open the record for dataset details and reuse information.

publicAug 2019View details →
dryad28/100

Data from: Rates of genomic divergence in humans, chimpanzees and their lice

Open the record for dataset details and reuse information.

publicDec 2014View details →
dryad28/100

Data from: Genomic DNA transposition induced by human PGBD5

Open the record for dataset details and reuse information.

publicSep 2015View details →
dryad28/100

Data from: Modeling human population separation history using physically phased genomes

Open the record for dataset details and reuse information.

publicNov 2017View details →
dryad28/100

Data from: Forward genetic screen of human transposase genomic rearrangements

Open the record for dataset details and reuse information.

publicJun 2017View details →
dryad28/100

Data from: Genomic characterization of large heterochromatic gaps in the human genome assembly

Open the record for dataset details and reuse information.

publicApr 2015View details →
geo24/100

Genome wide identification of p63 binding sites in human neonatal foreskin keratinocytes

GEO Series GSE32061. Homo sapiens. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenJun 2012View details →
geo24/100

Elimination of mitochondrial DNA variants by nuclear genome transfer in human oocytes

GEO Series GSE42077. Homo sapiens. 11 samples. Type: Expression profiling by array.

openGEO-OpenDec 2012View details →
geo24/100

Whole-genome microarray analysis of human skin fibroblasts

GEO Series GSE69447. Homo sapiens. 9 samples. Type: Expression profiling by array.

openGEO-OpenApr 2016View details →
geo24/100

Genomic and transcriptomic profiles underlying the effect of SETD1A disruption on human neurodevelopmental trajectories

GEO Series GSE237165. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2025View details →
geo24/100

Genome-wide CRISPR screen for comprehensive identification of human factors involved in alternative polyadenylation based on differential localization of CD47 protein

GEO Series GSE288919. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2025View details →
geo24/100

Genome-wide integration of microRNA and the transcriptome during human alveolar epithelial cell transdifferentiation identifies SGK1 as novel target of miR- 424/503

GEO Series GSE140073. Homo sapiens. 27 samples. Type: Non-coding RNA profiling by array; Expression profiling by high throughput sequencing.

openGEO-OpenJul 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record