Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “NA12878”
NA12878 WES Benchmark dataset
<p>This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session <strong>NA12878 WES Benchmark </strong>files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. </p> <p>The <a href="https://usegalaxy.org/u/erinija/p/omim-genes-in-na12878-wes-benchmark">"Procedure and datasets to cross-reference OMIM genes with the genomic regions of interest"</a> Galaxy page on <strong>usegalaxy.org</strong> server's <em><strong>Shared Data Pages</strong></em> describes practical procedure and several possible use cases for this data set. This page can be accessed freely by users logged into their accounts on usegalaxy.org. Please register if you don't have an account on usegalaxy.org Galaxy server. </p> <p>All genomic variant calls in all VCF files of this data set were decomposed and normalized with vt. This dataset contains: </p> <ol> <li>Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : <ol> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz</li> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi</li> <li>GIAB_v3.3.2_NA12878_HC_regions.bed</li> </ol> </li> <li>HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : <ul> <li>ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser <ol> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz</li> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi</li> <li> ARUP_SeqCap_EZ_Exome.bed</li> </ol> </li> <li>UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser <ol> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz</li> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi</li> <li>UCSF_WES_Agilent_V4_Custom.bed</li> </ol> </li> <li>Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory <ol> <li>CHEO_NA12878_WES_S1dataset.vcf.gz</li> <li>CHEO_NA12878_WES_S1dataset.vcf.gz.tbi</li> <li>Agilent_CRE_v2.bed</li> </ol> </li> </ul> </li> <li>Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : <ul> <li>Omim_Genes.bed </li> </ul> </li> </ol>
Coriell Index NA12878
<p>Variant calls from the first human genome analysed at University Clinical Center in Gdansk. The sample was sequenced at Genomics Core Facility in Bergen, Norway.</p> <p>Technology: Illumina HiSeq 4000</p> <p>Reference: GRCh38</p> <p>Alignment: cgpwgs - cgpmap 2.1.1</p> <p>Variant Calling: DeepVariant 1.6.1</p> <p> </p>
NA12878, 22RV1 and NB4 sequence dataset
<p>Adaptive Nanopore PromethION multiplexed sample. </p> <p>NA12878 is Barcode 05.</p> <p>NB4 is Barcode 06.</p> <p>22Rv1 is Barcode 07.</p> <p>The FASTQ is split by whether the read was sequenced or actively rejected by readfish.</p> <p>Corresponds to the dataset used in the readfish dataset.</p>
NA12878 and MCF7 data for Profiling Chromatin Accessibility in Humans Using Adenine Methylation and Long-Read Sequencing
<p>This dataset includes 5mC and 6mA frequency data for NA12878 and MCF7 EcoGII-treated chromatin samples sequenced on nanopore r9.4.1.</p>
Genotypes, variants and pedigree from a human parent-offspring trio (NA12878)
<p>This dataset includes whole genome sequencing data, produced at the French National Research Center for Human Genomics (CNRGH), with known pedigree information:<br> - The VCF file contains genotypes and variants from a CEU parent-offspring trio comprising NA12878 (child), NA12891 (father) and NA12892 (mother).<br> For each member, 150-bp paired-end whole genome sequencing data were generated on the Illumina HiSeq X system, from a PCR-free library (40x).<br> - The TFAM file contains pedigree information of this CEU parent-offspring trio.</p>
Agilent v7 NA12878 VCF files
<p>Agilent v7 VCF files generated for the NGS-CN benchmarking with the development nextflow exomme pipeline version of the CCG.</p>
Agilent v7 exomes of NA12878
<p>FastQ files used in the NGS-CN benchmarking initiative</p>
NA12878 WGS 30x
<p>Normal sample used in combination with simulated tumor for benchmarking purposes.</p> <p>10.5281/zenodo.11203957</p> <p>Created for 1+MG WG9.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.