Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.9.0
Dataset results
12 results for “GIAB”
HG005_Son.R1.fastq.gz (GIAB_ChineseTrio_30X)
<p>This is a sub-sampled version (30X) of ChineseTrio HG005(son)_R1 fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> As the total size of these data amounts to more than 200GB, we also prepared 10X fastq files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. <em>et al.</em> High-coverage, long-read sequencing of Han Chinese trio reference samples. <em>Sci Data</em> <strong>6</strong>, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
Rescued Phased VCF for GIAB HG001
<p><strong>The GIAB VCF file contained anomalies that elicited run-time errors from both LRphase and WhatsHap. This VCF is noncompliant with the VCF 4.3 specification (https://samtools.github.io/hts-specs/VCFv4.3.pdf) in at least two ways: 1) Phase Set (PS) tags within genotype fields contain strings instead of 32-bit integers. 2) Sample columns in the VCF column header are labeled ”INTEGRATION” rather than containing the sample name. We also found ~25,000 records with malformed genotype records, which contained extra fields not defined in the format string. Finally, we encountered errors from WhatsHap that suggested at least a subset of indel records are also malformed. Neither program would run successfully without correcting these errors, but we were able to rescue the VCF using a custom Python script. Briefly, we transliterated PS tag strings to integer values by concatenating a unique integer with the chromosome number for each record (24, 25, and 26, for chromosomes X, Y, and M, respectively). This ensured that all phase sets have integer labels and only included variants on the same chromosome. We defined an additional format tag, OPS, under which we stored the original PS tag values. Likewise, the “INTEGRATION” label in the column name header was replaced with the sample name, “HG001”. Variants with malformed genotype fields were rescued by removing the fields not defined in the format string. Since we were unable to identify the direct cause for WhatsHap errors related to indel record parsing, we filtered out all records for indels and structural variants, leaving only SNV records in the VCF.</strong></p>
EGP Mitochondrial Genome Analysis on GIAB Whole-Genome Sequencing Data
<div> <p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the GIAB. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on GIAB</strong>: Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://raw.githubusercontent.com/genome-in-a-bottle/giab_data_indexes/refs/heads/master/AshkenazimTrio/alignment.index.AJtrio_Illumina300X_wgs_novoalign_GRCh37_GRCh38_NHGRI_07282015</code></p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>5eac6ec7d36307aa401fd5441b38a506</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>1151ae74c8e515f4f39bef816bb55d6a</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Variant Tables</td> <td>b3e342fe9827df2e399f5685f84cd4dc</td> </tr> <tr> <td>GIAB</td> <td>Mitochondrial Genome Copy Number</td> <td>3c45c19f76f71b3ccaad155565ead5e4</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div> <p> </p> <p> </p> </div> <h2> </h2>
HG005_Son.R2.fastq.gz (GIAB_ChineseTrio_30X)
<p>This is a sub-sampled version (30X) of ChineseTrio HG005(son)_R2 fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> As the total size of these data amounts to more than 200GB, we also prepared 10X fastq files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG006_Father.R1.fastq.gz (GIAB_ChineseTrio_30X)
<p>This is a sub-sampled version (30X) of ChineseTrio HG006(father)_R1 fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> As the total size of these data amounts to more than 200GB, we also prepared 10X fastq files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG007_Mother.R1.fastq.gz (GIAB_ChineseTrio_30X)
<p>This is a sub-sampled version (30X) of ChineseTrio HG007(mother)_R1 fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> As the total size of these data amounts to more than 200GB, we also prepared 10X fastq files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG006_Father.R2.fastq.gz (GIAB_ChineseTrio_30X)
<p>This is a sub-sampled version (30X) of ChineseTrio HG005(son)_R2 fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> As the total size of these data amounts to more than 200GB, we also prepared 10X fastq files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG007_Mother.R2.fastq.gz (GIAB_ChineseTrio_30X)
<p>This is a sub-sampled version (30X) of ChineseTrio HG007(mother)_R2 fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> As the total size of these data amounts to more than 200GB, we also prepared 10X fastq files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG006_Father (GIAB_ChineseTrio_10X)
<p>This is a sub-sampled version (10X) of ChineseTrio HG006(father) fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> These files are 10X fastq files. Though 30X fastq files are preferable for WGS analysis, we prepared 10X fastq files for reducing the burden of handling the large files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG007_Mother (GIAB_ChineseTrio_10X)
<p>This is a sub-sampled version (10X) of ChineseTrio HG007(mother) fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> These files are 10X fastq files. Though 30X fastq files are preferable for WGS analysis, we prepared 10X fastq files for reducing the burden of handling the large files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
HG005_Son (GIAB_ChineseTrio_10X)
<p>This is a sub-sampled version (10X) of ChineseTrio HG005(son) fastq files published by NIST's Genome in a Bottle project.<br> We created them for practicing variant-call processes.<br> These files are 10X fastq files. Though 30X fastq files are preferable for WGS analysis, we prepared 10X fastq files for reducing the burden of handling the large files.</p> <p>The original data is available under:<a href="https://github.com/genome-in-a-bottle/giab_data_indexes">https://github.com/genome-in-a-bottle/giab_data_indexes</a><br> The original article is below;<br> Wang, YC., Olson, N.D., Deikus, G. et al. High-coverage, long-read sequencing of Han Chinese trio reference samples. Sci Data 6, 91 (2019). <a href="https://doi.org/10.1038/s41597-019-0098-2">https://doi.org/10.1038/s41597-019-0098-2</a><br> We are permitted to published sub-sampled versions by Genome in a Bottle project.</p>
SWaveform resource GIAB HG002 data
<p>This is a part of data associated with SWaveform resource (swaveform.compbio.ru). The data encompasses depth of coverage (DOC) signals from Genome in a Bottle consortia (GIAB) [Zook, Justin M et al. 2016] sample HG002. As the latter provides variant calls and regions for use in benchmarking and validating variant calling pipelines and contains a smaller number of individuals, we envision it to be a demo collection to deploy the SWaveform locally and to test the accompanying toolset.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.