Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3 results for “GATK”

Learn how ShareScore rates datasets ↗
zenodo40/100

Raw BRCA1/2 variants in breast cancer patients and healthy relatives produced with GATK.

<p>Aligned sequencing data is available in the NCBI Sequence Read Archive (SRA, https://www.ncbi.nlm.nih.gov/sra/) under accession SRP095082. Variants were called using GATK HaplotypeCaller (version 3.6). After joint performing joint genotyping multi-sample vcf file was generated. Next, SNPs and indels were extracted into two different vcf files and specific set of filters were applied for each case.</p> <p> </p> <p><strong>File descriptions</strong></p> <p><strong><em>Datasets</em></strong></p> <p><strong>BRCA_SNVs.vcf</strong> - this file contains SNPs called with GATK and hard filters applied. Following filtering options were applied: "QD &lt; 2.0", "FS &gt; 60.0", "MQ &lt; 40.0",  "MQRankSum &lt; -12.5", "ReadPosRankSum &lt; -8.0", "SB &lt; -0.10" , "DP &lt; 10" , "GQ &lt; 30" , and "SOR &gt; 3.0"</p> <p><strong>BRCA_indels.vcf</strong> - This file contains indels called with GATK and hard filters applied. Following filtering options were applied: "QD &lt; 2.0", "FS &gt; 200.0", "ReadPosRankSum &lt; -20.0", "InbreedingCoeff &lt; -0.8", "SOR &gt; 10.0".</p> <p> </p> <p><strong><em>Scripts package (scritps.zip)</em></strong></p> <p>Scripts.zip file contains scripts and supporting files for genotype calling and filtering. </p> <p><strong>raw.variant.caling.sh </strong>– bam files preprocessing, alignment refining and raw genotype calling with HaplotypeCaller.</p> <p><strong>genotyping_and_filtering.sh </strong>– joint genotyping, variant hard filtering and callset refinement.</p> <p><strong>LIST.txt</strong> – supporting file that contains bam filenames containing aligned reads.</p> <p><strong>sample_order.txt</strong> – supporting file for sample renaming.</p> <p> </p> <p><strong><em>Reference files (hg19) used in variant calling scripts</em></strong></p> <p>Reference files can be downloaded from GATK bundle web-site at https://software.broadinstitute.org/gatk/download/bundle.  </p> <p><strong>ucsc.hg19.fasta</strong> - human genome assembly;</p> <p><strong>Mills_and_1000G_gold_standard.indels.hg19.sites.vcf.gz</strong> – set of known indels to be used for local realignment;</p> <p><strong>1000G_phase1.indels.hg19.sites.vcf.gz</strong> – set of known indels to be used for local realignment;</p> <p><strong>dbsnp_138.hg19.vcf.gz</strong> – a recent dbSNP release (build 138); </p> <p><strong>1000G_phase3_v4_20130502.hg19.lifted.sites.vcf</strong> – the latest set from 1000G phase 3 (v4) for genotype refinement.</p> <p> </p>

opencc-by-4.0Dec 2016View details →
zenodo40/100

VCF, pileup, and other files for Betula analyses using ebg & GATK

<p>These are the VCF, pileup, and input files files produced by the Genome Analysis ToolKit, SAMtools, and our own scripts, respectively, used for comparing the genotypes estimated by the GATK UnifiedGenotyper and the new model for allopolyploids that we introduced in our paper. Additional processing of the files was conducting using Python and R for filtering and extracting error values and read counts (scripts available on GitHub: https://github.com/pblischak/polyploid-genotyping).</p> <p><strong>Main files:</strong></p> <ul> <li>pendula-ug-filtered30.vcf: VCF file from GATK with all called variants for <em>Betula</em> <em>pendula</em> (diploid).</li> <li>filtered30-pendula.pileup: SAMtools pileup file for variant sites identified by GATK in <em>B</em>. <em>pendula</em>.</li> <li>pubescens-ug-filtered30.vcf: VCF file from GATK with all called variants for <em>Betula</em> <em>pubescens</em> (allotetraploid).</li> <li>filtered30-pubescens.pileup: SAMtools pileup file for variant sites identified by GATK in <em>B</em>. <em>pubescens</em>.</li> </ul> <p><strong>Processed files:</strong></p> <ul> <li>filtered30-variants.txt: tab delimited file of shared variants between B. pendula and B. pubescens.</li> <li>filtered30-vcf1.vcf: VCF file for B. pendula with variants extracted from filtered30-variants.txt.</li> <li>filtered30-vcf2.vcf: VCF file for B. pubescens with variants extracted from filtered30-variants.txt.</li> </ul> <p><strong>Input files</strong>:</p> <ul> <li>filtered30-pubescens-tot.txt, filtered30-pubescents-alt.txt, filtered30-pubescens-err.txt: input files for running the alloSNP model in ebg.</li> <li>filtered30-pubescens-alloSNP-freqs2.txt, filtered30-pubescens-alloSNP-g1.txt, filtered30-pubescens-alloSNP-g2.txt: output files from the alloSNP model.</li> <li>filtered30-pendula-tot.txt, filtered30-pendula-alt.txt, filtered30-pendula-err.txt: input files for running the hwe model in ebg.</li> <li>filtered30-pendula-hwe-freqs.txt, filtered30-pendula-hwe-genos.txt: output files from the hwe model. The allele frequency estimates here were also used as the reference panel for the alloSNP model.</li> </ul>

opencc-by-4.0Jul 2017View details →
zenodo28/100

JG2.1.0-Based GATK Resource Bundle

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record