Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
61
datasets available to search
ShareScore release 0.9.0
Dataset results
61 results for “variant calling”
PopDel identifies medium-size deletions jointly in tens of thousands of genomes - Variant call sets
<p>This data set contains the variant calls sets generated by different tools for the benchmarks in the paper <a href="https://www.nature.com/articles/s41467-020-20850-5">PopDel identifies medium-size deletions simultaneously in tens of thousands of genomes</a>. It includes the VCFs/BCFs for the following test cases:</p> <ul> <li>Random deletion simulation on up to 1000 chromosome 21 samples</li> <li>1000 Genomes Project deletions inserted into simulated chromosomes 17 to 22 of up to 500 samples</li> <li>HG001 (NA12878)</li> <li>Trio of <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/">HG002</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG003_NA24149_father/NIST_HiSeq_HG003_Homogeneity-12389378/">HG003</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG004_NA24143_mother/NIST_HiSeq_HG004_Homogeneity-14572558/">HG004</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Diversity-Cohort">Polaris Diversity cohort</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Kids-Cohort">Polaris Kids cohort</a></li> </ul> <p>Further, the long and short read reference call sets for HG001 are provided. For HG002 the reference call set and the high confidence regions by the Genome in a Bottle consortium are provided.</p> <p>For details on how the files have been created, please refer to the paper and the script repository on <a href="https://github.com/kehrlab/PopDel-scripts">GitHub</a>.</p>
Hepatocystis alignments and variant calls
<p>This contains BAM files of <em>Hepatocystis</em> reads identified in <em>Papio</em> and <em>Chlorocebus</em> samples mapped to the <em>Hepatocystis</em> reference genome as well as major and minor allele calls for cHEP and pHEP called with ANGSD, both individually and jointly. Nucleotide alignments have been added in the latest version.</p>
Phase I trial of CX-5461, a first-in-class G-quadruplex stabilizer in patients with advanced solid tumors enriched for DNA-repair deficiencies (CCTG IND.231) - Variant Calls
<p>Variant Calls from Phase I trial of CX-5461, a first-in-class G-quadruplex stabilizer in patients with advanced solid tumors enriched for DNA-repair deficiencies (CCTG IND.231)</p> <p>See publication for methodology.</p>
Variant calls for 'Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration'
<p>A VCF of SNP calls used for input to GWAS in https://elifesciences.org/articles/26255</p>
Assembly and variant calling results of strain mixtures of HCMV
<p>The tgz files contain the assembly contigs of 10 different assemblers and variant calling results of 6 callers on the strain mixture dataset of the HCMV virus. </p>
Fastq files for benchmarking somatic variant calling pipelines
<p>The <a href="https://download.imgag.de/public/validation_dataset_somatic/readme.html" target="_blank" rel="noopener">original .bam files</a> were provided by <a href="https://www.medizin.uni-tuebingen.de/de/das-klinikum/mitarbeiter/profil/3377" target="_blank" rel="noopener">Marc Sturm</a> from the <a href="https://www.medizin.uni-tuebingen.de/de/das-klinikum/einrichtungen/institute/medizinische-genetik-und-angewandte-genomik" target="_blank" rel="noopener">Institut für Medizinische Genetik und Angewandte Genomik at the University Clinic Tübingen.</a></p> <p>The files were transformed into .fq.gz files using the <a href="https://github.com/nf-core/bamtofastq" target="_blank" rel="noopener">nf-core/bamtofastq pipeline.</a> All information on the pipeline run can be found in the <a href="../api/records/10805134/draft/files/execution_report_2024-03-11_11-44-17.html/content" target="_blank" rel="noopener noreferrer">execution_report_2024-03-11_11-44-17.html</a>. The FASTQ files for the normal sample were directly uploaded to <a href="https://osf.io/cduyq/files/onedrive">the corresponding osf project</a>.</p> <p> </p> <p> </p>
Training data for 'Somatic variant calling' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial that demonstrates identification of somatic and germline variants from tumor and normal sample pairs.</p>
Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.
<p>Ebola genome is manually mutated to contain non-structural as well as structural variants. One ebola genome contains non-structural variants - 10 SNPs, 10 INDELs, 05 TRANSLOCATIONs, 05 INSERSIONs and their reverse complements. Similarly, other two set of mutated genomes contain structural variants. Each set contains seven mutated ebola genome each one for large deletion, insertion, duplication, translocation, inversion, complex variant1 (consecutive three mutations - insertion, duplication and deletion) and complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion.</p> <p> </p>
Performance and agreement between WGS variant calling pipelines used for bovine tuberculosis control: towards international standardisation
<p>This repository contains the simulated genomes and FASTQ files used for the analyses reported in the publication: <strong>Performance and agreement between WGS variant calling pipelines used for bovine tuberculosis control: towards international standardisation</strong>.</p> <p>This dataset is based on previously published data: <a href="https://doi.org/10.1099/mgen.0.000388">https://doi.org/10.1099/mgen.0.000388</a></p> <p>Processing scripts can be found in: <a href="https://github.com/Viloleal/bTB-pipeline-comparison-data-and-tools">https://github.com/Viloleal/bTB-pipeline-comparison-data-and-tools</a></p>
spindle cell variant diffuse large B-cell lymphoma (NGS annotation file; high confidence calls) hematolrep-2136295
<p>Diffuse large B-cell lymphoma with spindle cell morphology is a rare variant. We present the case of a 74-year-old male who initially presented with a right supraclavicular (lymph) node enlargement. Histological analysis showed a proliferation of spindle-shaped cells with narrow cytoplasms. An immunohistochemical panel was used to exclude other tumors, such as melanoma, carcinoma, and sarcoma. The lymphoma was characterized by a cell-of-origin subtype of germinal center B-cell-like (GCB) based on Hans’ classifier (CD10-negative, BCL6-positive, and MUM1-negative); EBER negativity, and the absence of BCL2, BCL6, and MYC rearrangements. Mutational profiling using a custom panel of 168 genes associated with aggressive B-cell lymphomas confirmed mutations in ACTB, ARID1B, DUSP2, DTX1, HLA-B, PTEN, and TNFRSF14. Based on the LymphGen 1.0 classification tool, this case had an ST2 subtype prediction. The immune microenvironment was characterized by moderate infiltration of M2-like tumor-associated macrophages (TMAs) with positivity of CD163, CSF1R, CD85A (LILRB3), and PD-L1; moderate PD-1 positive T cells, and low FOXP3 regulatory T lymphocytes (Tregs). Immunohistochemical expression of PTX3 and TNFRSF14 was absent. Interestingly, the lymphoma cells were positive for HLA-DP-DR, IL-10, and RGS1, which are markers associated with poor prognosis in DLBCL. The patient was treated with R-CHOP therapy, and achieved a metabolically complete response.</p> <p>Carreras J, Kikuti YY, Miyaoka M, Hiraiwa S, Tomita S, Ikoma H, Kondo Y, Ito A, Nagase S, Miura H, Roncador G, Colomo L, Hamoudi R, Campo E, Nakamura N. Mutational Profile and Pathological Features of a Case of Interleukin-10 and RGS1-Positive Spindle Cell Variant Diffuse Large B-Cell Lymphoma. <em>Hematology Reports</em>. 2023; 15(1):188-200. https://doi.org/10.3390/hematolrep15010020</p>
Variant calling in the Goldilocks Zone: how reference genome choice and read mapping stringency impact heterozygosity estimates and phylogenetic analyses
Open the record for dataset details and reuse information.
Variant call file for mountain yellow-legged frog (MYLF) selection analysis
Open the record for dataset details and reuse information.
Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.
<p>Ebola genome is manually mutated to contain non-structural variants - 10 SNPs, 10 INDELs, 05 TRANSLOCATIONs, 05 INSERSIONs and their reverse complements. Paired-end illumina RNASEQ reads are simulated in fastq format. </p>
Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.
<p>Ebola genome is manually mutated to contain structural variants. There are seven mutated ebola genomes each one for two large deletions, insertions, duplications, translocations, inversions, one complex variant1 (consecutive three mutations - insertion, duplication and deletion) and one complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion. Paired-end illumine RNASEQ reads are simulated in fastq format.</p>
Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.
<p>Ebola genome is manually mutated to contain structural variants. There are seven mutated ebola genomes each one for large deletion, insertion, duplication, translocation, inversion, complex variant1 (consecutive three mutations - insertion, duplication and deletion) and complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion. Paired-end illumine RNASEQ reads are simulated in fastq format.</p>
Experimental results of "Managing variant calling datasets the big data way"
<p>Tomatula was demonstrated for retrieving the allele frequencies for a given region in the data from Aflitos et al (2014). We developed scripts to retrieve allele frequencies, either from the VCF file storage or Apache Parquet. We executed a series of experiments, querying for a region of 2000 bases in the file of chromosome 6, that corresponds to the approximate length of a gene. We compared both storage formats (VCF files and Parquet), two input sizes (104 and 1144 individuals), different cluster sizes varying between 2 and 150 executor nodes, and HDFS replication factor was set to 3, 5, 7, and 9, in order to examine four main factors that<br> can affect the performance of a Big Data cluster: (a) the storage format, (b) the size of the input files, (c) the number of computing nodes of the cluster, and (d) the replication factor of HDFS. The block size of the HDFS was kept at the default value of 128MB. All experiments were executed five times and the detailed results are provided here, along with a script that produces the corresponding figures.</p>
Variant calls for HG001-07
Open the record for dataset details and reuse information.
Calling structural variants with confidence from short-read data in wild bird populations
<p>Comprehensive characterisation of structural variation in natural populations has only become feasible in the last decade. To investigate the population genomic nature of structural variation (SV), reproducible and high-confidence SV callsets are first required. We created a population-scale reference of the genome-wide landscape of structural variation across 33 Nordic house sparrows (<em>Passer domesticus</em>) individuals. To produce a consensus callset across all samples using short-read data, we compare heuristic-based quality filtering and visual curation (Samplot/PlotCritic and Samplot-ML) approaches. We demonstrate that curation of SVs is important for reducing putative false positives and that the time invested in this step outweighs the potential costs of analysing short-read discovered SV datasets that include many potential false positives. We find that even a lenient manual curation strategy (e.g. applied by a single curator) can reduce the proportion of putative false positives by up to 80%, thus enriching the proportion of high-confidence variants. Crucially, in applying a lenient manual curation strategy with a single curator, nearly all (>99%) variants rejected as putative false positives were also classified as such by a more stringent curation strategy using three additional curators. Furthermore, variants rejected by manual curation failed to reflect the expected population structure from SNPs, whereas variants passing curation did. Combining heuristic-based quality-filtering with rapid manual curation of structural variants in short-read data can therefore become a time- and cost-effective first step for functional and population genomic studies requiring high-confidence SV callsets.</p>
LYCEUM: Learning to call copy number variants on low coverage ancient genomes
<p>Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduce additional noise into sequencing data; and finally, (iii) the typically low coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high coverage read-depth signals, underperform under such conditions. To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high- confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection.</p>
GHGA_sarek3_variant_calling_Agilent200M_WES
<p>Using the sarek pipeline default values. Aligned to HG38 using bwa, and dragmap. Variants are called using either haplotypecaller, strelka, deepvariant, or freebayes. Data input was Agilent 200M WES reads. Sample A006850052 (GIAB: HG001)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.