Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
61
datasets available to search
ShareScore release 0.9.0
Dataset results
61 results for “variant calling”
Nitzschia sp. Nitz4 variant calling
<p>BAM file of read alignments used to call variants in the Nitzschia sp. Nitz4 genome. The BAM file was processed to mark duplicates and add read group names.</p>
CWL run of Somatic Variant Calling Workflow (CWLProv 0.5.0 Research Object)
<p>The somatic variant calling workflow included in this case study is designed by <a href="http://bcb.io/">Blue Collar Bioinformatics (bcbio)</a>, a community-driven initiative to develop best-practice pipelines for variant calling, RNA-seq and small RNA analysis workflows. According to the documentation, the goal of this project is to facilitate the automated analysis of high throughput data by making the resources quantifiable, analyzable, scalable, accessible and reproducible.</p> <p>All the underlying tools are containerized, facilitating software use in the workflow. The somatic variant calling workflow defined in CWL is available on GitHub and equipped with a well defined test dataset.</p> <p>This dataset folder is a CWLProv Research Object that captures the Common Workflow Language execution provenance, see <a href="https://w3id.org/cwl/prov/0.5.0">https://w3id.org/cwl/prov/0.5.0</a> or use <a href="https://pypi.org/project/cwlprov/">https://pypi.org/project/cwlprov/</a> to explore</p> <p><strong>Steps to reproduce</strong></p> <p>To build the research object again, use Python 3 on macOS. Built on:</p> <ul> <li>Processor 2.8GHz Intel Core i7</li> <li>Memory: 16GB</li> <li>OS: macOS High Sierra, Version 10.13.3</li> <li>Storage: 250GB</li> </ul> <p>To run the workflow:<br> </p> <pre><code class="language-bash">pip3 install cwltool==1.0.20180912090223 git clone https://github.com/FarahZKhan/bcbio_test_cwlprov cd bcbio_test_cwlprov/somatic/somatic-workflow/ cwltool --provenance somaticwf_0.5.0_mac main-somatic.cwl main-somatic-samples.json</code></pre> <p>To package the research object:<br> </p> <pre><code class="language-bash">zip -r somaticwf_0.5.0_mac.zip somaticwf_0.5.0_mac/ sha256sum somaticwf_0.5.0_mac.zip > somaticwf_0.5.0_mac.zip.sha256</code></pre> <p>The <a href="https://github.com/FarahZKhan/bcbio_test_cwlprov">cloned git repository</a> is a fork of <a href="https://github.com/bcbio/test_bcbio_cwl">https://github.com/bcbio/test_bcbio_cwl</a>. It was obtained using:</p> <pre><code class="language-bash">wget -O test_bcbio_cwl.tar.gz https://github.com/bcbio/test_bcbio_cwl/archive/master.tar.gz</code></pre> <p>The content is from an archived version from the documentation here: <a href="https://bcbio-nextgen.readthedocs.io/en/latest/contents/cwl.html#install-bcbio-vm-with-containers">https://bcbio-nextgen.readthedocs.io/en/latest/contents/cwl.html#install-bcbio-vm-with-containers</a></p>
Human variant calling data
<p>Short read data from the exome of chromosome 22 of a single human individual. There are one million 76bp reads in the dataset, produced on an Illumina GAIIx from exome-enriched DNA. This data was generated as part of the 1000 Genomes project.</p>
NGSAP-VC : Genomic Variant Calling as an Installable GALAXY Workflow Using NGS data.
<p>Implementation of genomic variants calling as an installable GALAXY workflows using NGS data. Repository contains two separate sets of simulated ebola test data. One for SNPs and INDELs calling and another for Structural Variants calling.</p>
The result of variants calling with the Japonica reference
<p>We used two major varities of the rice genome, Indica and Japonica, to build different variant calling models that differ in the composition of samples from the two varities. </p>
The result of variants calling with the Indica reference
<p>We used two major varities of the rice genome, Indica and Japonica, to build different variant calling models that differ in the composition of samples from the two varities.</p>
From Forensics to Clinical Research: Expanding the Variant Calling Pipeline for the Precision ID mtDNA Whole Genome Panel
<p>In this dataset we provide the 1000 Genomes Project's samples processed in our manuscript:</p> <p>Cortes-Figueiredo, F.; Carvalho, F.S.; Fonseca, A.C.; Paul, F.; Ferro, J.M.; Schönherr, S.; Weissensteiner, H.; Morais, V.A. From Forensics to Clinical Research: Expanding the Variant Calling Pipeline for the Precision ID mtDNA Whole Genome Panel. <em>Int. J. Mol. Sci</em>. <strong>2021</strong>, <em>22</em>, 12031. <a href="https://doi.org/10.3390/ijms222112031">https://doi.org/10.3390/ijms222112031</a><em>.</em></p> <p><strong>Abstract</strong></p> <p>Despite a multitude of methods for the sample preparation, sequencing, and data analysis of mitochondrial DNA (mtDNA), the demand for innovation remains, particularly in comparison with nuclear DNA (nDNA) research. The Applied Biosystems™ Precision ID mtDNA Whole Genome Panel (Thermo Fisher Scientific, USA) is an innovative library preparation kit suitable for degraded samples and low DNA input. However, its bioinformatic processing occurs in the enterprise Ion Torrent Suite™ Software (TSS), yielding BAM files aligned to an unorthodox version of the revised Cambridge Reference Sequence (rCRS), with a heteroplasmy threshold level of 10%. Here, we present an alternative customizable pipeline, the PrecisionCallerPipeline (PCP), for processing samples with the correct rCRS output after Ion Torrent sequencing with the Precision ID library kit. Using 18 samples (3 original samples and 15 mixtures) derived from the 1000 Genomes Project, we achieved overall improved performance metrics in comparison with the proprietary TSS, with optimal performance at a 2.5% heteroplasmy threshold. We further validated our findings with 50 samples from an ongoing independent cohort of stroke patients, with PCP finding 98.31% of TSS’s variants (TSS found 57.92% of PCP’s variants), with a significant correlation between the variant levels of variants found with both pipelines.</p> <p><br> Please refer to our the github page <a href="https://github.com/filcfig/PCP.git">filcfig/PCP</a>, for more details on running the PrecisionCalllerPipeline.</p>
Supporting data for "Software pipelines for RNA-Seq, ChIP-Seq and Germline Variant calling analyses in Common Workflow Language (CWL)"
<p>Datasets produced during the validation of CWL-based pipelines, designed for the analysis of data from RNA-Seq, ChIP-Seq and germline variant calling experiments. Specifically, the workflows were tested using publicly available High-throughput (HTS) data from published studies on Chronic Lymphocytic Leukemia (CLL) (accession numbers: E-MTAB-6962, GSE115772) and Genome in a Bottle (GIAB) project samples (accession numbers: SRR6794144, SRR22476789, SRR22476790, SRR22476791).</p> <p>The supporting data include:</p> <ul> <li>Differential transcript and gene expression results produced during the analysis with the CWL-based RNA-Seq pipeline</li> <li>Bigwig and narrowPeak files, differential binding results, table of consensus peaks and read counts of EZH2 and H3K27me3, produced during the analysis with the CWL-based ChIP-Seq pipeline</li> <li>VCF files containing the detected and filtered variants, along with the respective hap.py () results regarding comparisons against the GIAB golden standard truth sets for both CWL-based germline variant calling pipelines</li> </ul>
Calling structural variants with confidence from short-read data in wild bird populations
Open the record for dataset details and reuse information.
Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa
<p>A comprehensive understanding of the degree to which genomic variation is maintained by selection versus drift and gene flow is lacking in many important species such as <em>Cannabis</em> <em>sativa </em>(<em>C. sativa</em>), one of the oldest known crops to be cultivated by humans worldwide. We generated whole genome resequencing data across diverse samples of feralized (escaped domesticated lineages) and domesticated lineages of <em>C. sativa</em>. We performed analyses to examine population structure, and genome wide scans for FST, balancing selection, and positive selection. Our analyses identified evidence for sub-population structure and further support the Asian origin hypothesis of this species. Feral plants sourced from the U.S. exhibited broad regions on chromosomes 4 and 10 with high <span>𝐹̅</span>ST which may indicate chromosomal inversions maintained at high frequency in this sub-population. Both our balancing and positive selection analyses identified loci that may reflect differential selection for traits favored by natural selection and artificial selection in feral versus domesticated sub-populations. In the U.S. feral sub-population, we found six loci related to stress response under balancing selection and one gene involved in disease resistance under positive selection, suggesting local adaptation to new climates and biotic interactions. In the marijuana sub-population, we identified the gene <em>SMALLER TRICHOMES</em> <em>WITH VARIABLE BRANCHES 2 </em>to be under positive selection which suggests artificial selection for increased tetrahydrocannabinol yield. Overall the data generated, and results obtained from our study help to form a better understanding of the evolutionary history in <em>C. sativa</em>.</p>
Training material for Calling variants in non-diploid systems
<p>The majority of life on Earth is non-diploid and represented by prokaryotes, viruses and their derivatives such as our own mitochondria or plant’s chloroplasts. In non-diploid systems allele frequencies can range anywhere between 0 and 100% and there could be multiple (not just two) alleles per locus. The main challenge associated with non-diploid variant calling is the difficulty in distinguishing between sequencing noise (abundant in all NGS platforms) and true low frequency variants. </p>
Variant Calling Files supporting leptospira analysis
<p>This data set contains the results of a variant calling analysis performed on 20 samples named X_SX for sample 1 to 16 and named B3288_2 B3288_8 B3288_10 and B3288_16 for samples named 17 to 20.</p> <p> </p> <p>All variant calling analysis were performed with Sequana variant calling pipeline (https://github.com/sequana/variant_calling) but only the final VCF files are provided in this dataset. One VCF file per sample and per reference genome used. There were 273 core genomes used; Therefore we have here 273 time 20 VCF files. In the summary directory one can also find the CSV files for SNPs and INDELs that were extracted from the VCF files using several filtering (depth of 10, strand balance >0.2, min frequency >=0.5).</p> <p>More information and notebooks generating and using those files can be found here: https://github.com/biomics-pasteur-fr/manuscript_capture_leptospira/ A tagged version of this github repository is provided as version 1 : manuscript_capture_leptospira-1.tar.gz</p>
Raw genotyped total called structural variant (SV)
Open the record for dataset details and reuse information.
Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa
Open the record for dataset details and reuse information.
Data from: De novo sequencing and variant calling with nanopores using PoreSeq
The accuracy of sequencing single DNA molecules with nanopores is continually improving, but de novo genome sequencing and assembly using only nanopore data remain challenging. Here we describe PoreSeq, an algorithm that identifies and corrects errors in nanopore sequencing data and improves the accuracy of de novo genome assembly with increasing coverage depth. The approach relies on modeling the possible sources of uncertainty that occur as DNA transits through the nanopore and finds the sequence that best explains multiple reads of the same region. PoreSeq increases nanopore sequencing read accuracy of M13 bacteriophage DNA from 85% to 99% at 100× coverage. We also use the algorithm to assemble Escherichia coli with 30× coverage and the λ genome at a range of coverages from 3× to 50×. Additionally, we classify sequence variants at an order of magnitude lower coverage than is possible with existing methods.
VCF of structural variant calls of Nanopore data aligned to dm6 reference genome
<p>Heterozygous chromosome inversions suppress meiotic crossover (CO) formation within an inversion, potentially because they lead to gross chromosome rearrangements that produce inviable gametes. Interestingly, COs are also severely reduced in regions nearby but outside of inversion breakpoints even though COs in these regions do not result in rearrangements. Our mechanistic understanding of why COs are suppressed outside of inversion breakpoints is limited by a lack of data on the frequency of noncrossover gene conversions (NCOGCs) in these regions. To address this critical gap, we mapped the location and frequency of rare CO and NCOGC events that occurred outside of the <em>dl</em>-<em>49</em> <em>chrX</em> inversion in <em>D</em>. <em>melanogaster</em>. We created full-sibling wildtype and inversion stocks and recovered COs and NCOGCs in the syntenic regions of both stocks, allowing us to directly compare rates and distributions of recombination events. We show that COs are completely suppressed within 500 kb of inversion breakpoints, are severely reduced within 2 Mb of an inversion breakpoint, and increase above wildtype levels 2–4 Mb from the breakpoint. We find that NCOGCs occur evenly throughout the chromosome and, importantly, occur at wild-type levels near inversion breakpoints. We propose a model in which COs are suppressed by inversion breakpoints in a distance-dependent manner through mechanisms that influence DNA double-strand break repair outcome but not double-strand break location or frequency. We suggest that subtle changes in the synaptonemal complex and chromosome pairing might lead to unstable interhomolog interactions during recombination that permits NCOGC formation but not CO formation.</p>
Dataset for the manuscript "In silico evaluation of variant calling methods for bacterial whole genome sequencing"
<p>Input data and associated analysis code for reproducing results reported in the manuscript.</p>
Data from: De novo sequencing and variant calling with nanopores using PoreSeq
Open the record for dataset details and reuse information.
VCF of structural variant calls of Nanopore data aligned to dm6 reference genome
Open the record for dataset details and reuse information.
A practical evaluation of alignment algorithms for RNA variant calling analysis
GEO Series GSE110114. Homo sapiens. 13 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.