Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

61

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

61 results for “variant calling”

Learn how ShareScore rates datasets ↗
zenodo36/100

Nitzschia sp. Nitz4 variant calling

<p>BAM file of read alignments used to call variants in the Nitzschia sp. Nitz4 genome. The BAM file was processed to mark duplicates and add read group names.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

CWL run of Somatic Variant Calling Workflow (CWLProv 0.5.0 Research Object)

<p>The somatic variant calling workflow included in this case study is designed by <a href="http://bcb.io/">Blue Collar Bioinformatics (bcbio)</a>, a community-driven initiative to develop best-practice pipelines for variant calling, RNA-seq and small RNA analysis workflows. According to the documentation, the goal of this project is to facilitate the automated analysis of high throughput data by making the resources quantifiable, analyzable, scalable, accessible and reproducible.</p> <p>All the underlying tools are containerized, facilitating software use in the workflow. The somatic variant calling workflow defined in CWL is available on GitHub and equipped with a well defined test dataset.</p> <p>This dataset folder is a CWLProv Research Object that captures the Common Workflow Language execution provenance, see <a href="https://w3id.org/cwl/prov/0.5.0">https://w3id.org/cwl/prov/0.5.0</a> or use <a href="https://pypi.org/project/cwlprov/">https://pypi.org/project/cwlprov/</a> to explore</p> <p><strong>Steps to reproduce</strong></p> <p>To build the research object again, use Python 3 on macOS. Built on:</p> <ul> <li>Processor 2.8GHz Intel Core i7</li> <li>Memory: 16GB</li> <li>OS: macOS High Sierra, Version 10.13.3</li> <li>Storage: 250GB</li> </ul> <p>To run the workflow:<br> &nbsp;</p> <pre><code class="language-bash">pip3 install cwltool==1.0.20180912090223 git clone https://github.com/FarahZKhan/bcbio_test_cwlprov cd bcbio_test_cwlprov/somatic/somatic-workflow/ cwltool --provenance somaticwf_0.5.0_mac main-somatic.cwl main-somatic-samples.json</code></pre> <p>To package the research object:<br> &nbsp;</p> <pre><code class="language-bash">zip -r somaticwf_0.5.0_mac.zip somaticwf_0.5.0_mac/ sha256sum somaticwf_0.5.0_mac.zip &gt; somaticwf_0.5.0_mac.zip.sha256</code></pre> <p>The <a href="https://github.com/FarahZKhan/bcbio_test_cwlprov">cloned git repository</a> is a fork of <a href="https://github.com/bcbio/test_bcbio_cwl">https://github.com/bcbio/test_bcbio_cwl</a>. It was obtained using:</p> <pre><code class="language-bash">wget -O test_bcbio_cwl.tar.gz https://github.com/bcbio/test_bcbio_cwl/archive/master.tar.gz</code></pre> <p>The content is from an archived version from the documentation here: <a href="https://bcbio-nextgen.readthedocs.io/en/latest/contents/cwl.html#install-bcbio-vm-with-containers">https://bcbio-nextgen.readthedocs.io/en/latest/contents/cwl.html#install-bcbio-vm-with-containers</a></p>

openmit-licenseDec 2017View details →
zenodo36/100

Human variant calling data

<p>Short read data from the exome of chromosome 22 of a single human individual.&nbsp;There are one million 76bp reads in the dataset, produced on an Illumina GAIIx from exome-enriched DNA. This data was generated as part of the 1000 Genomes project.</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

NGSAP-VC : Genomic Variant Calling as an Installable GALAXY Workflow Using NGS data.

<p>Implementation of genomic variants calling as an installable GALAXY workflows using NGS data. Repository contains two separate sets of simulated ebola test data. One for SNPs and INDELs calling and another for Structural Variants calling.</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

The result of variants calling with the Japonica reference

<p>We used two major varities of the rice genome, Indica and Japonica, to build different variant calling models that differ in the composition of samples from the two varities.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

The result of variants calling with the Indica reference

<p>We used two major varities of the rice genome, Indica and Japonica, to build different variant calling models that differ in the composition of samples from the two varities.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

From Forensics to Clinical Research: Expanding the Variant Calling Pipeline for the Precision ID mtDNA Whole Genome Panel

<p>In this dataset we provide the 1000 Genomes Project&#39;s samples processed in our manuscript:</p> <p>Cortes-Figueiredo, F.; Carvalho, F.S.; Fonseca, A.C.; Paul, F.; Ferro, J.M.; Sch&ouml;nherr, S.; Weissensteiner, H.; Morais, V.A. From Forensics to Clinical Research: Expanding the Variant Calling Pipeline for the Precision ID mtDNA Whole Genome Panel. <em>Int. J. Mol. Sci</em>. <strong>2021</strong>, <em>22</em>, 12031. <a href="https://doi.org/10.3390/ijms222112031">https://doi.org/10.3390/ijms222112031</a><em>.</em></p> <p><strong>Abstract</strong></p> <p>Despite a multitude of methods for the sample preparation, sequencing, and data analysis of mitochondrial DNA (mtDNA), the demand for innovation remains, particularly in comparison with nuclear DNA (nDNA) research. The Applied Biosystems&trade; Precision ID mtDNA Whole Genome Panel (Thermo Fisher Scientific, USA) is an innovative library preparation kit suitable for degraded samples and low DNA input. However, its bioinformatic processing occurs in the enterprise Ion Torrent Suite&trade; Software (TSS), yielding BAM files aligned to an unorthodox version of the revised Cambridge Reference Sequence (rCRS), with a heteroplasmy threshold level of 10%. Here, we present an alternative customizable pipeline, the PrecisionCallerPipeline (PCP), for processing samples with the correct rCRS output after Ion Torrent sequencing with the Precision ID library kit. Using 18 samples (3 original samples and 15 mixtures) derived from the 1000 Genomes Project, we achieved overall improved performance metrics in comparison with the proprietary TSS, with optimal performance at a 2.5% heteroplasmy threshold. We further validated our findings with 50 samples from an ongoing independent cohort of stroke patients, with PCP finding 98.31% of TSS&rsquo;s variants (TSS found 57.92% of PCP&rsquo;s variants), with a significant correlation between the variant levels of variants found with both pipelines.</p> <p><br> Please refer to our the github page <a href="https://github.com/filcfig/PCP.git">filcfig/PCP</a>, for more details on running the PrecisionCalllerPipeline.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Supporting data for "Software pipelines for RNA-Seq, ChIP-Seq and Germline Variant calling analyses in Common Workflow Language (CWL)"

<p>Datasets produced during the validation of CWL-based pipelines, designed for the analysis of data from&nbsp;RNA-Seq, ChIP-Seq and germline variant calling experiments. Specifically, the workflows were tested using publicly available High-throughput (HTS) data from published studies&nbsp;on Chronic Lymphocytic Leukemia (CLL) (accession numbers: E-MTAB-6962, GSE115772) and Genome in a Bottle (GIAB) project samples (accession numbers: SRR6794144, SRR22476789, SRR22476790, SRR22476791).</p> <p>The supporting data include:</p> <ul> <li>Differential transcript and gene expression results produced during the analysis with the CWL-based RNA-Seq pipeline</li> <li>Bigwig and narrowPeak files, differential binding results, table of consensus peaks and read counts of EZH2 and H3K27me3,&nbsp;produced during the analysis with the CWL-based ChIP-Seq pipeline</li> <li>VCF files containing the detected and filtered variants, along with the respective&nbsp;hap.py () results regarding&nbsp;comparisons&nbsp;against the GIAB golden standard truth sets for both CWL-based&nbsp;germline variant calling pipelines</li> </ul>

opencc-by-4.0Jul 2023View details →
dryad36/100

Calling structural variants with confidence from short-read data in wild bird populations

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad32/100

Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa

<p>A comprehensive understanding of the degree to which genomic variation is maintained by selection versus drift and gene flow is lacking in many important species such as <em>Cannabis</em> <em>sativa </em>(<em>C. sativa</em>), one of the oldest known crops to be cultivated by humans worldwide. We generated whole genome resequencing data across diverse samples of feralized (escaped domesticated lineages) and domesticated lineages of <em>C. sativa</em>. We performed analyses to examine population structure, and genome wide scans for FST, balancing selection, and positive selection. Our analyses identified evidence for sub-population structure and further support the Asian origin hypothesis of this species. Feral plants sourced from the U.S. exhibited broad regions on chromosomes 4 and 10 with high <span>𝐹̅</span>ST which may indicate chromosomal inversions maintained at high frequency in this sub-population. Both our balancing and positive selection analyses identified loci that may reflect differential selection for traits favored by natural selection and artificial selection in feral versus domesticated sub-populations. In the U.S. feral sub-population, we found six loci related to stress response under balancing selection and one gene involved in disease resistance under positive selection, suggesting local adaptation to new climates and biotic interactions. In the marijuana sub-population, we identified the gene <em>SMALLER TRICHOMES</em> <em>WITH VARIABLE BRANCHES 2 </em>to be under positive selection which suggests artificial selection for increased tetrahydrocannabinol yield. Overall the data generated, and results obtained from our study help to form a better understanding of the evolutionary history in <em>C. sativa</em>.</p>

opencc-zeroAug 2022View details →
zenodo32/100

Training material for Calling variants in non-diploid systems

<p>The majority of life on Earth is non-diploid and represented by prokaryotes, viruses and their derivatives such as our own mitochondria or plant&rsquo;s chloroplasts. In non-diploid systems allele frequencies can range anywhere between 0 and 100% and there could be multiple (not just two) alleles per locus. The main challenge associated with non-diploid variant calling is the difficulty in distinguishing between sequencing noise (abundant in all NGS platforms) and true low frequency variants.&nbsp;</p>

opencc-by-4.0May 2018View details →
zenodo32/100

Variant Calling Files supporting leptospira analysis

<p>This data set contains the results of a variant calling analysis performed on 20 samples named X_SX for sample 1 to 16 and named B3288_2 B3288_8 B3288_10 and B3288_16 for samples named 17 to 20.</p> <p>&nbsp;</p> <p>All variant calling analysis were performed with Sequana variant calling pipeline (https://github.com/sequana/variant_calling) but only the final VCF files are provided in this dataset. One VCF file per sample and per reference genome used. There were 273 core genomes used; Therefore we have here 273 time 20 VCF files. In the summary directory one can also find the CSV files for SNPs and INDELs that were extracted from the VCF files using several filtering (depth of 10, strand balance &gt;0.2, min frequency &gt;=0.5).</p> <p>More information and notebooks generating and using those files can be found here:&nbsp;https://github.com/biomics-pasteur-fr/manuscript_capture_leptospira/&nbsp; &nbsp;A tagged version of this github repository is provided as version 1 : manuscript_capture_leptospira-1.tar.gz</p>

opencc-by-4.0Jan 2023View details →
dryad32/100

Raw genotyped total called structural variant (SV)

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad32/100

Variant Call File (VCF) for Genome-wide polymorphism and genic selection in feral and domesticated lineages of Cannabis sativa

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad28/100

Data from: De novo sequencing and variant calling with nanopores using PoreSeq

The accuracy of sequencing single DNA molecules with nanopores is continually improving, but de novo genome sequencing and assembly using only nanopore data remain challenging. Here we describe PoreSeq, an algorithm that identifies and corrects errors in nanopore sequencing data and improves the accuracy of de novo genome assembly with increasing coverage depth. The approach relies on modeling the possible sources of uncertainty that occur as DNA transits through the nanopore and finds the sequence that best explains multiple reads of the same region. PoreSeq increases nanopore sequencing read accuracy of M13 bacteriophage DNA from 85% to 99% at 100× coverage. We also use the algorithm to assemble Escherichia coli with 30× coverage and the λ genome at a range of coverages from 3× to 50×. Additionally, we classify sequence variants at an order of magnitude lower coverage than is possible with existing methods.

opencc-zeroDec 2014View details →
dryad28/100

VCF of structural variant calls of Nanopore data aligned to dm6 reference genome

<p>Heterozygous chromosome inversions suppress meiotic crossover (CO) formation within an inversion, potentially because they lead to gross chromosome rearrangements that produce inviable gametes. Interestingly, COs are also severely reduced in regions nearby but outside of inversion breakpoints even though COs in these regions do not result in rearrangements. Our mechanistic understanding of why COs are suppressed outside of inversion breakpoints is limited by a lack of data on the frequency of noncrossover gene conversions (NCOGCs) in these regions. To address this critical gap, we mapped the location and frequency of rare CO and NCOGC events that occurred outside of the <em>dl</em>-<em>49</em> <em>chrX</em> inversion in <em>D</em>. <em>melanogaster</em>. We created full-sibling wildtype and inversion stocks and recovered COs and NCOGCs in the syntenic regions of both stocks, allowing us to directly compare rates and distributions of recombination events. We show that COs are completely suppressed within 500 kb of inversion breakpoints, are severely reduced within 2 Mb of an inversion breakpoint, and increase above wildtype levels 2–4 Mb from the breakpoint. We find that NCOGCs occur evenly throughout the chromosome and, importantly, occur at wild-type levels near inversion breakpoints. We propose a model in which COs are suppressed by inversion breakpoints in a distance-dependent manner through mechanisms that influence DNA double-strand break repair outcome but not double-strand break location or frequency. We suggest that subtle changes in the synaptonemal complex and chromosome pairing might lead to unstable interhomolog interactions during recombination that permits NCOGC formation but not CO formation.</p>

opencc-zeroMar 2023View details →
zenodo28/100

Dataset for the manuscript "In silico evaluation of variant calling methods for bacterial whole genome sequencing"

<p>Input data and associated analysis code for reproducing results reported in the manuscript.</p>

opencc-by-4.0Jun 2023View details →
dryad28/100

Data from: De novo sequencing and variant calling with nanopores using PoreSeq

Open the record for dataset details and reuse information.

publicSep 2015View details →
dryad28/100

VCF of structural variant calls of Nanopore data aligned to dm6 reference genome

Open the record for dataset details and reuse information.

publicMar 2023View details →
geo24/100

A practical evaluation of alignment algorithms for RNA variant calling analysis

GEO Series GSE110114. Homo sapiens. 13 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record