Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
43
datasets available to search
ShareScore release 0.7.1
Dataset results
43 results for “Indel”
SNP and indel discovery and genotyping in next-generation sequencing data
<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels >50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Semi-simulated dataset of indel calling tool evaluation
<p>The semi-simulated datasets used in Ning Wang, et al. </p> <p>100bp_5X_500_Venter_read1.fq.gz & 100bp_5X_500_Venter_read2.fq.gz: 5X coverage, 100bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>100bp_30X_500_Venter_read1.fq.gz & 100bp_30X_500_Venter_read2.fq.gz: 30X coverage, 100bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>250bp_30X_500_Venter_read1.fq.gz & 250bp_30X_500_Venter_read2.fq.gz: 30X coverage, 250bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>100bp_60X_500_Venter_read1.fq.gz & 100bp_60X_500_Venter_read2.fq.gz: 60X coverage, 100bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>Haplotype_1_no_gap.fa & Haplotype_2_no_gap.fa: the semi-simulated dipoid human genome hg19 chromosome 1 and chromosome 2, including HuRef indels used in Ning Wang, et al. In order to use ART (fastq simulator) to generate simulated fastq fiels, the gaps (Ns) of genome were removed.</p> <p>chr1_chr2_variants_truthset.txt: types, position, size and genotype of HuRef indels used in Ning Wang, et al. </p>
Filtered and annotated SNV and indel variants in the PC3 and LNCaP human prostate cancer cell lines
<p>150bp paired-end reads (insert size 350bp) were obtained using the Illumina HiSeqX sequencer. Samtools v1.3.1 mpileup and bcftools were used to interrogate indexed BAM files, from whole-genome reads aligned to human reference genome GRCh38 build 82, and generate a VCF (Variant Call Format) file of single nucleotide variants (SNVs) and short indel variants. Variants private, or unique to a particular cell line, or shared by both were next identified. Variants (likely to be common germline variants) present in HapMap, 1000 genomes phase 3 (2,504 human genomes), and the National Heart Lung and Blood Institute’s Exome Sequencing Project (ESP) (bundled variant data file available at https://goo.gl/mEogvD) were excluded. Variant files (VCF) were filtered using SnpSift with the following parameters: 'QUAL \textgreater= 200 \&\& DP \textgreater= 30', where QUAL denotes minimum variance confidence and DP total depth threshold. Filtered variants were annotated using SnpEff v4.3g. Please see https://github.com/sciseim/PCaWGS for associated scripts.</p> <p> </p>
ARPIP: Ancestral sequence Reconstruction with insertions and deletions under the Poisson Indel Process
<p>Modern phylogenetic methods allow inference of ancestral molecular sequences given an alignment and phylogeny relating present day sequences. This provides insight into the evolutionary history of molecules, helping to understand gene function and to study biological processes such as adaptation and convergent evolution across a variety of applications. Here we propose a dynamic programming algorithm for fast joint likelihood-based reconstruction of ancestral sequences under the Poisson Indel Process (PIP). Unlike previous approaches, our method, named ARPIP, enables the reconstruction with insertions and deletions based on an explicit indel model. Consequently, inferred indel events have an explicit biological interpretation. Likelihood computation is achieved in linear time with respect to the number of sequences. Our method consists of two steps, namely finding the most probable indel points and reconstructing ancestral sequences. First, we find the most likely indel points and prune the phylogeny to reflect the insertion and deletion events per site. Second, we infer the ancestral states on the pruned subtree in a manner similar to FastML. We applied ARPIP on simulated datasets and on real data from the Betacoronavirus genus. ARPIP reconstructs both the indel events and substitutions with a high degree of accuracy. Our method fares well when compared to established state-of-the-art methods such as FastML and PAML. Moreover, the method can be extended to explore both optimal and suboptimal reconstructions, include rate heterogeneity through time and more. We believe it will expand the range of novel applications of ancestral sequence reconstruction.</p>
Cas9-induced large deletions and small indels are controlled in a convergent fashion
<p>Demultiplexed reads assembled from 150PE Illumina using Pear with following settings: pear-0.9.10-bin-64 -f read1.fq -r read2.fq -n 20 -p 0.01. Unassembled PE reads (estimated at 0.1% of all reads) and reads that failed demux not included. 960 files total, corresponding to 96 wells of the library (95 control or knock-out mES clones and one well intentionally left empty) times two biological replicates times five test gRNAs (non-targetting control g33, chromosome X targetting g15, g48 and gU48 and autosome targetting g148). mES cells are derived from CAST x BL6 cross. Non-targeting control g33 fastq files are not locus-demultiplexed ie would map to g15, g48 and g148 loci. </p>
A non-coding indel polymorphism in the fruitless gene of Drosophila melanogaster exhibits antagonistically pleiotropic fitness effects
Open the record for dataset details and reuse information.
Supplementary table, figures and DNA sequences of sorghum gene models SbiRTx430.01G455400 and SbiRTx.02G006600 that feature primers, gRNAs and indels created
Open the record for dataset details and reuse information.
Calanus InDel genotypes from: No evidence for hybridization between Calanus finmarchicus and C. glacialis in a subarctic area of sympatry
Open the record for dataset details and reuse information.
ARPIP: Ancestral sequence Reconstruction with insertions and deletions under the Poisson Indel Process
Open the record for dataset details and reuse information.
Data from: The Cumulative Indel Model: fast and accurate statistical evolutionary alignment
Sequence alignment is essential for phylogenetic and molecular evolution inference, as well as in many other areas of bioinformatics and evolutionary biology. Inaccurate alignments can lead to severe biases in most downstream statistical analyses. Statistical alignment based on probabilistic models of sequence evolution addresses these issues by replacing heuristic score functions with evolutionary model-based probabilities. However, score-based aligners and fixed-alignment phylogenetic approaches are still more prevalent than methods based on evolutionary indel models, mostly due to computational convenience. Here, I present new techniques for improving the accuracy and speed of statistical evolutionary alignment. The "cumulative indel model" approximates realistic evolutionary indel dynamics using differential equations. "Adaptive banding" reduces the computational demand of most alignment algorithms without requiring prior knowledge of divergence levels or pseudo-optimal alignments. Using simulations, I show that these methods lead to fast and accurate pairwise alignment inference. Also, I show that it is possible, with these methods, to align and infer evolutionary parameters from a single long synteny block (approximately 530kbp) between the human and chimp genomes. The cumulative indel model and adaptive banding can therefore improve the performance of alignment and phylogenetic methods.
Genotypes of 6 InDel markers for species identification from the Calanus culture at the EMBRC-ERIC laboratory for low-level trophic interactions, NTNU SeaLab
<p><span><span><span><span><span><span><span><span><span><span><span>Late developmental stages of marine copepods in the genus <i>Calanus</i> can spend extended periods in a dormant stage (diapause). During the growth season, copepods must accumulate sufficient lipid stores to survive diapause. Predation risk is often overlooked as a potential diapause-inducing cue. We tested experimentally if predation risk in combination with high or low food availability leads to differences in lipid metabolism, and potentially diapause initiation. Expression of lipid metabolism genes showed that food availability influences the copepods' ability to cope with predator stress. Predation caused upregulation of lipid catabolism with high food, and downregulation with low food. Stage development and molecular markers demonstrated that copepods did not enter diapause, instead, development occurred faster in copepods with predator stress. This study demonstrates that lipid metabolism may be a sensitive endpoint for changes in the environment. Our findings can contribute towards understanding the mechanisms behind diapause timing. </span></span></span></span></span></span></span></span></span></span></span></p>
SNPs and InDels identified in 390 peanut accessions
<p>SNPs and InDels identified in 390 peanut accessions</p>
FIGURE 1 in Indels ascertain the phylogenetic position of Coleodactylus elizae Gonçalves, Torquato, Skuk & Sena, 2012 (Gekkota: Sphaerodactylidae)
FIGURE 1. Bayesian tree topology (consensus by majority rule 50%) of sphaerodactyl geckos obtained from 1828 base pairs of concatenated nuclear genes (RAG1 and PTPN12). Closed circles indicate nodes with posterior probabilities ≥ 0.95. Indels from protein-coding regions are indicated along with the gene name, PTPN12 in Chatogekko clade and RAG1 in Coleodactylus clade. RAG1 possessed multiple unique indels and each is numbered sequentially in 5′–3′ direction following Gamble et al. 2011b. Taxon names are shown on the right followed by GenBank accession number of PTPN12 and RAG1 genes respectively.
FIGURE 1 in Printed, or just indelible? On the earliest legitimate names, authorship and typification of the taxa described from Italy by Huter, Porta and/or Rigo
FIGURE 1. Lectotype of the name Tanacetum tridactylites A.Kern. & Huter ex Porta & Rigo, conserved at NAP.
DNA matrix combined (nuclear and indels coded) datasets for Hyptidinae (Lamiaceae)
<p class="CxSpFirst">Hyptidinae, ca. 400 species, is an important component of Neotropical vegetation formations. Members of the subtribe possess flowers arranged in variously modified bracteolate cymes and nutlets with an expanded areole and all share a unique explosive mechanism of pollen release, except for <i>Asterohyptis</i>. In a recent phylogenetic study, the group had its generic delimitations rearranged with the recognition of 19 genera in the subtribe. Although the previous phylogenetic analysis covered almost all the higher taxa in the subtribe, it lacked a broader sampling at the species level. Here we present a new expanded phylogenetic analysis for the subtribe comprising 153 accessions of Hyptidinae sequenced for the nuclear nrITS, nrETS, and waxy regions and the plastid markers<i> trnL-F, trnS-G, trnD-T, </i>and<i> matK</i>. Our results widely support the previous phylogenetic results with some changes in the support and relationship between genera. It also uncovers the need for a new combination of <i>Eriope machrisae </i>in <i>Hypenia</i> and the phylogenetic position of <i>Hyptis</i> sect. <i>Rhytidea</i>, which was demonstrated to be part of <i>Mesosphaerum</i>. The generic delimitation in Hyptidinae is discussed, and we recommend that further studies with more markers are needed to confirm the monophyly of <i>Hyptidendron</i> and <i>Mesosphaerum</i>, as well as to support taxonomic changes on the infrageneric delimitation within <i>Hyptis </i>s. s.</p>
Genotypes of 6 InDel markers for species identification from the Calanus culture at the EMBRC-ERIC laboratory for low-level trophic interactions, NTNU SeaLab
Open the record for dataset details and reuse information.
DNA matrix combined (nuclear and indels coded) datasets for Hyptidinae (Lamiaceae)
Open the record for dataset details and reuse information.
Data from: The Cumulative Indel Model: fast and accurate statistical evolutionary alignment
Open the record for dataset details and reuse information.
Simulated read data analysed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph"
<p>Simulated read data analyzed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph".</p> <p><strong>1) Human sequence data</strong></p> <p><strong>HO_chr11_50bp_sliding_window*fq.gz:</strong><br> All possible 50 bp reads overlapping chromosome 11 SNPs in the Human Origins dataset. Files with the word "alternate" in their filename carry the alternate allele, otherwise, they carry the reference allele. Deamination has been added into these simulated reads using gargammel (Renaud 2016) based on empirically estimated post-mortem damage in a dataset of 102 ancient genomes (Allentoft et al., 2015).</p> <p><strong>2) microbial data</strong></p> <p><strong>simulation_*_s.fq.gz:</strong><br> Simulated microbial read data from a set of microbial reference genomes identified in the ancient Clovis genome (Rasmussen 2014), using gargammel.</p>
Data from: Not all sequence tags are created equal: designing and validating sequence identification tags robust to indels
Ligating adapters with unique synthetic oligonucleotide sequences (sequence tags) onto individual DNA samples before massively parallel sequencing is a popular and efficient way to obtain sequence data from many individual samples. Tag sequences should be numerous and sufficiently different to ensure sequencing, replication, and oligonucleotide synthesis errors do not cause tags to be unrecoverable or confused. However, many design approaches only protect against substitution errors during sequencing and extant tag sets contain too few tag sequences. We developed an open-source software package to validate sequence tags for conformance to two distance metrics and design sequence tags robust to indel and substitution errors. We use this software package to evaluate several commercial and non-commercial sequence tag sets, design several large sets (maxcount=7,198) of edit metric sequence tags having different lengths and degrees of error correction, and integrate a subset of these edit metric tags to polymerase chain reaction (PCR) primers and sequencing adapters. We validate a subset of these edit metric tagged PCR primers and sequencing adapters by sequencing on several platforms and subsequent comparison to commercially available alternatives. We find that several commonly used sets of sequence tags or design methodologies used to produce sequence tags do not meet the minimum expectations of their underlying distance metric, and we find that PCR primers and sequencing adapters incorporating edit metric sequence tags designed by our software package perform as well as their commercial counterparts. We suggest that researchers evaluate sequence tags prior to use or evaluate tags that they have been using. The sequence tag sets we design improve on extant sets because they are large, valid across the set, and robust to the suite of substitution, insertion, and deletion errors affecting massively parallel sequencing workflows on all currently used platforms.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.