Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

43

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

43 results for “Indel”

Learn how ShareScore rates datasets ↗
zenodo44/100

SNP and indel discovery and genotyping in next-generation sequencing data

<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels &gt;50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

Semi-simulated dataset of indel calling tool evaluation

<p>The semi-simulated datasets used in Ning Wang, et al.&nbsp;</p> <p>100bp_5X_500_Venter_read1.fq.gz &amp;&nbsp;100bp_5X_500_Venter_read2.fq.gz: 5X coverage, 100bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>100bp_30X_500_Venter_read1.fq.gz &amp; 100bp_30X_500_Venter_read2.fq.gz: 30X coverage, 100bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>250bp_30X_500_Venter_read1.fq.gz &amp; 250bp_30X_500_Venter_read2.fq.gz: 30X coverage, 250bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>100bp_60X_500_Venter_read1.fq.gz &amp; 100bp_60X_500_Venter_read2.fq.gz: 60X coverage, 100bp read length semi-simulated paired-end sequencing FASTQ files.</p> <p>Haplotype_1_no_gap.fa &amp;&nbsp;Haplotype_2_no_gap.fa: the semi-simulated dipoid human genome hg19 chromosome 1 and chromosome 2, including HuRef indels used in Ning Wang, et al. In order to use ART (fastq simulator) to generate simulated fastq fiels, the gaps (Ns) of genome were removed.</p> <p>chr1_chr2_variants_truthset.txt: types, position, size and genotype of HuRef indels used in Ning Wang, et al.&nbsp;</p>

opencc-by-4.0Mar 2021View details →
zenodo36/100

Filtered and annotated SNV and indel variants in the PC3 and LNCaP human prostate cancer cell lines

<p>150bp paired-end reads (insert size 350bp) were obtained using the Illumina HiSeqX sequencer. Samtools v1.3.1 mpileup and bcftools were used to interrogate indexed BAM files, from whole-genome reads aligned to human reference genome GRCh38 build 82, and generate a VCF (Variant Call Format) file of single nucleotide variants (SNVs) and short indel variants. Variants private, or unique to a particular cell line, or shared by both were next identified. Variants (likely to be common germline variants) present in HapMap, 1000 genomes phase 3 (2,504 human genomes), and the National Heart Lung and Blood Institute’s Exome Sequencing Project (ESP) (bundled variant data file available at https://goo.gl/mEogvD) were excluded. Variant files (VCF) were filtered using SnpSift  with the following parameters: 'QUAL \textgreater= 200 \&amp;\&amp; DP \textgreater= 30', where QUAL denotes minimum variance confidence and DP total depth threshold. Filtered variants were annotated using SnpEff v4.3g. Please see https://github.com/sciseim/PCaWGS for associated scripts.</p> <p> </p>

opencc-by-4.0Jan 2017View details →
dryad36/100

ARPIP: Ancestral sequence Reconstruction with insertions and deletions under the Poisson Indel Process

<p>Modern phylogenetic methods allow inference of ancestral molecular sequences given an alignment and phylogeny relating present day sequences. This provides insight into the evolutionary history of molecules, helping to understand gene function and to study biological processes such as adaptation and convergent evolution across a variety of applications. Here we propose a dynamic programming algorithm for fast joint likelihood-based reconstruction of ancestral sequences under the Poisson Indel Process (PIP). Unlike previous approaches, our method, named ARPIP, enables the reconstruction with insertions and deletions based on an explicit indel model. Consequently, inferred indel events have an explicit biological interpretation. Likelihood computation is achieved in linear time with respect to the number of sequences. Our method consists of two steps, namely finding the most probable indel points and reconstructing ancestral sequences. First, we find the most likely indel points and prune the phylogeny to reflect the insertion and deletion events per site. Second, we infer the ancestral states on the pruned subtree in a manner similar to FastML. We applied ARPIP on simulated datasets and on real data from the Betacoronavirus genus. ARPIP reconstructs both the indel events and substitutions with a high degree of accuracy. Our method fares well when compared to established state-of-the-art methods such as FastML and PAML. Moreover, the method can be extended to explore both optimal and suboptimal reconstructions, include rate heterogeneity through time and more. We believe it will expand the range of novel applications of ancestral sequence reconstruction.</p>

opencc-zeroJul 2022View details →
zenodo36/100

Cas9-induced large deletions and small indels are controlled in a convergent fashion

<p>Demultiplexed reads assembled from 150PE Illumina using Pear with following settings: pear-0.9.10-bin-64 -f read1.fq -r read2.fq -n 20 -p 0.01. Unassembled PE reads (estimated at 0.1% of all reads) and reads that failed demux not included. 960 files total, corresponding to 96 wells of the library (95 control or knock-out mES clones and one well intentionally left empty) times two biological replicates times five test gRNAs (non-targetting control g33, chromosome X targetting g15, g48 and gU48 and autosome targetting g148). mES cells are derived from CAST x BL6 cross. Non-targeting control g33 fastq files are not locus-demultiplexed ie would map to g15, g48 and g148 loci.&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

A non-coding indel polymorphism in the fruitless gene of Drosophila melanogaster exhibits antagonistically pleiotropic fitness effects

Open the record for dataset details and reuse information.

publicApr 2021View details →
dryad36/100

Supplementary table, figures and DNA sequences of sorghum gene models SbiRTx430.01G455400 and SbiRTx.02G006600 that feature primers, gRNAs and indels created

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad36/100

Calanus InDel genotypes from: No evidence for hybridization between Calanus finmarchicus and C. glacialis in a subarctic area of sympatry

Open the record for dataset details and reuse information.

publicApr 2020View details →
dryad36/100

ARPIP: Ancestral sequence Reconstruction with insertions and deletions under the Poisson Indel Process

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad32/100

Data from: The Cumulative Indel Model: fast and accurate statistical evolutionary alignment

Sequence alignment is essential for phylogenetic and molecular evolution inference, as well as in many other areas of bioinformatics and evolutionary biology. Inaccurate alignments can lead to severe biases in most downstream statistical analyses. Statistical alignment based on probabilistic models of sequence evolution addresses these issues by replacing heuristic score functions with evolutionary model-based probabilities. However, score-based aligners and fixed-alignment phylogenetic approaches are still more prevalent than methods based on evolutionary indel models, mostly due to computational convenience. Here, I present new techniques for improving the accuracy and speed of statistical evolutionary alignment. The "cumulative indel model" approximates realistic evolutionary indel dynamics using differential equations. "Adaptive banding" reduces the computational demand of most alignment algorithms without requiring prior knowledge of divergence levels or pseudo-optimal alignments. Using simulations, I show that these methods lead to fast and accurate pairwise alignment inference. Also, I show that it is possible, with these methods, to align and infer evolutionary parameters from a single long synteny block (approximately 530kbp) between the human and chimp genomes. The cumulative indel model and adaptive banding can therefore improve the performance of alignment and phylogenetic methods.

opencc-zeroAug 2020View details →
dryad32/100

Genotypes of 6 InDel markers for species identification from the Calanus culture at the EMBRC-ERIC laboratory for low-level trophic interactions, NTNU SeaLab

<p><span><span><span><span><span><span><span><span><span><span><span>Late developmental stages of marine copepods in the genus <i>Calanus</i> can spend extended periods in a dormant stage (diapause). During the growth season, copepods must accumulate sufficient lipid stores to survive diapause. Predation risk is often overlooked as a potential diapause-inducing cue. We tested experimentally if predation risk in combination with high or low food availability leads to differences in lipid metabolism, and potentially diapause initiation. Expression of lipid metabolism genes showed that food availability influences the copepods' ability to cope with predator stress. Predation caused upregulation of lipid catabolism with high food, and downregulation with low food. Stage development and molecular markers demonstrated that copepods did not enter diapause, instead, development occurred faster in copepods with predator stress. This study demonstrates that lipid metabolism may be a sensitive endpoint for changes in the environment. Our findings can contribute towards understanding the mechanisms behind diapause timing. </span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroDec 2020View details →
zenodo32/100

SNPs and InDels identified in 390 peanut accessions

<p>SNPs and InDels identified in 390 peanut accessions</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

FIGURE 1 in Indels ascertain the phylogenetic position of Coleodactylus elizae Gonçalves, Torquato, Skuk & Sena, 2012 (Gekkota: Sphaerodactylidae)

FIGURE 1. Bayesian tree topology (consensus by majority rule 50%) of sphaerodactyl geckos obtained from 1828 base pairs of concatenated nuclear genes (RAG1 and PTPN12). Closed circles indicate nodes with posterior probabilities ≥ 0.95. Indels from protein-coding regions are indicated along with the gene name, PTPN12 in Chatogekko clade and RAG1 in Coleodactylus clade. RAG1 possessed multiple unique indels and each is numbered sequentially in 5′–3′ direction following Gamble et al. 2011b. Taxon names are shown on the right followed by GenBank accession number of PTPN12 and RAG1 genes respectively.

opennotspecifiedFeb 2016View details →
zenodo32/100

FIGURE 1 in Printed, or just indelible? On the earliest legitimate names, authorship and typification of the taxa described from Italy by Huter, Porta and/or Rigo

FIGURE 1. Lectotype of the name Tanacetum tridactylites A.Kern. &amp; Huter ex Porta &amp; Rigo, conserved at NAP.

opennotspecifiedJul 2018View details →
dryad32/100

DNA matrix combined (nuclear and indels coded) datasets for Hyptidinae (Lamiaceae)

<p class="CxSpFirst">Hyptidinae, ca. 400 species, is an important component of Neotropical vegetation formations. Members of the subtribe possess flowers arranged in variously modified bracteolate cymes and nutlets with an expanded areole and all share a unique explosive mechanism of pollen release, except for <i>Asterohyptis</i>. In a recent phylogenetic study, the group had its generic delimitations rearranged with the recognition of 19 genera in the subtribe. Although the previous phylogenetic analysis covered almost all the higher taxa in the subtribe, it lacked a broader sampling at the species level. Here we present a new expanded phylogenetic analysis for the subtribe comprising 153 accessions of Hyptidinae sequenced for the nuclear nrITS, nrETS, and waxy regions and the plastid markers<i> trnL-F, trnS-G, trnD-T, </i>and<i> matK</i>. Our results widely support the previous phylogenetic results with some changes in the support and relationship between genera. It also uncovers the need for a new combination of <i>Eriope machrisae </i>in <i>Hypenia</i> and the phylogenetic position of <i>Hyptis</i> sect. <i>Rhytidea</i>, which was demonstrated to be part of <i>Mesosphaerum</i>. The generic delimitation in Hyptidinae is discussed, and we recommend that further studies with more markers are needed to confirm the monophyly of <i>Hyptidendron</i> and <i>Mesosphaerum</i>, as well as to support taxonomic changes on the infrageneric delimitation within <i>Hyptis </i>s. s.</p>

opencc-zeroJul 2021View details →
dryad32/100

Genotypes of 6 InDel markers for species identification from the Calanus culture at the EMBRC-ERIC laboratory for low-level trophic interactions, NTNU SeaLab

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad32/100

DNA matrix combined (nuclear and indels coded) datasets for Hyptidinae (Lamiaceae)

Open the record for dataset details and reuse information.

publicJul 2021View details →
dryad32/100

Data from: The Cumulative Indel Model: fast and accurate statistical evolutionary alignment

Open the record for dataset details and reuse information.

publicAug 2020View details →
zenodo28/100

Simulated read data analysed in "Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph"

<p>Simulated read data analyzed in &quot;Removing reference bias and improving indel calling in ancient DNA data analysis by mapping to a sequence variation graph&quot;.</p> <p><strong>1) Human sequence data</strong></p> <p><strong>HO_chr11_50bp_sliding_window*fq.gz:</strong><br> All possible 50 bp reads overlapping chromosome 11 SNPs in the Human Origins dataset. Files with the word &quot;alternate&quot; in their filename carry the alternate allele, otherwise, they carry the reference allele. Deamination has been added into these simulated reads using gargammel (Renaud 2016) based on empirically estimated post-mortem damage in a dataset of 102 ancient genomes (Allentoft et al., 2015).</p> <p><strong>2) microbial data</strong></p> <p><strong>simulation_*_s.fq.gz:</strong><br> Simulated microbial read data&nbsp;from a set of microbial reference genomes identified in the ancient Clovis genome (Rasmussen 2014), using gargammel.</p>

opencc-by-4.0Jul 2020View details →
dryad28/100

Data from: Not all sequence tags are created equal: designing and validating sequence identification tags robust to indels

Ligating adapters with unique synthetic oligonucleotide sequences (sequence tags) onto individual DNA samples before massively parallel sequencing is a popular and efficient way to obtain sequence data from many individual samples. Tag sequences should be numerous and sufficiently different to ensure sequencing, replication, and oligonucleotide synthesis errors do not cause tags to be unrecoverable or confused. However, many design approaches only protect against substitution errors during sequencing and extant tag sets contain too few tag sequences. We developed an open-source software package to validate sequence tags for conformance to two distance metrics and design sequence tags robust to indel and substitution errors. We use this software package to evaluate several commercial and non-commercial sequence tag sets, design several large sets (maxcount=7,198) of edit metric sequence tags having different lengths and degrees of error correction, and integrate a subset of these edit metric tags to polymerase chain reaction (PCR) primers and sequencing adapters. We validate a subset of these edit metric tagged PCR primers and sequencing adapters by sequencing on several platforms and subsequent comparison to commercially available alternatives. We find that several commonly used sets of sequence tags or design methodologies used to produce sequence tags do not meet the minimum expectations of their underlying distance metric, and we find that PCR primers and sequencing adapters incorporating edit metric sequence tags designed by our software package perform as well as their commercial counterparts. We suggest that researchers evaluate sequence tags prior to use or evaluate tags that they have been using. The sequence tag sets we design improve on extant sets because they are large, valid across the set, and robust to the suite of substitution, insertion, and deletion errors affecting massively parallel sequencing workflows on all currently used platforms.

opencc-zeroDec 2011View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record