Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
175
datasets available to search
ShareScore release 0.9.0
Dataset results
175 results for “sequence alignments”
Multiple sequence alignments and phylogenetic trees from: Co-option of the limb patterning program in cephalopod eye development
Open the record for dataset details and reuse information.
Prosopis laevigata microsatellite and sequence alignment data
Open the record for dataset details and reuse information.
Aligned DNA sequence matrix for phylogenetic analyses in the article "Description and phylogenetic relationships of a new trans-Andean species of Elachistocleis Parker 1927 (Amphibia, Anura, Microhylidae)"
<p>Aligned DNA sequence matrix for phylogenetic analyses in the article "Description and phylogenetic relationships of a new trans-Andean species of <em>Elachistocleis</em> Parker 1927 (Amphibia, Anura, Microhylidae)"</p> <p>Gene partitions are arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>16S = 1-1165;<br> BDNFcodonPos1 = 1166 - 1874\3;<br> BDNFcodonPos2 = 1167 - 1875\3;<br> BDNFcodonPos3 = 1168 - 1876\3;<br> cmyccodonPos1 = 1878 - 2319\3;<br> cmyccodonPos2 = 1879 - 2320\3;<br> cmyccodonPos3 = 1877 - 2318\3;<br> CO1codonPos1 = 2321 - 2981\3;<br> CO1codonPos2 = 2322 - 2979\3;<br> CO1codonPos3 = 2323 - 2980\3;<br> histcodonPos1 = 2983 - 3307\3;<br> histcodonPos2 = 2984 - 3308\3;<br> histcodonPos3 = 2982 - 3309\3;<br> siacodonPos1 = 3311 - 3704\3;<br> siacodonPos2 = 3312 - 3705\3;<br> siacodonPos3 = 3310 - 3706\3;<br> tyrcodonPos1 = 3708 - 4263\3;<br> tyrcodonPos2 = 3709 - 4264\3;<br> tyrcodonPos3 = 3707 - 4262\3;<br> 28S = 4265-5084;<br> 12S = 5085-6171;</p> <p> </p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Three new species of frogs of the genus Pristimantis (Anura: Strabomantidae) with a redefinition of the P. lacrimosus species group"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "Three new species of frogs of the genus Pristimantis (Anura: Strabomantidae) with a redefinition of the P. lacrimosus species group". The matrix is in NEXUS format.</p> <p>Gene partitions are arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>RAG1 codon position 1 = 1-625\3;<br> RAG1 codon position 2 = 2-626\3;<br> RAG1 codon position 3 = 3-627\3;<br> 12S rRNA = 628-1440;<br> 16S rRNA = 1441-3071;<br> ND1 codon position 1 = 3072-3972\3;<br> ND1 codon position 2 = 3073-3973\3;<br> ND1 codon position 3 = 3074-3974\3;</p>
Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods
<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+Γ. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+Γ: six rates and one Γ shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p> </p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p> </p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1] </code> integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5] </code> frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10] </code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] </code> Γ shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without Γ) specified to <em>INDELible</em>,</li> <li><code>[12] </code> length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] </code> length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] </code> no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] </code> no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] </code> observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] </code> aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] </code> aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment
Reduced-representation genome sequencing such as RADseq aids the analysis of genomes by reducing the quantity of data, thereby lowering both sequencing costs and computational burdens. RADseq was initially designed for studying genetic variation across genomes at the population level, but has also proved to be suitable for interspecific phylogeny reconstruction. RADseq data pose challenges for standard phylogenomic methods, however, due to incomplete coverage of the genome and large amounts of missing data. Alignment-free methods are both efficient and accurate for phylogenetic reconstructions with whole genomes and are especially practical for non-model organisms; nonetheless, alignment-free methods have not been applied with reduced genome sequencing data. Here, we test a full-genome assembly and alignment-free method, AAF, in application to RADseq data and propose two procedures for reads selection to remove reads from restriction sites that were not found in taxa being compared. We validate these methods using both simulations and real datasets. Reads selection improved the accuracy of phylogenetic construction in every simulated scenario and the two real datasets, making AAF as good or better than a comparable alignment-based method, even though AAF had much lower computational burdens. We also investigated the sources of missing data in RADseq and their effects on phylogeny reconstruction using AAF. The AAF pipeline modified for RADseq or other reduced-representation sequencing data, phyloRAD, is available on github (https://github.com/fanhuan/phyloRAD).
Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning
Reconstructing the phylogenetic relationships between species is one of the most formidable tasks in evolutionary biology. Multiple methods exist to reconstruct phylogenetic trees, each with their own strengths and weaknesses. Both simulation and empirical studies have identified several "zones" of parameter space where accuracy of some methods can plummet, even for four-taxon trees. Further, some methods can have undesirable statistical properties such as statistical inconsistency and/or the tendency to be positively misleading (i.e. assert strong support for the incorrect tree topology). Recently, deep learning techniques have made inroads on a number of both new and longstanding problems in biological research. Here we designed a deep convolutional neural network (CNN) to infer quartet topologies from multiple sequence alignments. This CNN can readily be trained to make inferences using both gapped and ungapped data. We show that our approach is highly accurate on simulated data, often outperforming traditional methods, and is remarkably robust to bias-inducing regions of parameter space such as the Felsenstein zone and the Farris zone. We also demonstrate that the confidence scores produced by our CNN can more accurately assess support for the chosen topology than bootstrap and posterior probability scores from traditional methods. While numerous practical challenges remain, these findings suggest that deep learning approaches such as ours have the potential to produce more accurate phylogenetic inferences.
Data from: In silico site-directed mutagenesis informs species-specific predictions of chemical susceptibility derived from the Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool
Chemical hazard assessment requires extrapolation of information from model organisms to all species of concern. The Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool was developed as a rapid, cost effective method to aid cross-species extrapolation of susceptibility to chemicals acting on specific protein targets through evaluation of protein structural similarities and differences. The greatest resolution for extrapolation of chemical susceptibility across species involves comparisons of individual amino acid residues at key positions involved in protein-chemical interactions. However, a lack of understanding of whether specific amino acid substitutions among species at key positions in proteins affect interaction with chemicals made manual interpretation of alignments time consuming and potentially inconsistent. Therefore, this study used in silico site-directed mutagenesis coupled with docking simulations of computational models for acetylcholinesterase (AChE) and ecdysone receptor (EcR) to investigate how specific amino acid substitutions impact protein-chemical interaction. This study found that computationally derived substitutions in identities of key amino acids caused no change in protein-chemical interaction if residues share the same side chain functional properties and have comparable molecular dimensions, while differences in these characteristics can change protein-chemical interaction. These findings were considered in the development of capabilities for automatically generated species-specific predictions of chemical susceptibility in SeqAPASS. These predictions for AChE and EcR were shown to agree with SeqAPASS predictions comparing the primary sequence and functional domain sequence of proteins for more than 90 % of the investigated species, but also identified dramatic species-specific differences in chemical susceptibility that align with results from standard toxicity tests. These results provide a compelling line-of-evidence for use of SeqAPASS in deriving screening level, species-specific, susceptibility predictions across broad taxonomic groups for application to human and ecological hazard assessment.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Multiple sequence aligners typically work by progressively aligning the most closely related sequences or group of sequences according to guide trees. In PNAS, Boyce et al. report that alignments reconstructed using simple chained trees (i.e., comb-like topologies) with random leaf assignment performed better in protein structure-based benchmarks than those reconstructed using phylogenies estimated from the data as guide trees. The authors state that this result could turn decades of research in the field on its head. In light of this statement, it is important to check immediately whether their result holds under evolutionary criteria: recovery of homologous sequence residues and inference of phylogenetic trees from the alignments. We have done this and the results are entirely opposed to Boyce et al.'s findings.
Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference
Phylogenetic inference is generally performed on the basis of multiple sequence alignments (MSA). Because errors in an alignment can lead to errors in tree estimation, there is a strong interest in identifying and removing unreliable parts of the alignment. In recent years several automated filtering approaches have been proposed, but despite their popularity, a systematic and comprehensive comparison of different alignment filtering methods on real data has been lacking. Here, we extend and apply recently introduced phylogenetic tests of alignment accuracy on a large number of gene families and contrast the performance of unfiltered versus filtered alignments in the context of single-gene phylogeny reconstruction. Based on multiple genome-wide empirical and simulated data sets, we show that the trees obtained from filtered MSAs are on average worse than those obtained from unfiltered MSAs. Furthermore, alignment filtering often leads to an increase in the proportion of well-supported branches that are actually wrong. We confirm that our findings hold for a wide range of parameters and methods. Although our results suggest that light filtering (up to 20% of alignment positions) has little impact on tree accuracy and may save some computation time, contrary to widespread practice, we do not generally recommend the use of current alignment filtering methods for phylogenetic inference. By providing a way to rigorously and systematically measure the impact of filtering on alignments, the methodology set forth here will guide the development of better filtering algorithms.
Data from: Diversity measures in environmental sequences are highly dependent on alignment quality—data from ITS and new LSU primers targeting basidiomycetes
The ribosomal DNA comprised of the ITS1-5.8S-ITS2 regions is widely used as a fungal marker in molecular ecology and systematics but cannot be aligned with confidence across genetically distant taxa. In order to study the diversity of Agaricomycotina in forest soils, we designed primers targeting the more alignable 28S (LSU) gene, which should be more useful for phylogenetic analyses of the detected taxa. This paper compares the performance of the established ITS1F/4B primer pair, which targets basidiomycetes, to that of two new pairs. Key factors in the comparison were the diversity covered, off-target amplification, rarefaction at different Operational Taxonomic Unit (OTU) cutoff levels, sensitivity of the method used to process the alignment to missing data and insecure positional homology, and the congruence of monophyletic clades with OTU assignments and BLAST-derived OTU names. The ITS primer pair yielded no off-target amplification but also exhibited the least fidelity to the expected phylogenetic groups. The LSU primers give complementary pictures of diversity, but were more sensitive to modifications of the alignment such as the removal of difficult-to align stretches. The LSU primers also yielded greater numbers of singletons but also had a greater tendency to produce OTUs containing sequences from a wider variety of species as judged by BLAST similarity. We introduced some new parameters to describe alignment heterogeneity based on Shannon entropy and the extent and contents of the OTUs in a phylogenetic tree space. Our results suggest that ITS should not be used when calculating phylogenetic trees from genetically distant sequences obtained from environmental DNA extractions and that it is inadvisable to define OTUs on the basis of very heterogeneous alignments.
Data from: GHOST: Recovering Historical Signal from Heterotachously-evolved Sequence Alignments
<p><span>Molecular sequence data that have evolved under the influence of heterotachous evolutionary processes are known to mislead phylogenetic inference. We introduce the General Heterogeneous evolution On a Single Topology (GHOST) model of sequence evolution, implemented under a maximum-likelihood framework in the phylogenetic program IQ-TREE (</span><a class="link link-uri" href="http://www.iqtree.org/">http://www.iqtree.org</a><span>). Simulations show that using the GHOST model, IQ-TREE can accurately recover the tree topology, branch lengths, and substitution model parameters from heterotachously evolved sequences. We investigate the performance of the GHOST model on empirical data by sampling phylogenomic alignments of varying lengths from a plastome alignment. We then carry out inference under the GHOST model on a phylogenomic data set composed of 248 genes from 16 taxa, where we find the GHOST model concurs with the currently accepted view, placing turtles as a sister lineage of archosaurs, in contrast to results obtained using traditional variable rates-across-sites models. Finally, we apply the model to a data set composed of a sodium channel gene of 11 fish taxa, finding that the GHOST model is able to elucidate a subtle component of the historical signal, linked to the previously established convergent evolution of the electric organ in two geographically distinct lineages of electric fish. We compare inference under the GHOST model to partitioning by codon position and show that, owing to the minimization of model constraints, the GHOST model offers unique biological insights when applied to empirical data.</span></p>
Alignments of Sequence Data for Phylogenetic Analysis of Damsel
<p>Initially described in 1882, <i>Chromis enchrysurus</i>, the Yellowtail Reeffish, was redescribed in 1982 to account for an observed color morph that possesses a white tail instead of a yellow one, but morphological and geographic boundaries between the two color morphs were not well understood. Taking advantage of newly collected material from submersible studies of deep reefs and photographs from rebreather dives, we sought to determine whether the white-tailed <i>Chromis</i> is actually a color morph of <i>Chromis enchrysurus</i> or a distinct species. These alignments for mitochondrial genes cytochrome b and cytochrome c oxidase subunit I were used to generate phylogenetic trees that separated <i>Chromis enchrysurus</i> and the white-tailed <i>Chromis</i> into two reciprocally monophyletic clades. Genetic, morphological, and biogeographic data all indicate that the white-tailed <i>Chromis</i> is a distinct species, herein described as <i>Chromis vanbebberae </i>sp. nov. The discovery of a new species within a conspicuous group such as damselfishes in a well-studied region of the world highlights the importance of deep-reef exploration in documenting undiscovered biodiversity.</p>
Molecular systematics of the tribe Physarieae (Brassicaceae) based on the nuclear ITS, LUMINIDEPENDENS, and chloroplast ndhF: Sequence alignments, trees, and supplemental figures
<p>Physarieae is a small tribe of herbaceous annual and woody perennial mustards that are mostly endemic to North America, with its members including a large amount of variation in floral, fruit and chromosomal variation. Building on a previous study of Physarieae based on morphology and <i>ndhF</i> plastid DNA, we reconstructed the evolutionary history of the tribe using new sequence data from two nuclear markers, and compared the new topologies against previously published cpDNA-based phylogenetic hypotheses. The novel analyses included ca. 420 new sequences of ITS and <i>LUMINIDEPENDENS</i> (<i>LD</i>) markers for 39 and 47 species, respectively, with sampling accounting for all seven genera of Physarieae, including nomenclatural type species, and 11 outgroup taxa. Maximum parsimony, maximum likelihood, and Bayesian analyses showed that these additional markers were largely consistent with the previous <i>ndh</i>F data that supported the monophyly of Physarieae and resolved two major clades within the tribe, i.e. DDNLS (<i>Dithyrea</i>, <i>Dimorphocarpa</i>, <i>Nerisyrenia</i>, <i>Lyrocarpa</i>, and <i>Synthlipsis</i>) and PP (<i>Paysonia</i> and <i>Physaria</i>). New analyses also increased internal resolution for some closely related species and lineages within both clades. The monophyly<i> </i>of <i>Dithyrea</i> and the sister relationship of <i>Paysonia</i> to <i>Physaria</i> was consistent in all trees, with the sister relationship of <i>Nerisyrenia </i>to <i>Lyrocarpa </i>supported by <i>ndhF</i> and <i>ITS</i>, and the positions of <i>Dimorphocarpa</i> and <i>Synthlipsis</i> shifted within the DDNLS Clade depending on the employed data set. Finally, using the strong, new phylogenetic framework of combined cpDNA + nDNA data, we discussed standing hypotheses of trichome evolution in the tribe suggested by <i>ndhF</i>.</p>
Dataset and code for Fast sequence alignment for centromere with RaMA
<p>Dataset and code for Fast sequence alignment for centromere with RaMA</p>
FORAlign: Accelerating gap-affine DNA pairwise sequence alignment using FOR-blocks based on FOur Russians approach with linear space complexity
Open the record for dataset details and reuse information.
Dastset and code for Fast sequence alignment for centromere with RaMA
<p>Dastset and code for Fast sequence alignment for centromere with RaMA</p>
Stockholm sequence-structure alignments for 9 Bacteroides sRNAs
<p>Sequence alignment supporting the manuscript "Comparative genomics provides structural and functional insights into Bacteroides RNA biology".</p>
A hepatitis B virus (HBV) sequence variation graph improves alignment and sample-specific consensus sequence construction
Open the record for dataset details and reuse information.
Alignment underpinning the molecular phylogeny of the Streptophyta based on SSU rDNA and rbcL sequence comparisons presented in Figure S1 of the article "Phylogenomic insights into the first multicellular streptophyte"
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.