Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
44
datasets available to search
ShareScore release 0.7.1
Dataset results
44 results for “multiple sequence alignment”
Cephalopod retinal development shows vertebrate-like mechanisms of neurogenesis: Multiple sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Phylogenetic and recombination analysis of adenovirus isolates reveals discordance between serotype and phylogeny: Multiple sequence alignments
Open the record for dataset details and reuse information.
Multiple alignment of DNA-B sequences from EACMV, EACMKV, EACMMV, EACMZV, SACMV (5 "species")
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Multiple alignment of ICMV and SLCMV DNA-B sequences
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl and AliView.</p> <p>Note added 2020-09-07: AJ575821 is listed in the file as ICMV based on its assignment in the NCBI Taxonomy database (taxa 341701 and 31600) but it is better classified as SLCMV.</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Multiple sequence alignment of DNA-A sequences from ACMBFV, ACMV, CMMGV, EACMCV, EACMKV, EACMMV, EACMV, EACMZV, SACMV, ICMV, SLCMV (11 species)
<p>All full-length DNA-A sequences available in GenBank as of July 2019 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were manually adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p>
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy has better precision and recall (with respect to the true alignments) than the other alignment methods on the simulated data sets, but has consistently lower recall on the biological benchmarks (with respect to the reference alignments) than many of the other methods. In other words, we find that BAli-Phy systematically under-aligns when operating on biological sequence data, but shows no sign of this on simulated data. There are several potential causes for this change in performance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments, and future research is needed to determine the most likely explanation. We conclude with a discussion of the potential ramifications for each of these possibilities.
TAPER: Pinpointing errors in multiple sequence alignments despite varying rates of evolution
<p>Datasets related to "TAPER: Pinpointing errors in multiple sequence alignments despite varying rates of evolution."</p>
Multiple sequence alignments and phylogenetic trees from: Co-option of the limb patterning program in cephalopod eye development
<p>Background</p> <p><span>Across the Metazoa, similar genetic programs are found in the development of analogous, independently evolved, morphological features. The functional significance of this reuse and the underlying mechanisms of co-option remain unclear. Cephalopods have evolved a highly acute visual system with a cup shaped retina and a novel refractive lens in the anterior, important for a number of sophisticated behaviors including predation, mating and camouflage. Almost nothing is known about the molecular-genetics of lens development in the cephalopod.</span></p> <p><span>Results</span></p> <p><span>Here we identify the co-option of the canonical bilaterian limb pattering program during cephalopod lens development, a functionally unrelated structure. We show radial expression of transcription factors <i>SP6-9/sp1, Dlx/dll, </i><i>Pbx/exd, Meis/hth, </i>and a <i>Prdl</i> homolog in the squid <i>Doryteuthis pealeii</i>, similar to expression required in <i>Drosophila</i> limb development.<i> </i>We assess the role of Wnt signaling in the cephalopod lens, a positive regulator in the developing <i>Drosophila </i>limb, and find the regulatory relationship reversed, with ectopic Wnt signaling leading to lens loss. </span></p> <p><span>Conclusion</span></p> <p><span>This regulatory divergence suggests that duplication of SP6-9 in cephalopods may mediate the co-option of the limb patterning program. Thus our study suggests that the limb network could perform a more universal developmental function in radial pattering and highlights how canonical genetic programs are repurposed in novel structures.</span></p>
Response_reg and ABC_tran: monster benchmark families for multiple sequence alignments
<p>The data set contains two benchmark families: Response_reg (1.8 million sequences) and ABC_tran (3.5 million sequences). The sets were constructed by combining Homstrad reference alignments (November 2022) with Pfam 35 UniProt complete families (PF00072 and PF00005).</p>
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
Open the record for dataset details and reuse information.
Multiple Sequence Alignments (MSA) for comparative phylogeography of four lizard taxa within an Oceanic Island
Open the record for dataset details and reuse information.
Multiple sequence alignments and phylogenetic trees from: Co-option of the limb patterning program in cephalopod eye development
Open the record for dataset details and reuse information.
Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning
Reconstructing the phylogenetic relationships between species is one of the most formidable tasks in evolutionary biology. Multiple methods exist to reconstruct phylogenetic trees, each with their own strengths and weaknesses. Both simulation and empirical studies have identified several "zones" of parameter space where accuracy of some methods can plummet, even for four-taxon trees. Further, some methods can have undesirable statistical properties such as statistical inconsistency and/or the tendency to be positively misleading (i.e. assert strong support for the incorrect tree topology). Recently, deep learning techniques have made inroads on a number of both new and longstanding problems in biological research. Here we designed a deep convolutional neural network (CNN) to infer quartet topologies from multiple sequence alignments. This CNN can readily be trained to make inferences using both gapped and ungapped data. We show that our approach is highly accurate on simulated data, often outperforming traditional methods, and is remarkably robust to bias-inducing regions of parameter space such as the Felsenstein zone and the Farris zone. We also demonstrate that the confidence scores produced by our CNN can more accurately assess support for the chosen topology than bootstrap and posterior probability scores from traditional methods. While numerous practical challenges remain, these findings suggest that deep learning approaches such as ours have the potential to produce more accurate phylogenetic inferences.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Multiple sequence aligners typically work by progressively aligning the most closely related sequences or group of sequences according to guide trees. In PNAS, Boyce et al. report that alignments reconstructed using simple chained trees (i.e., comb-like topologies) with random leaf assignment performed better in protein structure-based benchmarks than those reconstructed using phylogenies estimated from the data as guide trees. The authors state that this result could turn decades of research in the field on its head. In light of this statement, it is important to check immediately whether their result holds under evolutionary criteria: recovery of homologous sequence residues and inference of phylogenetic trees from the alignments. We have done this and the results are entirely opposed to Boyce et al.'s findings.
Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference
Phylogenetic inference is generally performed on the basis of multiple sequence alignments (MSA). Because errors in an alignment can lead to errors in tree estimation, there is a strong interest in identifying and removing unreliable parts of the alignment. In recent years several automated filtering approaches have been proposed, but despite their popularity, a systematic and comprehensive comparison of different alignment filtering methods on real data has been lacking. Here, we extend and apply recently introduced phylogenetic tests of alignment accuracy on a large number of gene families and contrast the performance of unfiltered versus filtered alignments in the context of single-gene phylogeny reconstruction. Based on multiple genome-wide empirical and simulated data sets, we show that the trees obtained from filtered MSAs are on average worse than those obtained from unfiltered MSAs. Furthermore, alignment filtering often leads to an increase in the proportion of well-supported branches that are actually wrong. We confirm that our findings hold for a wide range of parameters and methods. Although our results suggest that light filtering (up to 20% of alignment positions) has little impact on tree accuracy and may save some computation time, contrary to widespread practice, we do not generally recommend the use of current alignment filtering methods for phylogenetic inference. By providing a way to rigorously and systematically measure the impact of filtering on alignments, the methodology set forth here will guide the development of better filtering algorithms.
Data from: Multiple sequence alignment averaging improves phylogeny reconstruction
Open the record for dataset details and reuse information.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Open the record for dataset details and reuse information.
Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning
Open the record for dataset details and reuse information.
Data from: SATé-II: very fast and accurate simultaneous estimation of multiple sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.