Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
175
datasets available to search
ShareScore release 0.9.0
Dataset results
175 results for “sequence alignments”
Data from: In silico site-directed mutagenesis informs species-specific predictions of chemical susceptibility derived from the Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool
Open the record for dataset details and reuse information.
Data from: Diversity measures in environmental sequences are highly dependent on alignment quality—data from ITS and new LSU primers targeting basidiomycetes
Open the record for dataset details and reuse information.
Data from: Multiple sequence alignment averaging improves phylogeny reconstruction
Open the record for dataset details and reuse information.
Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks
Open the record for dataset details and reuse information.
Data from: Generalized bootstrap supports for phylogenetic analyses of protein sequences incorporating alignment uncertainty
Open the record for dataset details and reuse information.
Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning
Open the record for dataset details and reuse information.
Alignments of Sequence Data for Phylogenetic Analysis of Damsel
Open the record for dataset details and reuse information.
Data from: SATé-II: very fast and accurate simultaneous estimation of multiple sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference
Open the record for dataset details and reuse information.
Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment
Open the record for dataset details and reuse information.
Data from: GHOST: Recovering Historical Signal from Heterotachously-evolved Sequence Alignments
Open the record for dataset details and reuse information.
Sequence alignment of Stizophyllum for the markers: Ndhf, rpl32-trnl and pepc
Open the record for dataset details and reuse information.
Molecular systematics of the tribe Physarieae (Brassicaceae) based on the nuclear ITS, Luminidependens, and chloroplast ndhF: Sequence alignments, trees, and supplemental figures
Open the record for dataset details and reuse information.
Sequence alignments and tree file for wsp gene.
<p>Alignment of <em>wsp</em> dataset was constructed on the GUIDANCE2 Server based on codons using the MAFFT algorithm, and ambiguous alignments with the confidence score below 0.7 were excluded. The maximum likelihood tree was built using <em>RAxML</em> program.</p>
Simulated nucleotide sequences for testing alignment-free genome distance estimates
<p>This repository contains (12×500=)6,000 pairs of nucleotide sequences that have been simulated for testing alignment-free genome distance estimates, as described in <a href="https://riojournal.com/article/36178/">Criscuolo (2019)</a>. Given an evolutionary distance <em>d</em> varying from 0.05 to 0.60 (step = 0.05), the program <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> was used to simulate the evolution of 500 nucleotide sequence pairs with <em>d</em> substitution events per character (GTR+Γ evolutionary model).</p> <p>For each of the 12 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 0.60, an XZ-compressed file containing 500 lines is available. Each line contains 18 fields separated by blank spaces:<br> [1] seed value used during simulation,<br> [2] true evolutionary distance <em>d</em> between the two simulated sequences,<br> [3] total number of simulated characters,<br> [4] number of non-indel characters with nucleotide mismatch,<br> [5] number of non-indel characters,<br> [6-9] A, C, G, T frequencies used during simulation,<br> [10-15] GTR parameters used during simulation,<br> [16] Γ distribution parameter used during simulation,<br> [17-18] two simulated sequences with indel events as gaps.</p> <p>Of note, each pair of aligned sequences without gaps can be regenerated using <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> v1.3.4 with parameters from fields [1,3,6-16] and the following two-leaf model tree:</p> <pre>(t1:d,t2:0.000);</pre> <p>where <em>d</em> is given in field [2].</p> <p>___</p> <p>Criscuolo A (2019) <em>A fast alignment-free bioinformatics procedure to infer accurate distance-based phylogenetic trees from genome assemblies</em>. Research Ideas and Outcomes, 5:e36178. doi:<a href="https://doi.org/10.3897/rio.5.e36178">10.3897/rio.5.e36178</a></p>
Sequence alignments for ITS, rpl16, and trnL-trnF
<p>The two data files consist of sequence alignments in FASTA format. The file 'Distichium_alignment_ITS.txt' is the alignment used for the analysis resulting in the network in Fig. 1A in the paper and the file 'Distichium_alignment_3 markers.txt' is the alignment used for the analysis resulting in the network in Fig. 1B. The GenBank numbers corresponding with the sequence numbers can be found in Appendix 1 in the paper.</p>
Multiple Sequence Alignments of primate proteomes
<p>This image dataset contains Multiple Sequence Alignments of protein sequences extracted from the Uniprot reference proteomes and RefSeq databases. </p>
Tree and sequence alignment used for diet reconstruction of Laurasiatheria
<p><b>Background:</b> Laurasiatheria contains taxa with diverse diets, while the molecular basis and evolutionary history underlying their dietary diversification are less clear.</p> <p><b>Results:</b> In this study, we used the recently developed molecular phyloecological approach to examine the adaptive evolution of digestive system-related genes across both carnivorous and herbivorous mammals within Laurasiatheria. Our results show an intensified selection of fat and/or protein utilization across all examined carnivorous lineages, which is consistent with their high-protein and high-fat diets. Intriguingly, for herbivorous lineages (ungulates), which have a high-carbohydrate diet, they show a similar selection pattern as that of carnivorous lineages. Our results suggest that for the ungulates, which have a specialized digestive system, the selection intensity of their digestive system-related genes does not necessarily reflect loads of the nutrient components in their diets but appears to be positively related to the loads of the nutrient components that are capable of being directly utilized by the herbivores themselves. Based on these findings, we reconstructed the dietary evolution within Laurasiatheria, and our results reveal the dominant carnivory during the early diversification of Laurasiatheria. In particular, our results suggest that the ancestral bats and the common ancestor of ruminants and cetaceans may be carnivorous as well. We also found evidence of the convergent evolution of one fat utilization-related gene, <i>APOB</i>, across carnivorous taxa.</p> <p><b>Conclusions: </b>Our molecular phyloecological results suggest that digestive system-related genes can be used to determine the molecular basis of diet differentiations and to reconstruct ancestral diets. </p>
alignment of potassium channels sequences
<p>This aligment (MAFFT) was used to perform a phylogenetic analysis (PhyML) to determine the position of the potassium channel sequences of Oopsacas minuta.</p>
Aligned and reduced 5'-COI sequence alignment from Steinke et al. 2009
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.