Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

175

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

175 results for “sequence alignments”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: In silico site-directed mutagenesis informs species-specific predictions of chemical susceptibility derived from the Sequence Alignment to Predict Across Species Susceptibility (SeqAPASS) tool

Open the record for dataset details and reuse information.

publicJul 2018View details →
dryad28/100

Data from: Diversity measures in environmental sequences are highly dependent on alignment quality—data from ITS and new LSU primers targeting basidiomycetes

Open the record for dataset details and reuse information.

publicMar 2012View details →
dryad28/100

Data from: Multiple sequence alignment averaging improves phylogeny reconstruction

Open the record for dataset details and reuse information.

publicMay 2018View details →
dryad28/100

Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad28/100

Data from: Generalized bootstrap supports for phylogenetic analyses of protein sequences incorporating alignment uncertainty

Open the record for dataset details and reuse information.

publicDec 2017View details →
dryad28/100

Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad28/100

Alignments of Sequence Data for Phylogenetic Analysis of Damsel

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad28/100

Data from: SATé-II: very fast and accurate simultaneous estimation of multiple sequence alignments and phylogenetic trees

Open the record for dataset details and reuse information.

publicJun 2011View details →
dryad28/100

Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference

Open the record for dataset details and reuse information.

publicMay 2015View details →
dryad28/100

Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment

Open the record for dataset details and reuse information.

publicMay 2018View details →
dryad28/100

Data from: GHOST: Recovering Historical Signal from Heterotachously-evolved Sequence Alignments

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad28/100

Sequence alignment of Stizophyllum for the markers: Ndhf, rpl32-trnl and pepc

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad28/100

Molecular systematics of the tribe Physarieae (Brassicaceae) based on the nuclear ITS, Luminidependens, and chloroplast ndhF: Sequence alignments, trees, and supplemental figures

Open the record for dataset details and reuse information.

publicDec 2021View details →
zenodo24/100

Sequence alignments and tree file for wsp gene.

<p>Alignment of <em>wsp</em> dataset was constructed on the GUIDANCE2 Server based on codons using the MAFFT algorithm, and ambiguous alignments with the confidence score below 0.7 were excluded. The maximum likelihood&nbsp;tree was built using <em>RAxML</em>&nbsp;program.</p>

opencc-by-4.0Oct 2019View details →
zenodo24/100

Simulated nucleotide sequences for testing alignment-free genome distance estimates

<p>This repository contains (12&times;500=)6,000 pairs of nucleotide sequences that have been simulated for testing alignment-free genome distance estimates, as described in <a href="https://riojournal.com/article/36178/">Criscuolo (2019)</a>. Given an evolutionary distance <em>d</em> varying from 0.05 to 0.60 (step = 0.05), the program <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> was used to simulate the evolution of 500 nucleotide sequence pairs with <em>d</em> substitution events per character (GTR+&Gamma; evolutionary model).</p> <p>For each of the 12 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 0.60, an XZ-compressed file containing 500 lines is available. Each line contains 18 fields separated by blank spaces:<br> &nbsp; [1] &nbsp; &nbsp; seed value used during simulation,<br> &nbsp; [2] &nbsp; &nbsp; true evolutionary distance <em>d</em> between the two simulated sequences,<br> &nbsp; [3] &nbsp; &nbsp; total number of simulated characters,<br> &nbsp; [4] &nbsp; &nbsp; number of non-indel characters with nucleotide mismatch,<br> &nbsp; [5] &nbsp; &nbsp; number of non-indel characters,<br> &nbsp; [6-9] &nbsp; A, C, G, T frequencies used during simulation,<br> &nbsp; [10-15] &nbsp; GTR parameters used during simulation,<br> &nbsp; [16] &nbsp; &nbsp; &Gamma; distribution parameter used during simulation,<br> &nbsp; [17-18] &nbsp; two simulated sequences with indel events as gaps.</p> <p>Of note, each pair of aligned sequences without gaps can be regenerated using <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> v1.3.4 with parameters from fields [1,3,6-16] and the following two-leaf model tree:</p> <pre>(t1:d,t2:0.000);</pre> <p>where <em>d</em> is given in field [2].</p> <p>___</p> <p>Criscuolo A (2019) <em>A fast alignment-free bioinformatics procedure to infer accurate distance-based phylogenetic trees from genome assemblies</em>. Research Ideas and Outcomes, 5:e36178. doi:<a href="https://doi.org/10.3897/rio.5.e36178">10.3897/rio.5.e36178</a></p>

opencc-by-4.0Sep 2020View details →
dryad24/100

Sequence alignments for ITS, rpl16, and trnL-trnF

<p>The two data files consist of sequence alignments in FASTA format. The file 'Distichium_alignment_ITS.txt' is the alignment used for the analysis resulting in the network in Fig. 1A in the paper and the file 'Distichium_alignment_3 markers.txt' is the alignment used for the analysis resulting in the network in Fig. 1B. The GenBank numbers corresponding with the sequence numbers can be found in Appendix 1 in the paper.</p>

opencc-zeroNov 2021View details →
zenodo24/100

Multiple Sequence Alignments of primate proteomes

<p>This image dataset contains Multiple Sequence Alignments of&nbsp;protein sequences extracted from the Uniprot reference proteomes and RefSeq databases.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
dryad24/100

Tree and sequence alignment used for diet reconstruction of Laurasiatheria

<p><b>Background:</b> Laurasiatheria contains taxa with diverse diets, while the molecular basis and evolutionary history underlying their dietary diversification are less clear.</p> <p><b>Results:</b> In this study, we used the recently developed molecular phyloecological approach to examine the adaptive evolution of digestive system-related genes across both carnivorous and herbivorous mammals within Laurasiatheria. Our results show an intensified selection of fat and/or protein utilization across all examined carnivorous lineages, which is consistent with their high-protein and high-fat diets. Intriguingly, for herbivorous lineages (ungulates), which have a high-carbohydrate diet, they show a similar selection pattern as that of carnivorous lineages. Our results suggest that for the ungulates, which have a specialized digestive system, the selection intensity of their digestive system-related genes does not necessarily reflect loads of the nutrient components in their diets but appears to be positively related to the loads of the nutrient components that are capable of being directly utilized by the herbivores themselves. Based on these findings, we reconstructed the dietary evolution within Laurasiatheria, and our results reveal the dominant carnivory during the early diversification of Laurasiatheria. In particular, our results suggest that the ancestral bats and the common ancestor of ruminants and cetaceans may be carnivorous as well. We also found evidence of the convergent evolution of one fat utilization-related gene, <i>APOB</i>, across carnivorous taxa.</p> <p><b>Conclusions: </b>Our molecular phyloecological results suggest that digestive system-related genes can be used to determine the molecular basis of diet differentiations and to reconstruct ancestral diets.  </p>

opencc-zeroSep 2021View details →
zenodo24/100

alignment of potassium channels sequences

<p>This aligment (MAFFT) was used to perform a phylogenetic analysis (PhyML) to determine the position of the potassium channel sequences of Oopsacas minuta.</p>

opencc-by-4.0Jan 2023View details →
dryad24/100

Aligned and reduced 5'-COI sequence alignment from Steinke et al. 2009

Open the record for dataset details and reuse information.

publicMar 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record