Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

55

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

55 results for “multiple alignment”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference

Phylogenetic inference is generally performed on the basis of multiple sequence alignments (MSA). Because errors in an alignment can lead to errors in tree estimation, there is a strong interest in identifying and removing unreliable parts of the alignment. In recent years several automated filtering approaches have been proposed, but despite their popularity, a systematic and comprehensive comparison of different alignment filtering methods on real data has been lacking. Here, we extend and apply recently introduced phylogenetic tests of alignment accuracy on a large number of gene families and contrast the performance of unfiltered versus filtered alignments in the context of single-gene phylogeny reconstruction. Based on multiple genome-wide empirical and simulated data sets, we show that the trees obtained from filtered MSAs are on average worse than those obtained from unfiltered MSAs. Furthermore, alignment filtering often leads to an increase in the proportion of well-supported branches that are actually wrong. We confirm that our findings hold for a wide range of parameters and methods. Although our results suggest that light filtering (up to 20% of alignment positions) has little impact on tree accuracy and may save some computation time, contrary to widespread practice, we do not generally recommend the use of current alignment filtering methods for phylogenetic inference. By providing a way to rigorously and systematically measure the impact of filtering on alignments, the methodology set forth here will guide the development of better filtering algorithms.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Exploring data interaction and nucleotide alignment in a multiple gene analysis of Ips (Coleoptera: Scolytinae)

The possibility of gene tree incongruence in a species-level phylogenetic analysis of the genus Ips (Coleoptera: Scolytidae) was investigated based on mitochondrial 16S rRNA and nuclear Elongation factor-1α sequences, and existing Cytochrome Oxidase I and non-molecular data sets. Separate cladistic analyses of the data partitions resulted in partially discordant most-parsimonious trees but revealed only low conflict of the phylogenetic signal. Interactions among data partitions, which differed in the level of sequence divergence (COI > 16S > EF-1α), base composition, and homoplasy, revealed that much of the branch support only emerges in the simultaneous analysis, in particular for deeper nodes in the tree which are almost entirely supported due to "hidden support" (sensu Gatesy et al., 1999). Apparent incongruence between data partitions is in part due to suboptimal alignments and bias of character transformations, but there is little evidence to invoke incongruent phylogenetic histories of genetic loci. There is also no justification for eliminating or downweighting gene partitions based on their level of homoplasy or apparent incongruence with other partitions, as the signal only emerges in the interaction of all data. In comparison to the traditional taxonomy, the pini, plastographus and perturbatus groups are polyphyletic, whereas the grandicollis group is monophyletic except for the inclusion of the (monophyletic) calligraphus group. The latidens group and some European species are distantly related and closer to other genera within Ipini. Our robust cladogram was used to revise the classification of Ips. We provide new diagnoses for Ips and four subgeneric taxa.

opencc-zeroDec 2008View details →
zenodo28/100

Multiple whole genome alignment of 63 Nymphalidae (HAL file, alignment version 1.0), plus protein-coding annotations for each of the species in the hall file.

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
dryad28/100

Data from: Multiple sequence alignment averaging improves phylogeny reconstruction

Open the record for dataset details and reuse information.

publicMay 2018View details →
dryad28/100

Data from: Simple chained guide trees give poorer multiple sequence alignments than inferred trees in simulation and phylogenetic benchmarks

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad28/100

Data from: Exploring data interaction and nucleotide alignment in a multiple gene analysis of Ips (Coleoptera: Scolytinae)

Open the record for dataset details and reuse information.

publicJun 2009View details →
dryad28/100

Data from: Accurate inference of tree topologies from multiple sequence alignments using deep learning

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad28/100

Data from: SATé-II: very fast and accurate simultaneous estimation of multiple sequence alignments and phylogenetic trees

Open the record for dataset details and reuse information.

publicJun 2011View details →
dryad28/100

Data from: Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference

Open the record for dataset details and reuse information.

publicMay 2015View details →
zenodo24/100

Multiple Sequence Alignments of primate proteomes

<p>This image dataset contains Multiple Sequence Alignments of&nbsp;protein sequences extracted from the Uniprot reference proteomes and RefSeq databases.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
dryad24/100

Multiple alignments of Ficus benghalensis var. krishnae markers

<p><em>Ficus</em> <em>benghalensis</em> var. <em>krishnae</em> is a rare tree species considered to be endemic to India. Its peculiar feature is the presence of a cup-shaped structure at the base of its leaves. The taxonomic status of <em>Ficus</em> <em>benghalensis</em> and its variant <em>Ficus</em> <em>benghalensis</em> var. <em>krishnae</em> is currently under debate. In this study, we used two chloroplast DNA barcode markers (<em>matK</em> and <em>rbcL</em>) to explore the taxonomical relationship of <em>Ficus</em> <em>benghalensis</em> var. <em>krishnae</em> with other <em>Ficus</em> species.</p>

opencc-zeroAug 2023View details →
dryad24/100

Multiple alignments of Ficus benghalensis var. krishnae markers

Open the record for dataset details and reuse information.

publicAug 2023View details →
zenodo12/100

Candidate human enhancer multiple sequence alignments

<p>ENCODE SCREEN V3 Proximal and Distal Enhancers with homologs.</p>

restrictedcc-by-4.0Dec 2023View details →
zenodo12/100

Multiple sequence alignments of Treponema pallidum complete genomes using three different references for mapping NGS reads

<p>Each file corresponds to the multiple sequence alignment of 75 complete Treponema pallidum genome sequences using the genomes of strains Nichols, SS14, and CDC-2 as references for mapping. This is supplemental data to the manuscript &quot;Evolutionary processes in the emergence and recent spread of <em>Treponema pallidum</em>, the causative agent of syphilis&quot; by Marta Pla-D&iacute;az, Leonor S&aacute;nchez-Bus&oacute;, Lorenzo Giacani, David &Scaron;majs, Philipp P. Bosshard, Homayoun C. Bagheri, Verena J. Schuenemann, Kay Nieselt, Natasha Arora and Fernando Gonz&aacute;lez-Candelas</p>

restrictedAug 2021View details →
zenodo8/100

Multiple sequence alignment of CYP450-PF00067 family

<p>A MSA of all the CYP450 family members (pfam identifier PF00067; 1635 sequences as per UniProt at 25th November 2021)</p>

restrictedNov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record