Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
44
datasets available to search
ShareScore release 0.7.1
Dataset results
44 results for “multiple sequence alignment”
Multiple alignment of DNA-B sequences from CMMGV, EACMCV, EACMV, EACMKV, EACMMV, EACMZV, SACMV (7 "species")
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Multiple alignment of ACMV and ACMBFV DNA-B sequences
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method
<p>This dataset contains a GitHub repository containing all the data, analysis, Nextflow workflows and Jupyter notebooks to replicate the manuscript titled "Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method".</p> <p>It also contains the Multiple Sequence Alignments (MSAs) generated and well as the main figures and tables from the manuscript.</p> <p>The repository is also available at GitHub (https://github.com/cbcrg/dpa-analysis) release `v1.2`.</p> <p>For details on how to use the regressive alignment algorithm, see the T-Coffee software suite (https://github.com/cbcrg/tcoffee).</p>
Multiple Sequence Alignment of a diverse dataset with 1788 Mycobacterium tuberculosis isolates
<p><strong>Multiple Sequence Alignment of a diverse dataset with 1788 <em>Mycobacterium tuberculosis</em> isolates used for <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> benchmarking</strong></p> <p>The dataset comprises whole-genome sequence data published by <a href="https://doi.org/10.1016/S1473-3099(15)00062-6">Walker et al. 2015</a>. For the multiple sequence analysis, we proceeded as follows:</p> <ol> <li>Reads were downloaded from ENA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA282721">PRJNA282721</a> (accessed on March 16<sup>th</sup>, 2023) and trimmed using Trimmomatic (<a href="https://pubmed.ncbi.nlm.nih.gov/24695404/">Bolger et al., 2014</a>) with <a href="https://github.com/B-UMMI/INNUca">INNUca</a> default settings;</li> <li>Quality-processed reads were individually mapped against the H37Rv reference genome (Genbank accession: <a href="https://www.ncbi.nlm.nih.gov/nuccore/NC_000962.3/">NC_000962.3</a>) using <a href="https://github.com/tseemann/snippy">Snippy</a> v4.5.1 and SNP-calling was performed on variant sites with the following criteria: a minimum proportion of reads differing from the reference of 70%, a minimum mapping quality of 30 and a minimum coverage for SNP calling of 10;</li> <li>A full alignment was extracted using Snippy’s core module (snippy-core), with masking of SNPs falling within known <em>M. tuberculosis</em> genomic regions with high GC content, repetitive elements and resistance-associated positions (corresponding to ~8% of the genome), as previously described for surveillance purposes (<a href="https://pubmed.ncbi.nlm.nih.gov/30948181/">Macedo et al., 2019</a>);</li> <li><em>M. tuberculosis </em>lineages were determined using tb-profiler v4.4.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31234910/">Phelan et al., 2019</a>), with samples from the <em>M. tuberculosis</em> complex other than <em>M. tuberculosis</em>, representing a mix of multiple lineages, or with less than 95% of mapped positions in the reference, being excluded;</li> <li>A filtered alignment comprising the maximum number of informative sites (88,562 nucleotide sites with at least one mutation in a given sequence) was extracted from the full alignment using the alignment_processing.py v1.1.0 (default settings) of <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a>, and then used as input for the benchmarking.</li> </ol> <p>In this repository, we provide two alignment files:</p> <ul> <li>Core_MTB_1787_strs.full.aln: this corresponds to the full multiple sequence alignment comprising 1787 samples and the reference (corresponding to the point 4 of the methodology).</li> <li>MTb_original_align_profile.fasta: this corresponds to the multiple sequence alignment comprising 1787 samples and the reference and only presenting the alignment informative sites (corresponding to the point 5 of the methodology)</li> </ul>
Multiple sequence alignments of sensor histidine kinases and response regulators
<p>The two FASTA files contain multiple sequence alignments of sensor histidine kinase and response regulator sequences. The source sequences were obtained by BLAST, clustered with usearch and aligned with muscle. More details to be found in Multamäki et al. 2021.</p>
Text-fig. 5. Multiple sequence alignment of mtDNA from ancient bone and recent greyhound, (Gundry et al. 2007) primer pair A – F15.719 and R16.114. in Genetic Analysis Of Possibly The Oldest Greyhound Remains Within The Territory Of The Czech Republic As Proof Of A Local Elite Presence At Chotěbuz-Podobora Hillfort In The 8 -9 Century Ad
Text-fig. 5. Multiple sequence alignment of mtDNA from ancient bone and recent greyhound, (Gundry et al. 2007) primer pair A – F15.719 and R16.114.
CONCATENATING sample files to prepare for multiple sequence alignment in Galaxy
<p>These are a few sample files to practice the correct way to concatenate files with the reference strain at the top, in order to continue with the next step, which is doing a multiple sequence alignment.</p>
Data from: Multiple genotypes of Phelipanche ramosa indicate repeated introductions to the Americas: Sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Multiple sequence alignment for the native Norwegian vascular plant phylogeny
<p>Methods: We produced a multi-locus Maximum Likelihood (ML) phylogeny using a combination of newly produced DNA sequences from herbarium specimens and sequences available from public repositories. We combined the phylogeny with species occurrence data to estimate phylogenetic diversity and phylogenetic endemism across Norway, using a spatial randomization to judge statistical significance. We used multiple-model inference to identify environmental variables that contributed the most to the patterns of phylogenetic diversity. Finally, we estimated phylogenetic turnover and used this to identify Norwegian plant assemblages in terms of composition and evolutionary history.<br> <br> Results: Our ML phylogeny contained 87% of all currently described native Norwegian vascular plants. Assemblages were phylogenetically overdispersed in warmer and wetter regions of Norway, as well as in regions with a longer post-glacial history. In cold and dry regions, plant assemblages were phylogenetically clustered, and characterised by neo-endemism, while the mild and wet regions were characterised by both paleo- and neo-endemism. Phylogenetic diversity was positively correlated with summer temperature and habitat heterogeneity, and peaked in the southeast of Norway.<br> <br> Main conclusions: Both contemporary ecological factors (climate and habitat heterogeneity), and post-glacial history seem to have shaped the phylogenetic structure of the flora of Norway. The flora in the far north of Norway appear to be a result of recent diversification while the coastal regions are assemblages of deeper lineages. Our results suggest that there is an evolutionary signal in the distribution of the Norwegian vascular flora.</p>
Multiple sequence alignments of full-length L1 elements with evidence of retrotransposition activity.
<p>DNA sequences for full-length L1 elements showing evidence of retrotransposition activity were aligned using MUSCLE v3.8 with default number of iterations. Manual inspection of the multiple alignments was performed with Jalview v2.11 in order to remove upstream and downstream spurious sequences.</p> <ul> <li><em>selected_active_L1sOK_aligned_trimmed.fa</em> file includes aligned DNA sequences for 86 full-length L1Hs with medium-high activity and for 1 active L1Pt (outgroup).</li> <li><em> all_active_L1sOK_aligned_trimmed.fa</em> file includes aligned DNA sequences for 143 active L1Hs and for 1 active L1Pt (outgroup).</li> </ul>
Multiple sequence alignments: Detection and isolation of a new member of Burkholderiaceae‑related endofungal bacteria from Saksenaea boninensis sp. nov., a new thermotolerant fungus in Mucorales
<p><strong>Methods:</strong></p><p>Nucleotide sequences were aligned independently for each region using MAFFT v7.212 (Katoh and Standley, 2013). The obtained alignment blocks were subject to Gblocks 0.91b (Castresana, 2000) to remove poorly aligned positions with the relaxed selection setting described in Talavera & Castresana (2007) using the following parameters (-t = d -b2 = 9 -b3 = 10 -b4 = 5 -b5 = h). After automatically removing gaps, the alignment blocks were viewed using MEGA 6.06 software (Tamura et al., 2013) and poorly aligned positions at either end of the alignments were removed manually. Pairwise distances of the nucleotide sequences (ITS2, ITS1-5.8S-ITS2, LSU, and tef1) of the ex-type strains of seven <i>Saksenaea</i> spp. and the representative isolate <i>S. boninensis</i> Sak4 were calculated by MEGA 6.06 software (Tamura et al. 2013). Multiple sequence alignment of 16S rRNA gene of the family <i>Burkholderiaceae</i> was prepared for the phylogeny of a bacterial endosymbiont. Multiple sequence alignments of ITS, LSU, and tef1 genes of <i>Saksenaea</i> spp. (Mucorales) were separately prepared for the phylogeny of a fungal host. Concatenated dataset of these genes were also prepared. All nucleotide sequences were retrieved from GenBank (See "Sequence_ID.csv" and taxon names of each alignment). </p><p> </p><p><strong>Description of files:</strong></p><p><strong>A. Phylogeny of the family </strong><i><strong>Burkholderiaceae</strong></i><strong> (Bacterial endosymbiont):</strong></p><p>1. Burkholderiaceae_16S_RAW.fasta</p><p>Non-aligned dataset of 16S rRNA gene of the family <i>Burkholderiaceae</i>.</p><p> </p><p>2. Burkholderiaceae_16S_aligned.fasta</p><p>Aligned dataset of 16S rRNA gene of the family <i>Burkholderiaceae</i>.</p><p> </p><p><strong>B. Phylogenies of </strong><i><strong>Saksenaea</strong></i><strong> spp. (Fungal host):</strong></p><p>1. Sequence_ID_v2.csv</p><p>Taxon names, accession numbers, and sequence ID for the concatenated multiple sequence alignment are listed.</p><p> </p><p>2. Saksenaea_ITS_RAW_v2.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. </p><p> </p><p>3. Saksenaea_ITS_aligned_v2.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. Only used for ITS1-5.8S-ITS2 phylogeny.</p><p> </p><p>4. Saksenaea_LSU_RAW_v2.fasta</p><p>Non-aligned dataset of LSU gene region of <i>Saksenaea</i> spp. </p><p> </p><p>5. Saksenaea_LSU_aligned_v2.fasta</p><p>Aligned dataset of LSU gene region of <i>Saksenaea</i> spp. Only used for LSU phylogeny.</p><p> </p><p>6. Saksenaea_tef1_RAW_v2.fasta</p><p>Non-aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. </p><p> </p><p>7. Saksenaea_tef1_aligned_v2.fasta</p><p>Aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. Only used for tef1 phylogeny.</p><p> </p><p><strong><Concatenated dataset 1 (ITS2, LSU, tef1)></strong></p><p>8. Saksenaea_ITS2_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 1.</p><p> </p><p>9. Saksenaea_ITS2_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 1.</p><p> </p><p>10. Saksenaea_LSU_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of LSU gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>11. Saksenaea_LSU_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of LSU gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>12. Saksenaea_tef1_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>13. Saksenaea_tef1_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>14. Saksenaea_ITS2_LSU_tef1_concatenated_dataset1.fasta</p><p>Concatenated dataset of three multiple sequence alignments (9, 11, and 13). This concatenated dataset was used for the main phylogeny of <i>Saksenaea</i> spp.</p><p> </p><p><strong><Concatenated dataset 2 (ITS1-5.8S-ITS2, LSU, tef1)></strong></p><p>15. Saksenaea_ITS_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 2.</p><p> </p><p>16. Saksenaea_ITS_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 2.</p><p>Blank sequences were inserted for five isolates of <i>Saksenaea longicolla</i> after the alignment.</p><p> </p><p>17.Saksenaea_ITS_LSU_tef1_concatenated_dataset2.fasta</p><p>Concatenated dataset of three multiple sequence alignments (15, 11, and 13). This concatenated dataset was used for the main phylogeny of <i>Saksenaea</i> spp.</p><p> </p><p><strong>C. Pairwise distances of the ex-type strains of </strong><i><strong>Saksenaea</strong></i><strong> spp.</strong></p><p>1. Saksenaea_ITS_type_RAW.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>2. Saksenaea_ITS2_type_aligned.fasta</p><p>Aligned dataset of ITS2 region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>3.Saksenaea_ITS_type_aligned.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of the ex-type strains of <i>Saksenaea</i> spp. without <i>Saksenaea longicolla</i>.</p><p> </p><p>4. Saksenaea_LSU_type_RAW.fasta</p><p>Non-aligned dataset of LSU gene region of the ex-type strains of <i>Saksenaea </i>spp.</p><p> </p><p>5. Saksenaea_LSU_type_aligned.fasta</p><p>Aligned dataset of LSU gene region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>6. Saksenaea_tef1_type_RAW.fasta</p><p>Non-aligned dataset of tef1 gene region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>7. Saksenaea_tef1_type_aligned.fasta</p><p>Aligned dataset of tef1 gene region of the ex-type strains of <i>Saksenaea</i> spp.</p>
Imputed Multiple Sequence Alignment used in 'Estimating the relative proportions of SARS-CoV-2 strains from wastewater samples'
<p>Multiple Sequence Alignment of imputed SARS-CoV-2 sequences used in 'Estimating the relative proportions of SARS-CoV-2 strains from wastewater samples'</p>
Phylogenetic and recombination analysis of adenovirus isolates reveals discordance between serotype and phylogeny: Multiple sequence alignments
<p><strong>Background</strong></p> <p>Human adenovirus (HAdV) infections are caused by seven mastadenovirus species (A-G) and are the source for a variety of pathologies including gastrointestinal, respiratory, neurological, and ocular disease. While HAdV-D is the most common cause of adenovirus ocular infections, human adenoviruses B and E have also been isolated from the eye. </p> <p><strong>Results</strong></p> <p>In the course of classifying three new atypical ocular adenovirus samples, taken from the vitreous humor, we found that all three isolates were HAdV-B species, with isolate BP-AdV1 sorting with the B1 clade, and isolates BP-AdV2 and BP-AdV3 grouping into the B2 clade. The three Bascom Palmer HAdV-B genomes were then combined with over 300 HAdV-B genome sequences, including 9 ocular HAdV-B genome sequences. The whole genome phylogenetic analysis showed that 9 of the 11 ocular sequences grouped into the B1 clade, forming two clusters within B1. Attempts to categorize the penton, hexon and fiber serotypes using phylogeny of the three Bascom Palmer samples were inconclusive due to incongruence between serotype and phylogeny in the dataset. Recombination analysis using a subset of HAdV-B strains to generate a hybridization network detected recombination between non-human primate and human derived strains, recombination between one HAdV-B strain and the HAdV-E outgroup and limited recombination between the B1 and B2 clades. </p> <p><strong>Conclusions</strong></p> <p>The discordance between serotype and phylogeny detected in this study suggests that the current penton/hexon/fiber-based classification mechanism does not accurately describe the natural history and phylogenetic relationships amongst adenoviruses. A new adenovirus strain classification strategy may be beneficial to the field.</p>
Multiple Sequence Alignments for Octopus bocki
<p><span>Multiple sequence alignments were created in MEGA11: Molecular Evolutionary Genetics Analysis version 11 (<em>Whelan and Goldman, 2001</em>)</span><span> using Muscle default parameters (UPGMA cluster method with -2.9 gap open and 0 gap extension penalties)</span></p> <p><em><span>Whelan, S. and Goldman, N. (2001). A general empirical model of protein evolution derived from multiple protein families using a maximum-likelihood approach. Molecular Biology and Evolution 18:691-699.</span></em></p>
Multiple sequence alignment of USP Zf-UBD proteins
<p>Using Molsoft's ICM-Pro, a multiple sequence alignment of USP Zf-UBDs was done against HDAC6 Zf-UBD. </p>
Cephalopod retinal development shows vertebrate-like mechanisms of neurogenesis: Multiple sequence alignments and phylogenetic trees
<p>Coleoid cephalopods, including squid, cuttlefish and octopus, have large and complex nervous systems and camera-type eyes that are comparable only to features that have independently evolved in the vertebrate lineage. The changes in development that result in the evolution of nervous system size and diversity of neural cell-types are not well understood. Here, we have pioneered live-imaging techniques and performed functional interrogation to show the squid, <em>Doryteuthis</em> <em>pealeii</em>, utilizes mechanisms during retinal neurogenesis that are hallmarks of vertebrate processes. Given the convergent evolution of elaborate visual systems in cephalopods and vertebrates, these results reveal common mechanisms that underlie the growth of highly proliferative neurogenic primordia that may alter ontogenetic allometry and contribute to the evolution of complex nervous systems.</p>
Multiple sequence alignments of newly reconstructed and published cervid and human mtDNA
<p><span>Assigning prehistoric objects to specific individuals is usually impossible outside of burial contexts. Here we present a non-destructive method for gradually releasing DNA from ancient bone and tooth artifacts. Application of the method to an Upper Paleolithic deer tooth pendant from Denisova Cave (Russia) resulted in the recovery of DNA from both the deer and a female human individual. Genetic dates obtained from the deer and human mitochondrial genomes estimate the age of the pendant at approximately 20,000 to 24,000 years. Nuclear DNA from its presumed maker or wearer shows strong affinities to contemporaneous Ancient North Eurasian individuals previously found further east in Siberia. Our work opens up new possibilities for linking cultural and genetic records in prehistoric archaeology.</span></p>
extHomFam v37.0: structural benchmark for protein multiple sequence alignments
<p>extHomFam v37.0 was constructed by combining Homstrad reference alignments (2 December 2023 release) with Pfam 37.0 (UniProt release) families containing at least 200 sequences. Homstrad entries with less than 3 reference sequences and those pointing to dead Pfam families were discarded.</p> <p> </p>
Multiple sequence alignment for the native Norwegian vascular plant phylogeny
Open the record for dataset details and reuse information.
Multiple sequence alignments of newly reconstructed and published cervid and human mtDNA
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.