Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
175
datasets available to search
ShareScore release 0.9.0
Dataset results
175 results for “sequence alignments”
Fig. 6 in Total Evidence, Sequence Alignment, Evolution of Polychrotid Lizards, and a Reclassification of the Iguania (Squamata: Iguania)
Fig. 6. Phylogenetic hypotheses for the Polychrotidae of Etheridge and de Queiroz (1988), and Frost and Etheridge (1989). Circles represent hypothesized rooting points.
Fig. 1 in Total Evidence, Sequence Alignment, Evolution of Polychrotid Lizards, and a Reclassification of the Iguania (Squamata: Iguania)
Fig. 1. Single tree obtained on molecularonly data (length = 2332; CI = 0.23; RI = 0.68). Numbers on branches are Bremer values.
CONCATENATING sample files to prepare for multiple sequence alignment in Galaxy
<p>These are a few sample files to practice the correct way to concatenate files with the reference strain at the top, in order to continue with the next step, which is doing a multiple sequence alignment.</p>
pWCP alignment sequences
<p>Alignment sequences of pWCP across locations.</p>
Data from: Multiple genotypes of Phelipanche ramosa indicate repeated introductions to the Americas: Sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Scripts and data sets associated with: On testing homogeneity of the evolutionary process using alignments of homologous sequences
Open the record for dataset details and reuse information.
Aligned DNA sequence matrix for phylogenetic analyses in the article "New species of fossorial salamanders of the genus Oedipina (Plethodontidae) from the northwestern Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "New species of fossorial salamanders of the genus Oedipina (Plethodontidae) from the northwestern Ecuador". The matrix is in NEXUS format.</p> <p>Gene partitions are arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>16S = 4- 789 1761- 1891 ;<br> ND1-codonPos1 = 790-1759\3;<br> ND1-codonPos2 = 791-1760\3;<br> ND1-codonPos3 = 792-1758\3;<br> CytB-codonPos1 = 1893-2274\3;<br> CytB-codonPos2 = 1894-2275\3;<br> CytB-codonPos3 = 1892-2276\3;</p>
Multiple sequence alignment for the native Norwegian vascular plant phylogeny
<p>Methods: We produced a multi-locus Maximum Likelihood (ML) phylogeny using a combination of newly produced DNA sequences from herbarium specimens and sequences available from public repositories. We combined the phylogeny with species occurrence data to estimate phylogenetic diversity and phylogenetic endemism across Norway, using a spatial randomization to judge statistical significance. We used multiple-model inference to identify environmental variables that contributed the most to the patterns of phylogenetic diversity. Finally, we estimated phylogenetic turnover and used this to identify Norwegian plant assemblages in terms of composition and evolutionary history.<br> <br> Results: Our ML phylogeny contained 87% of all currently described native Norwegian vascular plants. Assemblages were phylogenetically overdispersed in warmer and wetter regions of Norway, as well as in regions with a longer post-glacial history. In cold and dry regions, plant assemblages were phylogenetically clustered, and characterised by neo-endemism, while the mild and wet regions were characterised by both paleo- and neo-endemism. Phylogenetic diversity was positively correlated with summer temperature and habitat heterogeneity, and peaked in the southeast of Norway.<br> <br> Main conclusions: Both contemporary ecological factors (climate and habitat heterogeneity), and post-glacial history seem to have shaped the phylogenetic structure of the flora of Norway. The flora in the far north of Norway appear to be a result of recent diversification while the coastal regions are assemblages of deeper lineages. Our results suggest that there is an evolutionary signal in the distribution of the Norwegian vascular flora.</p>
Multiple sequence alignments of full-length L1 elements with evidence of retrotransposition activity.
<p>DNA sequences for full-length L1 elements showing evidence of retrotransposition activity were aligned using MUSCLE v3.8 with default number of iterations. Manual inspection of the multiple alignments was performed with Jalview v2.11 in order to remove upstream and downstream spurious sequences.</p> <ul> <li><em>selected_active_L1sOK_aligned_trimmed.fa</em> file includes aligned DNA sequences for 86 full-length L1Hs with medium-high activity and for 1 active L1Pt (outgroup).</li> <li><em> all_active_L1sOK_aligned_trimmed.fa</em> file includes aligned DNA sequences for 143 active L1Hs and for 1 active L1Pt (outgroup).</li> </ul>
Improving phonetic alignment by handling secondary sequence structures
<p>Supplementary material accompanying the paper "Improving phonetic alignment by handling secondary sequence structures".</p> <p>The data consists of 5 files:</p> <ul> <li>gold_standard.psa : the gold standard used in the analysis in PSA format</li> <li> sca-secondary.psa : the output of the algorithm with the secondary extension</li> <li>sca-traditional.psa : the output of the traditional algorithm</li> <li>sca-secondary-diff.psa : the differences of the secondary extension compared to the GS</li> <li>sca-traditional-diff.psa : the differences of the traditional algorithm compared to the GS</li> </ul> <p>For a description of the file-format used in this dataset, please refer to the LingPy tutorial under http://lingpy.org.</p>
Neutralization Data and Aligned ENV Sequences for Predicting Antibody Affinities using Artificial Neural Networks
<p>Sample file with neutralization data (IC<sub>50</sub>) for different antibodies and viral strains, adapted from J. Huang, G. Ofek, L. Laub, M. K. Louder, N. A. Doria-Rose, N. S. Longo, H. Imamichi, R. T. Bailer, B. Chakrabarti, S. K. Sharma, S. M. Alam, T. Wang, Y. Yang, B. Zhang, S. A. Migueles, R. Wyatt, B. F. Haynes, P. D. Kwong, J. R. Mascola, and M. Connors, “Broad and potent neutralization of HIV-1 by a gp41-specific human antibody.,” <em>Nature</em>, vol. 491, no. 7424, pp. 406–12, Nov. 2012.</p> <p> </p> <p>Aligned ENV sequences downloaded from the HIV Sequence Database (www.hiv.lanl.gov/content/sequence/HIV/mainpage.html). There are 4907 sequences and the alignment length is 1369.</p>
Aligned and trimmed 16S and COI DNA sequences of Oceaniidae (Hydrozoa)
<p>Aligned and trimmed 16S and COI sequences of Oceaniidae (Hydrozoa) used for the study "The polyps of <em>Oceania armata</em> identified by DNA barcoding (Cnidaria, Hydrozoa)"</p> <p>Format is Fasta, files are text files</p>
Multiple sequence alignments: Detection and isolation of a new member of Burkholderiaceae‑related endofungal bacteria from Saksenaea boninensis sp. nov., a new thermotolerant fungus in Mucorales
<p><strong>Methods:</strong></p><p>Nucleotide sequences were aligned independently for each region using MAFFT v7.212 (Katoh and Standley, 2013). The obtained alignment blocks were subject to Gblocks 0.91b (Castresana, 2000) to remove poorly aligned positions with the relaxed selection setting described in Talavera & Castresana (2007) using the following parameters (-t = d -b2 = 9 -b3 = 10 -b4 = 5 -b5 = h). After automatically removing gaps, the alignment blocks were viewed using MEGA 6.06 software (Tamura et al., 2013) and poorly aligned positions at either end of the alignments were removed manually. Pairwise distances of the nucleotide sequences (ITS2, ITS1-5.8S-ITS2, LSU, and tef1) of the ex-type strains of seven <i>Saksenaea</i> spp. and the representative isolate <i>S. boninensis</i> Sak4 were calculated by MEGA 6.06 software (Tamura et al. 2013). Multiple sequence alignment of 16S rRNA gene of the family <i>Burkholderiaceae</i> was prepared for the phylogeny of a bacterial endosymbiont. Multiple sequence alignments of ITS, LSU, and tef1 genes of <i>Saksenaea</i> spp. (Mucorales) were separately prepared for the phylogeny of a fungal host. Concatenated dataset of these genes were also prepared. All nucleotide sequences were retrieved from GenBank (See "Sequence_ID.csv" and taxon names of each alignment). </p><p> </p><p><strong>Description of files:</strong></p><p><strong>A. Phylogeny of the family </strong><i><strong>Burkholderiaceae</strong></i><strong> (Bacterial endosymbiont):</strong></p><p>1. Burkholderiaceae_16S_RAW.fasta</p><p>Non-aligned dataset of 16S rRNA gene of the family <i>Burkholderiaceae</i>.</p><p> </p><p>2. Burkholderiaceae_16S_aligned.fasta</p><p>Aligned dataset of 16S rRNA gene of the family <i>Burkholderiaceae</i>.</p><p> </p><p><strong>B. Phylogenies of </strong><i><strong>Saksenaea</strong></i><strong> spp. (Fungal host):</strong></p><p>1. Sequence_ID_v2.csv</p><p>Taxon names, accession numbers, and sequence ID for the concatenated multiple sequence alignment are listed.</p><p> </p><p>2. Saksenaea_ITS_RAW_v2.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. </p><p> </p><p>3. Saksenaea_ITS_aligned_v2.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. Only used for ITS1-5.8S-ITS2 phylogeny.</p><p> </p><p>4. Saksenaea_LSU_RAW_v2.fasta</p><p>Non-aligned dataset of LSU gene region of <i>Saksenaea</i> spp. </p><p> </p><p>5. Saksenaea_LSU_aligned_v2.fasta</p><p>Aligned dataset of LSU gene region of <i>Saksenaea</i> spp. Only used for LSU phylogeny.</p><p> </p><p>6. Saksenaea_tef1_RAW_v2.fasta</p><p>Non-aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. </p><p> </p><p>7. Saksenaea_tef1_aligned_v2.fasta</p><p>Aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. Only used for tef1 phylogeny.</p><p> </p><p><strong><Concatenated dataset 1 (ITS2, LSU, tef1)></strong></p><p>8. Saksenaea_ITS2_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 1.</p><p> </p><p>9. Saksenaea_ITS2_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 1.</p><p> </p><p>10. Saksenaea_LSU_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of LSU gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>11. Saksenaea_LSU_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of LSU gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>12. Saksenaea_tef1_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>13. Saksenaea_tef1_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>14. Saksenaea_ITS2_LSU_tef1_concatenated_dataset1.fasta</p><p>Concatenated dataset of three multiple sequence alignments (9, 11, and 13). This concatenated dataset was used for the main phylogeny of <i>Saksenaea</i> spp.</p><p> </p><p><strong><Concatenated dataset 2 (ITS1-5.8S-ITS2, LSU, tef1)></strong></p><p>15. Saksenaea_ITS_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 2.</p><p> </p><p>16. Saksenaea_ITS_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 2.</p><p>Blank sequences were inserted for five isolates of <i>Saksenaea longicolla</i> after the alignment.</p><p> </p><p>17.Saksenaea_ITS_LSU_tef1_concatenated_dataset2.fasta</p><p>Concatenated dataset of three multiple sequence alignments (15, 11, and 13). This concatenated dataset was used for the main phylogeny of <i>Saksenaea</i> spp.</p><p> </p><p><strong>C. Pairwise distances of the ex-type strains of </strong><i><strong>Saksenaea</strong></i><strong> spp.</strong></p><p>1. Saksenaea_ITS_type_RAW.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>2. Saksenaea_ITS2_type_aligned.fasta</p><p>Aligned dataset of ITS2 region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>3.Saksenaea_ITS_type_aligned.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of the ex-type strains of <i>Saksenaea</i> spp. without <i>Saksenaea longicolla</i>.</p><p> </p><p>4. Saksenaea_LSU_type_RAW.fasta</p><p>Non-aligned dataset of LSU gene region of the ex-type strains of <i>Saksenaea </i>spp.</p><p> </p><p>5. Saksenaea_LSU_type_aligned.fasta</p><p>Aligned dataset of LSU gene region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>6. Saksenaea_tef1_type_RAW.fasta</p><p>Non-aligned dataset of tef1 gene region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>7. Saksenaea_tef1_type_aligned.fasta</p><p>Aligned dataset of tef1 gene region of the ex-type strains of <i>Saksenaea</i> spp.</p>
.bam alignment files of Illumina and ONT sequencing of pREF plasmid
<p>The expression of genes encompasses their transcription into mRNA followed by translation into protein. In recent years, next-generation sequencing and mass spectrometry methods have profiled DNA, RNA and protein abundance in cells. However, there are currently no reference standards that are compatible across these genomic, transcriptomic and proteomic methods, and provide an integrated measure of gene expression. Here, we use synthetic biology principles to engineer a multi-omics control, termed <em>pREF</em>, that can act as a universal molecular standard for next-generation sequencing and mass spectrometry methods. The <em>pREF</em> sequence encodes 21 synthetic genes that can be <em>in vitro</em> transcribed into spike-in mRNA controls, and <em>in vitro</em> translated to generate matched protein controls. The synthetic genes provide qualitative controls that can measure sensitivity and quantitative accuracy of DNA, RNA and peptide detection. We demonstrate the use of <em>pREF</em> in metagenome DNA sequencing and RNA sequencing experiments and evaluate the quantification of proteins using mass spectrometry. Unlike previous spike-in controls, <em>pREF</em> can be independently propagated and the synthetic mRNA and protein controls can be sustainably prepared by recipient laboratories using common molecular biology techniques. Together, this provides the first universal synthetic standard able to integrate genomic, transcriptomic and proteomic methods.</p>
Alignments of reindeer/caribou mitogenome sequences
<p>Climate warming at the end of the last glacial period had profound effects on the distribution of cold-adapted species. As their range shifted towards northern latitudes, they were able to colonise previously glaciated areas, including remote Arctic islands. However, there is still uncertainty about their colonisation routes and timings. At the end of the last ice age, reindeer/caribou (<em>Rangifer tarandus</em>) expanded to the Holarctic region and colonised the archipelagos of Svalbard and Franz Josef Land. Earlier studies have proposed two possible colonisation routes, either from the Eurasian mainland or from Canada via Greenland. Here, we used 174 ancient, historical, and modern mitogenomes to reconstruct the phylogeny of reindeer across its whole range and to infer the colonisation route of the Arctic islands. Our data shows a close affinity among Svalbard, Franz Josef Land, and Novaya Zemlya reindeer. We also found tentative evidence for positive selection in the mitochondrial gene ND4, which is possibly associated with increased heat production. Our results thus support a colonisation of Arctic archipelagos from the Eurasian mainland and provide some insights into the evolutionary history and adaptation of the species to its High Arctic habitat. </p>
Alignment of mitogenome sequences (FASTA file) for a paleogenomic investigation of overharvest implications in an endemic wild reindeer subspecies
<p>Overharvest can severely reduce the abundance and distribution of a species and thereby impact its genetic diversity and threaten its future viability. Overharvest remains an ongoing issue for Arctic mammals, which due to climate change now also confront one of the fastest changing environments on Earth. The high-Arctic Svalbard reindeer (<em>Rangifer tarandus platyrhynchus</em>), endemic to Svalbard, experienced a harvest-induced demographic bottleneck that occurred during the 17–20th centuries. Here we investigate changes in genetic diversity, population structure, and gene-specific differentiation during and after this overharvesting event. Using whole-genome shotgun sequencing, we generated the first ancient and historical nuclear (n = 11) and mitochondrial (n = 18) genomes from Svalbard reindeer (up to 4000 BP) and integrated these data with a large collection of modern genome sequences (n = 90), to infer temporal changes. We show that hunting resulted in major genetic changes and restructuring in reindeer populations. Near-extirpation followed by pronounced genetic drift have altered the allele frequencies of important genes contributing to diverse biological functions. Median heterozygosity was reduced by 23%, while the mitochondrial genetic diversity was reduced only to a limited extent, likely due to already low pre-harvest diversity and a complex post-harvest recolonization process. Such genomic erosion and genetic isolation of populations due to past anthropogenic disturbance will likely play a major role in metapopulation dynamics (i.e., extirpation, recolonization) under further climate change. Our results from a high-arctic case study therefore emphasize the need to understand the long-term interplay of past, current, and future stressors in wildlife conservation.</p>
The alignments of chloroplast genome sequences and nuclear ribosomal DNA fragments of six oak species sampled in the hot-dry valley of the Jinsha River, southwestern China
<p>Both chloroplast (cp) genome sequences and nuclear ribosomal (nr) DNA were assembled using GetOrganelle v.1.7.6.1 for 18 oak trees sampled in the Panzhihua Cycad National Nature Reserve, Sichuan Province, China. These trees belong to six oak species, including Quercus cocciferoides, Q. dolicholepis, Q. franchetii, Q. griffithii, Q. longispica, and Q. variabilis. We used PhyloSuite v.1.1.152 to extract coding sequences (CDSs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) of the 18 oak cp genomes. These sequences were aligned separately using MAFFT v.7.3.13 and manually adjusted with BioEdit v.7.2.5. Length variations in mononucleotide repeats were excluded and inversions were replaced with their reverse complements because of their tendency for homoplasy. Other indels were coded as binary characters according to the simple gap coding method using GapCoder. Separate assignments were concatenated according to their respective positions in the cp genome to obtain the alignments of LSC, SSC, IRb, and the whole cp genome.</p>
Imputed Multiple Sequence Alignment used in 'Estimating the relative proportions of SARS-CoV-2 strains from wastewater samples'
<p>Multiple Sequence Alignment of imputed SARS-CoV-2 sequences used in 'Estimating the relative proportions of SARS-CoV-2 strains from wastewater samples'</p>
Arabidopsis thaliana Col-CEN complete Chromosome 2 numt sequences and alignments
<p>Data associated with the assembly of complete chromosome 2 nuclear insertion of mitochondrial DNA (numt) from the <em>Arabidopsis thaliana</em> accession Columbia (Col-CEN). A full report of this project can be obtained in a manuscript titled "<strong>Complete sequence of a 641-kb insertion of mitochondrial DNA in the <em>Arabidopsis thaliana </em>nuclear genome</strong>".</p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Dos nuevas especies del grupo Pristimantis boulengeri (Anura: Strabomantidae) de la cuenca alta del río Napo, Ecuador" by Bejarano, et al.
<p>Matrix in nexus format that include sequences of 16S (1-1295), RAG1 (1296-1922), 12S (1923-3306), and COI (3307-3984) for 97 specimens belonging to the genus <em>Pristimantis</em>, in addition to specimens of <em>Strabomantis</em> and <em>Niceforonia</em> as outgroups.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.