Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
175
datasets available to search
ShareScore release 0.9.0
Dataset results
175 results for “sequence alignments”
Multiple alignment of DNA-B sequences from CMMGV, EACMCV, EACMV, EACMKV, EACMMV, EACMZV, SACMV (7 "species")
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Multiple alignment of ACMV and ACMBFV DNA-B sequences
<p>All sequences available in GenBank as of 2019-06-03 were downloaded via the Taxonomy Browser interface. Sequence names were normalized/simplified and orientations of these circular sequences were standardized to begin at the replication origin nick site. Sequences were aligned with MUSCLE and alignments were adjusted with SeAl (A. Rambaut) and AliView (A. Larsson).</p> <p>These results are described in a paper by Crespo-Bellido et al. (2021) https://doi.org/10.1128/JVI.00541-21</p>
Datasets of sequences, alignments and structural models generated for the structural prediction of complexes mediated by intrinsically disordered regions.
<p>This repository contains input and ouput files used and generated for the scanning of intrinsically disordered region and the prediction of their binding sites to receptor proteins using the <a href="https://github.com/i2bc/SCAN_IDR">SCAN_IDR</a> pipeline with AlphaFold2-Multimer.</p><p>It contains two archives: </p><ol><li><a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> dedicated to the analysis of a dataset of 42 protein complexes non redundant with the dataset used for AlphaFold2 training,</li><li><a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> dedicated to the analysis of 923 complexes from the ELM database.</li></ol><p>These data can be used to rerun specific sections of the pipeline and scripts provided in: <a href="https://github.com/i2bc/SCAN_IDR">https://github.com/i2bc/SCAN_IDR</a></p><h4><strong>Dataset of 42 non redundant complexes</strong></h4><p>The first archive <a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> contains 3 compressed directories and a README file detailing their contents :</p><ul><li>the initial raw sequence and alignment data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the input and output data of every Alphafold run for every complex -> DIRECTORY <strong>af2_runs/</strong></li><li>the native reference structures -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p>The protein-peptide complex cases have been assigned a distinct index number, from 1 to 42, consistent across the several directories of the archive. Their corresponding directories are labelled as <i><index>_<pdbcode></i>.</p><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.2</i></p><h4><strong>Dataset of 923 complexes selected from the ELM database</strong></h4><p>The second archive <a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> contains input and ouput files used and generated for the analysis of 923 Eukaryotic Linear Motifs (ELM) database entries.</p><p>Each ELM entry is indexed with specific integer id and is composed of a receptor and a ligand protein. </p><p>The archive contains a Table associating ELM indexes with the ELM entry information, 5 directories and a README file detailing their contents:</p><ul><li>the table describing ELM entries -> FILE <strong>Table_923ELM_uid_delimitations_info_for_archive.txt</strong></li><li>the initial raw sequence and multiple sequence alignment (MSA) data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the concatenated MSA model for every ELM complex and protocol used -> DIRECTORY <strong>af2_elm_coali_inputs/</strong></li><li>the best model of every AF2 protocol for every complex according to the AF2 -> DIRECTORY <strong>af2_elm_models/</strong></li><li>the best model cut in the ligand part to select only the ELM motifs as used for the evaluation of the models -> DIRECTORY <strong>elm_cut_models/</strong></li><li>the reference structures used for the evaluation of the models -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.3</i></p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Three new species of Torrent Treefrogs (Anura: Hylidae) of the Hyloscirtus bogotensis group from the eastern Andean slopes and the biogeographic history of the genus"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "Three new species of Torrent Treefrogs (Anura: Hylidae) of the Hyloscirtus bogotensis group from the Amazon foothills and the biogeographic history of the genus"</p> <p>The matrix is in NEXUS format and has 3259 bp and 25 terminals.</p> <p>Partitions are as follows:</p> <div>charset 12S = 1-955;</div> <div>charset ND1_nonCoding1 = 956-1279;</div> <div>charset ND1_Pos1 = 1280-2240\3;</div> <div>charset ND1_Pos2 = 1281-2241\3;</div> <div>charset ND1_Pos3 = 1282-2242\3;</div> <div>charset ND1_nonCoding2 = 2243-2361;</div> <div>charset cmyc_Pos1 = 2362-2779\3;</div> <div>charset cmyc_Pos2 = 2363-2780\3;</div> <div>charset cmyc_Pos3 = 2364-2781\3;</div> <div>charset Rag1_Pos1 = 2782-3415\3;</div> <div>charset Rag1_Pos2 = 2783-3416\3;</div> <div>charset Rag1_Pos3 = 2784-3417\3;</div>
Aligned DNA sequence matrix for phylogenetic analyses in the article "A new glassfrog of the genus Centrolene (Amphibia: Centrolenidae) from the Subandean Kutukú Cordillera, eastern Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "A new glassfrog of the genus Centrolene (Amphibia: Centrolenidae) from the Subandean Kutukú Cordillera, eastern Ecuador"</p> <p>The matrix is in NEXUS format and has 6626 bp and 239 terminals.</p> <p>Partitions are as follows:</p> <div> <div>charset 12S = 1-967;</div> <div>charset 16S = 968-2130;</div> <div> </div> <div>charset BNDFcodonPos1 = 2133-2829\3;</div> <div>charset BNDFcodonPos2 = 2131-2830\3;</div> <div>charset BNDFcodonPos3 = 2132-2828\3;</div> <div> </div> <div> </div> <div>charset ND1codonPos1 = 2832-3786\3;</div> <div>charset ND1codonPos2 = 2833-3787\3;</div> <div>charset ND1codonPos3 = 2831-3788\3;</div> <div> </div> <div> </div> <div>charset CXCR4codonPos1 = 3790-4144\3;</div> <div>charset CXCR4codonPos2 = 3791-4142\3;</div> <div>charset CXCR4codonPos3 = 3789-4143\3;</div> <div> </div> <div> </div> <div>charset cmyccodonPos1 = 4145-4547\3;</div> <div>charset cmyccodonPos2 = 4146-4548\3;</div> <div>charset cmyccodonPos3 = 4147-4549\3;</div> <div> </div> <div> </div> <div>charset POMCcodonPos1 = 4551-5160\3;</div> <div>charset POMCcodonPos2 = 4552-5161\3;</div> <div>charset POMCcodonPos3 = 4550-5162\3;</div> <div> </div> <div> </div> <div>charset RAG1codonPos1 = 5163-5616\3;</div> <div>charset RAG1codonPos2 = 5164-5617\3;</div> <div>charset RAG1codonPos3 = 5165-5618\3;</div> <div> </div> <div> </div> <div>charset SLC8A1codonPos1 = 5620-6160\3;</div> <div>charset SLC8A1codonPos2 = 5621-6158\3;</div> <div>charset SLC8A1codonPos3 = 5619-6159\3;</div> <div> </div> <div> </div> <div> </div> <div>charset SLC8A3codonPos1 = 6162-6627\3;</div> <div>charset SLC8A3codonPos2 = 6163-6625\3;</div> <div>charset SLC8A3codonPos3 = 6161-6626\3;</div> </div> <p> </p>
Alignments and ML trees of cassava brown streak virus and Ugandana cassava brown streak virus polyprotein nucleotide sequences
<p>Alignments of full and nearly full polyprotein-length nucleotide sequences from GenBank for the two ipomoviruses that cause cassava brown streak disease, in fasta format. Separate alignments for 67 cassava brown streak virus sequences and 81 Ugandan cassava brown streak virus sequences are provided, as well as a combined alignment of 148 sequences. Alignments were created with MUSCLE and then modified by eye in AliView.</p> <p>Also, two tree files (in nexus) format are supplied, resulting from a maximum likelihood analysis with IQTree on each of the two single-species datasets. Support for nodes with aLRT and 100 actual bootstrap replicates are provided (aLRT/bootstrap).</p>
Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method
<p>This dataset contains a GitHub repository containing all the data, analysis, Nextflow workflows and Jupyter notebooks to replicate the manuscript titled "Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method".</p> <p>It also contains the Multiple Sequence Alignments (MSAs) generated and well as the main figures and tables from the manuscript.</p> <p>The repository is also available at GitHub (https://github.com/cbcrg/dpa-analysis) release `v1.2`.</p> <p>For details on how to use the regressive alignment algorithm, see the T-Coffee software suite (https://github.com/cbcrg/tcoffee).</p>
Aligned Cox1 and Cob sequences from Oikopleura dioica and other tunicates.
<p>Supporting data for the manuscript « <em>Widespread use of the “ascidian” mitochondrial genetic code in tunicates</em> » containing a) assemblies of the cytochrome oxidase subunit 1 (Cox1) and cytochrome b (Cob) mitochondrial genes from ESTs downloaded from the Oikobase database and b) protein alignment of these sequences and other tunicate sequences to study which genetic code is used in tunicate mitochondria.</p>
Multiple Sequence Alignment of a diverse dataset with 1788 Mycobacterium tuberculosis isolates
<p><strong>Multiple Sequence Alignment of a diverse dataset with 1788 <em>Mycobacterium tuberculosis</em> isolates used for <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> benchmarking</strong></p> <p>The dataset comprises whole-genome sequence data published by <a href="https://doi.org/10.1016/S1473-3099(15)00062-6">Walker et al. 2015</a>. For the multiple sequence analysis, we proceeded as follows:</p> <ol> <li>Reads were downloaded from ENA BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA282721">PRJNA282721</a> (accessed on March 16<sup>th</sup>, 2023) and trimmed using Trimmomatic (<a href="https://pubmed.ncbi.nlm.nih.gov/24695404/">Bolger et al., 2014</a>) with <a href="https://github.com/B-UMMI/INNUca">INNUca</a> default settings;</li> <li>Quality-processed reads were individually mapped against the H37Rv reference genome (Genbank accession: <a href="https://www.ncbi.nlm.nih.gov/nuccore/NC_000962.3/">NC_000962.3</a>) using <a href="https://github.com/tseemann/snippy">Snippy</a> v4.5.1 and SNP-calling was performed on variant sites with the following criteria: a minimum proportion of reads differing from the reference of 70%, a minimum mapping quality of 30 and a minimum coverage for SNP calling of 10;</li> <li>A full alignment was extracted using Snippy’s core module (snippy-core), with masking of SNPs falling within known <em>M. tuberculosis</em> genomic regions with high GC content, repetitive elements and resistance-associated positions (corresponding to ~8% of the genome), as previously described for surveillance purposes (<a href="https://pubmed.ncbi.nlm.nih.gov/30948181/">Macedo et al., 2019</a>);</li> <li><em>M. tuberculosis </em>lineages were determined using tb-profiler v4.4.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31234910/">Phelan et al., 2019</a>), with samples from the <em>M. tuberculosis</em> complex other than <em>M. tuberculosis</em>, representing a mix of multiple lineages, or with less than 95% of mapped positions in the reference, being excluded;</li> <li>A filtered alignment comprising the maximum number of informative sites (88,562 nucleotide sites with at least one mutation in a given sequence) was extracted from the full alignment using the alignment_processing.py v1.1.0 (default settings) of <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a>, and then used as input for the benchmarking.</li> </ol> <p>In this repository, we provide two alignment files:</p> <ul> <li>Core_MTB_1787_strs.full.aln: this corresponds to the full multiple sequence alignment comprising 1787 samples and the reference (corresponding to the point 4 of the methodology).</li> <li>MTb_original_align_profile.fasta: this corresponds to the multiple sequence alignment comprising 1787 samples and the reference and only presenting the alignment informative sites (corresponding to the point 5 of the methodology)</li> </ul>
Multiple sequence alignments of sensor histidine kinases and response regulators
<p>The two FASTA files contain multiple sequence alignments of sensor histidine kinase and response regulator sequences. The source sequences were obtained by BLAST, clustered with usearch and aligned with muscle. More details to be found in Multamäki et al. 2021.</p>
PubMLST allele profiles and sequence alignments of 382 carbapenem-resistant Pseudomonas aeruginosa isolates collected from Japanese hospitals in 2019-2020
<p>This dataset provides PubMLST allele profiles, sequence alignments, and Mash distance data used in the study of "Nationwide genome surveillance of carbapenem-resistant Pseudomonas aeruginosa in Japan". </p>
Pterula and Pterulicium concatenated sequence alignment
Open the record for dataset details and reuse information.
Concatenated alignment of 12S and 16S sequences of the genus Pristimantis.
<p>A new species of Pristimantis (Amphibia, Anura, Strabomantidae) from a montane forest of the Pui Pui Protected Forest in central Peru</p>
Alignment files for coverage benchmarks: Illumina and Nanopore sequencing datasets
<ul> <li><strong>cpara-illumina-noseq.bam</strong> and <strong>cpara-ont-noseq.bam</strong>: BAM files produced aligning the raw reads produced respectively by Illumina NextSeq and ONT Nanopore sequencing of an isolate of <em>C. parapsilosis</em> to evaluate the coverage calculations using real datasets.*</li> <li><strong>HG00258.bam</strong>: Exome sequencing from the 1000 Genomes Project (Clarke et al 2016 <a href="https://doi.org/10.1093/nar/gkw829">https://doi.org/10.1093/nar/gkw829</a>).</li> <li><strong>panel_01.bam</strong>: targeted sequencing of a Human gene panel of 16 genes.*</li> </ul> <p>* Sequences and qualities have been removed</p>
Aligned gene sequences
<p>Aligned gene sequences used for the phylogenetic analysis</p>
DNA sequences alignements for 27 species of ticks (SCO50 matrix)
<p>The two files contain respectively the concatenation of DNA sequences alignements for 27 species of ticks (SCO50 matrix, n=952 genes) and to the partition file indicating the positions of each gene in the concatenation.</p> <p>The data set corresponds to the article "A transcriptome-based phylogenetic study of hard ticks (Ixodidae)" to be published in Scientific Reports, by N Pierre Charrier, Axelle Hermouet, Caroline Hervet, Albert Agoulon,<br> Stephen Barker, Dieter Heylen, Céline Toty, Karen McCoy, Olivier Plantard, Claude Rispe.</p>
Text-fig. 5. Multiple sequence alignment of mtDNA from ancient bone and recent greyhound, (Gundry et al. 2007) primer pair A – F15.719 and R16.114. in Genetic Analysis Of Possibly The Oldest Greyhound Remains Within The Territory Of The Czech Republic As Proof Of A Local Elite Presence At Chotěbuz-Podobora Hillfort In The 8 -9 Century Ad
Text-fig. 5. Multiple sequence alignment of mtDNA from ancient bone and recent greyhound, (Gundry et al. 2007) primer pair A – F15.719 and R16.114.
Fig. 5 in Total Evidence, Sequence Alignment, Evolution of Polychrotid Lizards, and a Reclassification of the Iguania (Squamata: Iguania)
Fig. 5. The same consensus as shown in figure 4, with stems and taxa identified for ready comparison with evolutionary changes presented in appendix 4 (change list). Stems 5, 6, 16, 20, 23, 24 represent alternative relationships not rejected by the data.
Fig. 3. Sensitivity analysis graphic. Y in Total Evidence, Sequence Alignment, Evolution of Polychrotid Lizards, and a Reclassification of the Iguania (Squamata: Iguania)
Fig. 3. Sensitivity analysis graphic. Yaxis represents the logarithm of the ratio of transversion: transition weights. The xaxis represents the logarithm of the ratio of the indel cost versus the maximal cost of a molecular change. The colors represent the zaxis, which is congruence between the molecular and morphological data partitions. Red is good, blue is bad.
Fig. 2 in Total Evidence, Sequence Alignment, Evolution of Polychrotid Lizards, and a Reclassification of the Iguania (Squamata: Iguania)
Fig. 2. Strict consensus of 192 equally parsimonious trees based on morphology alone (length = 276; CI = 0.443; RI = 0.74). Numbers on branches are Bremer values.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.