Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
175
datasets available to search
ShareScore release 0.9.0
Dataset results
175 results for “sequence alignments”
FIGURE 42. Maximum Likelihood consensus tree inferred from the 16S rDNA sequence alignment representing a in Monographic revision of the endemic Helix mazzullii De Cristofori & Jan, 1832 complex from Sicily and re-introduction of the genus Erctella Monterosato, 1894 (Pulmonata, Stylommatophora, Helicidae)
FIGURE 42. Maximum Likelihood consensus tree inferred from the 16S rDNA sequence alignment representing a possible reconstruction of Helicidae phylogeny. Initial trees for the heuristic search were obtained automatically. A GTR + Γ model (alpha= 0.29) was employed. The analysis involved 58 nucleotide sequences. All positions containing gaps and missing data were eliminated.
Aligned DNA sequence matrix for phylogenetic analyses in the article "Molecular and Morphological Assessment of Rain Frogs in the Pristimantis orestes Species Group (Amphibia: Anura: Strabomantidae) with the Description of Three New Cryptic Species from Southern Ecuador"
<p>The aligned matrix is in fasta format. Genes are arranged as follows:</p> <p>12S = 1–901</p> <p>16S = 902–2094</p> <p>RAG-1 = 2095–2733</p>
alignment of cathepsin and silicatein sequences
<p>alignment used to build the cathepsin/silicatein tree provided in supplementary figures</p>
Synthetic data for Aligning Distant Sequences to Graphs using Long Seed Sketches
<p>Each directory inside the folder correponds to the number of levels used to generate the dataset (for more details, see the description written in the publication). Inside each directory, the files "reference_X" and "mutated_X" correpond to the sequences reference and mutated at rate X, respectively. </p>
Response_reg and ABC_tran: monster benchmark families for multiple sequence alignments
<p>The data set contains two benchmark families: Response_reg (1.8 million sequences) and ABC_tran (3.5 million sequences). The sets were constructed by combining Homstrad reference alignments (November 2022) with Pfam 35 UniProt complete families (PF00072 and PF00005).</p>
Genome Alignment of Cancer Sequencing Data
<p>Part of the GTN Cancer Analysis learning Pathway based on the Bioinformatics.ca Cancer Workshop</p>
EPSAPG: A Pipeline Combining MMseqs2 and PSI-BLAST to Quickly Generate Extensive Protein Sequence Alignment Profiles
<p>This repository contains all data, queries, and search results used in the analysis of</p> <ul> <li>Arab, Issar. “<strong>EPSAPG: A Pipeline Combining MMseqs2 and PSI-BLAST to Quickly Generate</strong><br><strong>Extensive Protein Sequence Alignment Profiles</strong>.” IEEE/ACM 10th International Conference on<br>Big Data Computing, Applications and Technologies (BDCAT ’23), December 4–7, 2023,<br>Taormina, Messina, Italy. <a href="https://doi.org/10.1145/3632366.3632384">doi.org/10.1145/3632366.3632384</a></li> </ul>
Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.
<p>List of 200 most abundant VOG descriptions.</p>
Fig. 4. Amino acid sequences alignment between TCS1 and candidate N in Discovery and Biochemical Characterization of N-methyltransferase Genes Involved in Purine Alkaloid Biosynthetic Pathway of Camellia gymnogyna Hung T.Chang (Theaceae) from Dayao Mountain
Fig. 4. Amino acid sequences alignment between TCS1 and candidate N-methyltransferase genes (GCS1, GCS2, and GCS3).
Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data
<p>Source data for the paper "Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data"</p>
Sequence alignment for 7 gene regions for new Phytophthora species in clade 2a
Open the record for dataset details and reuse information.
Aligned and curated mtDNA sequences from: Ancient DNA reveals interstadials as a driver of common vole population dynamics during the last glacial period
Open the record for dataset details and reuse information.
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
Open the record for dataset details and reuse information.
Two new species of Aphyllon from northeastern Mexico: Sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Coregonus spp. opsin amplicon sequence alignments
Open the record for dataset details and reuse information.
Sequence alignments of Corallicolids, apicomplexan symbionts of coral
Open the record for dataset details and reuse information.
Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)
Open the record for dataset details and reuse information.
Multiple Sequence Alignments (MSA) for comparative phylogeography of four lizard taxa within an Oceanic Island
Open the record for dataset details and reuse information.
Data from: Aligner optimization increases accuracy and decreases compute times in multi-species sequence data
Open the record for dataset details and reuse information.
Aligned DNA sequences of Vanilla
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.