Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
175
datasets available to search
ShareScore release 0.9.0
Dataset results
175 results for “sequence alignments”
A Framework for High-throughput Sequence Alignment using Real Processing-in-Memory Systems
<p>Sequence alignment is a fundamentally memory bound computation whose performance in modern systems is limited by the memory bandwidth bottleneck. Processing-in-memory architectures alleviate this bottleneck by providing the memory with computing competencies. We propose Alignment-in-Memory (AIM), a framework for high-throughput sequence alignment using processing-in-memory, and evaluate it on UPMEM, the first publicly-available general-purpose programmable processing-in-memory system.</p>
Phylogenetic and recombination analysis of adenovirus isolates reveals discordance between serotype and phylogeny: Multiple sequence alignments
<p><strong>Background</strong></p> <p>Human adenovirus (HAdV) infections are caused by seven mastadenovirus species (A-G) and are the source for a variety of pathologies including gastrointestinal, respiratory, neurological, and ocular disease. While HAdV-D is the most common cause of adenovirus ocular infections, human adenoviruses B and E have also been isolated from the eye. </p> <p><strong>Results</strong></p> <p>In the course of classifying three new atypical ocular adenovirus samples, taken from the vitreous humor, we found that all three isolates were HAdV-B species, with isolate BP-AdV1 sorting with the B1 clade, and isolates BP-AdV2 and BP-AdV3 grouping into the B2 clade. The three Bascom Palmer HAdV-B genomes were then combined with over 300 HAdV-B genome sequences, including 9 ocular HAdV-B genome sequences. The whole genome phylogenetic analysis showed that 9 of the 11 ocular sequences grouped into the B1 clade, forming two clusters within B1. Attempts to categorize the penton, hexon and fiber serotypes using phylogeny of the three Bascom Palmer samples were inconclusive due to incongruence between serotype and phylogeny in the dataset. Recombination analysis using a subset of HAdV-B strains to generate a hybridization network detected recombination between non-human primate and human derived strains, recombination between one HAdV-B strain and the HAdV-E outgroup and limited recombination between the B1 and B2 clades. </p> <p><strong>Conclusions</strong></p> <p>The discordance between serotype and phylogeny detected in this study suggests that the current penton/hexon/fiber-based classification mechanism does not accurately describe the natural history and phylogenetic relationships amongst adenoviruses. A new adenovirus strain classification strategy may be beneficial to the field.</p>
Alignment of ITS sequences of Erysiphe quercicola and related Erysiphe species
<p>Internal transcribed spacer sequences (ITS) of the nuc rDNA are given by their scientific name, the specimen and GenBank accession numbers, the host name and country in an alignment composed by MUSCLE implemented in MEGA X.</p>
Aligned_sequences
<p>A csv file of aligned sequence for the pre-training of DNABERT-S. Used during the master thesis 'Transformer-based AI models as classification tool for DNA-metabarcoding' by JSchreijer. <span><br></span></p>
Multiple Sequence Alignments for Octopus bocki
<p><span>Multiple sequence alignments were created in MEGA11: Molecular Evolutionary Genetics Analysis version 11 (<em>Whelan and Goldman, 2001</em>)</span><span> using Muscle default parameters (UPGMA cluster method with -2.9 gap open and 0 gap extension penalties)</span></p> <p><em><span>Whelan, S. and Goldman, N. (2001). A general empirical model of protein evolution derived from multiple protein families using a maximum-likelihood approach. Molecular Biology and Evolution 18:691-699.</span></em></p>
Multiple sequence alignment of USP Zf-UBD proteins
<p>Using Molsoft's ICM-Pro, a multiple sequence alignment of USP Zf-UBDs was done against HDAC6 Zf-UBD. </p>
AFLP dataset and sequence alignments from a study of Rhododendron ferrugineum across the whole species range.
<p>STRUCTURE DATASET</p> <p>The raw data are in the files:</p> <ol> <li>Rfe_KAR_AAT-CAC2_2017-10-18-16-55-16.zip</li> <li>Rfe_KAR_ATC-CAC_2017-10-25-12-39-28.zip</li> <li>Rfe_KAR_ATG-CTG2_2017-10-24-09-40-25.zip</li> </ol> <p>Population data: populations.txt</p> <p>The final structure file: Rfe.str</p> <p>We analysed the obtained results in GeneMapper (Applied Biosystems) with default settings for AFLP analysis. The combinations of selective starters amplified the following numbers of polymorphic markers: AAT-CAC=127 (mean=40,4; SD=8,1), ATC-CAC=119 (mean=35,3; SD=5,4), ATG-CTG=144 (mean=44,4; SD=5,0). The blank samples yielded 0, 6 and 5 markers, respectively, which were removed from the final data matrix. We also removed the markers present in only one individual. Three samples (one from each of the populations L22, K27 and S05) failed to amplify in one of the selective primer reactions and were thus removed from the final matrix. The estimated error rate was 2.87% and the final matrix comprises of 87 individuals and 254 polymorphic loci.</p> <p>SEQUENCE ALIGNMENTS</p> <ol> <li>Rfe_ITS_align.fas</li> <li>Rfe_rpl32-trnL_align.fas </li> <li>Rfe_RPS12-RPL20_align.fas </li> <li>Rfe_trnfM-trnS_align.fas </li> <li>Rfe_trnLF_align.fas </li> </ol>
Aligned DNA sequence matrix for phylogenetic analyses of Scinax in the article "Advertisement calls and DNA sequences reveal a new species of Scinax (Anura: Hylidae) on the Pacific Lowlands of Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of Scinax in the article "Advertisement calls and DNA sequences reveal a new species of Scinax (Anura: Hylidae) on the Pacific Lowlands of Ecuador". The matrix is in NEXUS format.</p> <p>Gene partition as arranged as follows (tRNAs are included as part of larger adjacent genes):</p> <p>12S RNA: 1-1033</p> <p>16S RNA: 1034-2832</p> <p>NADH dehydrogenase subunit 1: 2833-3793</p> <p>Cytochrome Oxidase sub-unit I: 3965-4654</p> <p>Cytochrome B: 4655-5447</p> <p> </p>
Aligned DNA sequence matrix for phylogenetic analyses of the article "A bizarre new species of Lynchius (Amphibia, Anura, Strabomantidae) from the Andes of Ecuador and first report of Lynchius parkeri in Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "A bizarre new species of <em>Lynchius</em> (Amphibia, Anura, Strabomantidae) from the Andes of Ecuador and first report of <em>Lynchius parkeri</em> in Ecuador"</p> <p>The matrix is in NEXUS format. Genes are arranged as follows:</p> <p>RAG1: 1-652<br> Tyrosinase: 653-1195<br> 12S RNA: 1196-2242<br> tRNA Val: 2243-2313<br> 16S RNA: 2314-3994<br> tRNA Leu = 3995-4065<br> ND1: 4066-5026<br> tRNA Ile: 5027-5144;</p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Systematics of Huicundomantis, a new subgenus of Pristimantis (Anura, Strabomantidae) with extraordinary cryptic diversity and eleven new species"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "Systematics of Huicundomantis, a new subgenus of Pristimantis (Anura, Strabomantidae) with extraordinary cryptic diversity and eleven new species"</p> <p>The matrix is in NEXUS format. Genes are arranged as follows:</p> <p>16S RNA, tRNA-Leu = 1–1408<br> ND1 codon position 1 = 1409–2369\3;<br> ND1 codon position 2 = 1410–2367\3;<br> ND1 codon position 3 = 1411–2368\3;<br> tRNA-Ile, tRNA-Gln, rRNA-Met = 2370–2544; <br> RAG1 codon position 3 = 2545–3199\3;<br> RAG1 codon position 1 = 2546–3197\3<br> RAG1 codon position 2 = 2547–3198\3;</p>
Dataset: EF1-α sequence alignement of Lecanostica acicola
<p>Sequence alignement of elongation factor 1-α (EF1-α) gene. The sequence dataset was aligned using the MUSCLE algorithm in MEGA 6.0 (Tamura <em>et al.</em>, 2013).</p>
Towards a rapid sequencing-based molecular surveillance and mosaicism investigation of Toxoplasma gondii (nucleotide alignment dataset)
<p>This dataset includes the nucleotide alignment of eight Toxoplasma gondii genome loci (Sag1 / Chromossome VIII, Gra6 / Chromossome X, PK1 / Chromossome VI, Sag3 / Chromossome XII, L363 / Chromossome VIIb, CB21-4 / Chromossome III, M102 / Chromossome VIIa, Sag2 / Chromossome VIII). Each alignment includes sequences from T. gondii reference strains (retrieved from ToxoDB) as well as sequences from multiple clinical strains (obtained by Sanger / Next-generation sequencing) of the collection of the National Reference Laboratory of Parasitic and Fungal Infections, Department of Infectious Diseases, National Institute of Health Dr. Ricardo Jorge, Portugal. </p>
Aligned DNA sequence matrixes for the study of the divergent times of phytoplasmas
<p>Sequence alignments of 16S rRNA and methionine aminopeptidase (map) are provided in FASTA files “Cao_et_al_16S.fas” and “Cao_et_al_map.fas”, respectively. Detailed information of the data matrixes is as follows:</p> <p> </p> <p>File name: Cao_et_al_16S.fas</p> <p>Number of taxa: 220</p> <p>Number of characters: 1655</p> <p>Gap: -</p> <p> </p> <p>File name: Cao_et_al_map.fas</p> <p>Number of taxa: 83</p> <p>Number of characters: 564</p> <p>Gap: -</p>
Aligned DNA sequence matrixes for the study of the divergent times of phytoplasmas
<p>Sequence alignments of 16S rRNA and methionine aminopeptidase (map) are provided in FASTA files “Cao_et_al_16S.fas” and “Cao_et_al_map.fas”, respectively. Detailed information of the data matrixes is as follows:</p> <p> </p> <p>File name: Cao_et_al_16S.fas</p> <p>Number of taxa: 220</p> <p>Number of characters: 1655</p> <p>Gap: -</p> <p> </p> <p>File name: Cao_et_al_map.fas</p> <p>Number of taxa: 83</p> <p>Number of characters: 564</p> <p>Gap: -</p>
Data file with manuscript titled 'A Structurally Validated Sequence Alignment of 497 Human Protein Kinase Domains'
<p>The files used in different analysis reported in the manuscript titled - 'A Structurally-Validated Multiple Sequence Alignment of 497 Human Protein Kinase Domains' are shared at two locations. Following is a brief description of these files.</p> <p>Location - https://github.com/DunbrackLab/Kinases<br> 1. HMM profile files - HMM files for each of the nine groups computed separately labeled as Groupname.hmm, like AGC.hmm<br> 2. HMM profile file - HMM file computed from the full alignment including all the sequences - Human-PK.hmm<br> 3. Score files - HMM scores of each kinase sequence against all the groupwise HMMs both for iteration1 (HMM-iter1-scores-tables.txt) and iteration2 (HMM-iter1-scores-tables.txt)<br> 4. Jalview session file - Kinase alignment with sequences colored by secondary structure information from PDB file if the structure is known; or predicted secondary structure if the experimental structure is not known. The file could be opened in Jalview - kinases-PDB-SSPred.jvp</p> <p>Location - https://zenodo.org/record/3445533<br> 1. The file contains list of residue pairs aligned in pairwise structural alignments of 272 human protein kinases which were used as a benchmark in the study. The alignments were created by FATCAT and optimized by SE program.</p>
Human sequence alignment data set used for analysis of SPDI algorithm and tools
<p>Collection of alignment segments produced on October 30, 2019. The ADS currently consists of over 2,680,000 pairwise alignment segments generated from over 350,000 distinct input sequences. </p> <ul> <li> <p>Old assembly to current Genome Reference Consortium (GRC) <a href="http://f1000.com/work/citation?ids=111899&pre=&suf=&sa=0">(Church et al., 2011)</a> primary assemblies (e.g. GRCh36(hg18) or GRCh37(hg19) with GRCh38(hg38))</p> </li> </ul> <ul> <li> <p>Patches, alternative loci, or pseudoautosomal regions (PAR) to GRC primary assembly</p> </li> <li> <p>RefSeq <a href="http://f1000.com/work/citation?ids=2599029&pre=&suf=&sa=0">(O’Leary et al., 2016)</a> and select GenBank <a href="http://f1000.com/work/citation?ids=6183037&pre=&suf=&sa=0">(Benson et al., 2018)</a> transcripts to selected RefSeq genomic regions, also known as RefSeqGene (NG), a member of the Locus Reference Genome (LRG) collaboration <a href="http://f1000.com/work/citation?ids=3225699&pre=&suf=&sa=0">(Dalgleish et al., 2010)</a>.</p> </li> <li> <p>Current RefSeq transcripts (NM/NR/XM/XR) and RefSeq genomic regions (NG) to the latest Assembly</p> </li> <li> <p>Previous versions of NG and RefSeq transcripts (NM/NR) to GRC primary assembly</p> </li> </ul>
Aligned DNA sequence matrix for phylogenetic analyses in the article "A new glassfrog of the genus Nymphargus (Anura: Centrolenidae) from Cordillera del Cóndor, Ecuador"
<p>Aligned DNA sequence matrix for phylogenetic analyses of the article "<span>A new glassfrog of the genus Nymphargus</span> (Anura<span>:</span> Centrolenidae) <span>from Cordillera del Cóndor, Ecuador</span>"</p> <p>The matrix is in NEXUS format and has 6613 bp and 102 terminals.</p> <p>Partitions are as follows:</p> <div> <div>charset 12S = 1-944;</div> <div>charset 16S = 945-2093;</div> <div>charset BNDFcodonPos1 = 2096-2792\3;</div> <div>charset BNDFcodonPos2 = 2094-2793\3;</div> <div>charset BNDFcodonPos3 = 2095-2791\3;</div> <div>charset ND1codonPos1 = 2795-3749\3;</div> <div>charset ND1codonPos2 = 2796-3750\3;</div> <div>charset ND1codonPos3 = 2794-3751\3;</div> <div>charset CXCR4codonPos1 = 3753-4107\3;</div> <div>charset CXCR4codonPos2 = 3754-4105\3;</div> <div>charset CXCR4codonPos3 = 3752-4106\3;</div> <div>charset cmyccodonPos1 = 4108-4507\3;</div> <div>charset cmyccodonPos2 = 4109-4508\3;</div> <div>charset cmyccodonPos3 = 4110-4509\3;</div> <div>charset POMCcodonPos1 = 4511-5120\3;</div> <div>charset POMCcodonPos2 = 4512-5121\3;</div> <div>charset POMCcodonPos3 = 4510-5122\3;</div> <div>charset RAG1codonPos1 = 5123-5576\3;</div> <div>charset RAG1codonPos2 = 5124-5577\3;</div> <div>charset RAG1codonPos3 = 5125-5578\3;</div> <div>charset SLC8A1codonPos1 = 5580-6120\3;</div> <div>charset SLC8A1codonPos2 = 5581-6118\3;</div> <div>charset SLC8A1codonPos3 = 5579-6119\3;</div> <div>charset SLC8A3codonPos1 = 6122-6587\3;</div> <div>charset SLC8A3codonPos2 = 6123-6585\3;</div> <div>charset SLC8A3codonPos3 = 6121-6586\3;</div> </div>
DNA sequences of transgenes detected via environmental DNA (raw ABI files, processed FASTA files, and reference alignments)
We demonstrate that simple, non-invasive environmental DNA (eDNA) methods can detect transgenes of genetically modified (GM) animals from terrestrial and aquatic sources in invertebrate and vertebrate systems. We detected transgenic fragments between 82-234 bp through targeted PCR amplification of environmental DNA extracted from food media of GM fruit flies (<i>Drosophila melanogaster</i>), feces, urine, and saliva of GM laboratory mice (<i>Mus musculus</i>), and aquarium water of GM tetra fish (<i>Gymnocorymbus ternetzi</i>). With rapidly growing accessibility of genome-editing technologies such as CRISPR, the prevalence and diversity of GM animals will increase dramatically. GM animals have already been released into the wild with more releases planned in the future. eDNA methods have the potential to address the critical need for sensitive, accurate, and cost-effective detection and monitoring of GM animals and their transgenes in nature.
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
<p>The profiling of epigenetic marks like DNA methylation has become a central aspect of studies in evolution and ecology. Bisulfite sequencing is commonly used for assessing genome-wide DNA methylation at single nucleotide resolution but these data can also provide information on genetic variants like single nucleotide polymorphisms (SNPs). However, bisulfite conversion causes unmethylated cytosines to appear as thymines, complicating the alignment and subsequent SNP calling. Several tools have been developed to overcome this challenge, but there is no independent evaluation of such tools for non-model species, which often lack genomic references. Here, we used whole-genome bisulfite sequencing (WGBS) data from four female great tits (<i>Parus major</i>) to evaluate the performance of seven tools for SNP calling from bisulfite sequencing data. We used SNPs from whole-genome resequencing data of the same samples as baseline SNPs to assess common performance metrics like sensitivity, precision, and the number of true positive, false positive, and false negative SNPs for the full range of variant and genotype quality values. We found clear differences between the tools in either optimizing precision (Bis-SNP), sensitivity (biscuit), or a compromise between both (all other tools). Overall, the choice of SNP caller strongly depends on which performance parameter should be maximized and whether ascertainment bias should be minimized to optimize downstream analysis, highlighting the need for studies that assess such differences.</p>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Our unknown neighbor: a new species of rain frog of the genus Pristimantis (Amphibia: Anura: Strabomantidae) from the city of Loja, southern Ecuador"
<p>The aligned matrix is in fasta format. Genes are arranged as follows:</p> <p>12S = 1–909</p> <p>16S = 910–1787</p> <p>RAG-1 = 1788–2399</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.