Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,199
datasets available to search
ShareScore release 0.9.0
Dataset results
1,199 results for “Alignement”
45 fish genome alignment
<p>A visual representation of an all-to-all alignment of 45 fish genomes. Generated in two days on a single 48-core compute node using wfmash v0.7.0+2bbd135. Plot generated with pafplot -s 20000. Represented genomes:</p> <p>fArcCen1<br> fNotCel1.pri<br> sPriPec2.1.pri<br> fCycLum1.pri<br> fHipHip1.pri<br> fPerMag1.pri<br> fEsoLuc1.pri<br> fAngAng1.pri<br> fAntMac1.pri<br> fEleEle1.pri<br> fMegCyp1.pri<br> fAnaAna1.pri<br> fXenCan1.pri<br> fPygNat1.pri<br> fSebUmb1.pri<br> sCarCar2.pri<br> fAplTae1.pri<br> fMelBoe1.pri<br> fCheRos1.pri<br> fToxJac2.pri<br> fAstCal1.3<br> fAnaTes1.3<br> fMasArm1.3<br> fCotGob3.1<br> fParRan2.2<br> fGouWil2.2<br> fBetSpl5.3<br> fDenClu1.2<br> fErpCal1.2<br> fSpaAur1.2<br> fEcheNa1.2<br> fSclFor1.1<br> fTakRub1.3<br> fSalTru1.2<br> fSynAcu1.2<br> fSalaFa1.1<br> fSphaOr1.1<br> fMyrMur1.1<br> gadMor3.0<br> fChaCha1.1<br> fThaAma1.1<br> fDreTuH1.2<br> fDreABH1.1<br> fAcaLat1.1<br> fTraTra1.1<br> </p>
Supporting trees and alignments for the publication: Cryptic and abundant marine viruses at the evolutionary origins of Earth's RNA virome
<div class="page"> <div class="layoutArea"> <div class="column"> <p>Whereas DNA viruses are known to be abundant, diverse, and commonly key ecosystem players, RNA viruses are relatively understudied outside disease settings. Here, we analyzed ≈28 terabases of Global Ocean RNA sequences to expand Earth's RNA virus catalogues and their taxonomy, investigate their evolutionary origins, and assess their marine biogeography from pole to pole. Using new approaches to optimize discovery and classification, we identified RNA viruses that necessitate substantive revisions of taxonomy (doubling phyla and adding >50% new classes) and evolutionary understanding. "Species"-rank abundance determination revealed that viruses of new phyla<span> </span><em>"Taraviricota"</em><span>, </span>a missing link in early RNA virus evolution, and<span> </span><em>"Arctiviricota"</em><span> </span>are widespread and dominant in the oceans. These efforts provide foundational knowledge critical to integrating RNA viruses into ecological and epidemiological models.</p> </div> </div> </div>
Aligned DNA sequence matrix for phylogenetic analyses in the article "Dos nuevas especies del grupo Pristimantis boulengeri (Anura: Strabomantidae) de la cuenca alta del río Napo, Ecuador" by Bejarano, et al.
<p>Matrix in nexus format that include sequences of 16S (1-1295), RAG1 (1296-1922), 12S (1923-3306), and COI (3307-3984) for 97 specimens belonging to the genus <em>Pristimantis</em>, in addition to specimens of <em>Strabomantis</em> and <em>Niceforonia</em> as outgroups.</p>
Alignments used for the phylogenetic analysis of sweet cucumber (Solanum muricatum, Solanaceae)
<p><span><span>Sweet cucumber <em>(Solanum muricatum) </em>sect. <em>Basarthrum, </em>is a neglected native crop to the Andean region. It is naturally distributed very close to potatoes (<em>Solanum </em>sect. <em>Petota</em>) and tomatoes (<em>Solanum</em> sect. <em>Lycopersicon</em>), two groups with great economic importance. We obtained the first complete chloroplast (cp) genome of sweet cucumber and compared with seven Solanacea species. Pair-end clean reads were obtained by PE 150 library and the Illumina HiSeq 2500 platform. The complete cp genome of <em>S. muricatum</em> had a 155,681 bp with typical quadripartite structure, containing a large single copy (LSC) region (86,182 bp) and a small single-copy (SSC) region (18,360 bp), separated by two inverted repeat (IR) regions (25,568 bp). The annotation of chloroplast genome predicted 88 protein-coding genes (CDS), 8 ribosomal RNA (rRNA) genes, 37 transfer RNA (tRNA) genes, and one pseudogene. A total of 48 perfect microsatellites were identified, divided in mononucleotide repeats (32), followed by tetranucleotide (6) and dinucleotides (5). SSRs with trinucleotides repeats (3), pentanucleotide (1) and hexanucleotide (1) repeats motifs in these genomes were identified in lower quantity. Most of these repeats were distributed in the noncoding regions. Whole chloroplast genome comparison with the other seven Solanaceae species revealed that the small single copy and large single copy regions showed more divergence than inverted regions. Finally, phylogenetic analyses resolved that <em>S. muricatum</em> is a sister species to members of sections <em>Petota</em> + <em>Lycopersicum</em> + <em>Etuberosum</em>. This study reports for the first time the genome organization, gene content, and structural features of the chloroplast genome of <em>S. muricatum</em>. Also, this study </span><span>may provide the basis for evaluating genetic diversity within the Solanaceae species, and will be useful to examine the evolutionary processes in sweet cucumber landraces.</span></span></p>
Specimen alignment with limited point-based homology: 3D morphometrics of disparate bivalve shells (Mollusca: Bivalvia)
<p>Supplemental data and code for Edie, Collins, and Jablonski, Specimen alignment with limited point-based homology: 3D morphometrics of disparate bivalve shells (Mollusca: Bivalvia).</p>
Comparison of Cosine, Modified Cosine, and Neutral Loss Based Spectral Alignment For Discovery of Structurally Related Molecules
<p>Spectral libraries and analysis results of the evaluation between the cosine similarity, modified cosine similarity, and neutral loss matching for the discovery of structurally related molecules.</p> <p>Spectral libraries used as input:<br> - MassIVE-KB peptide spectral library (version 2018/06/15): LIBRARY_CREATION_AUGMENT_LIBRARY_TEST-82c0124b-download_filtered_mgf_library-main.mgf<br> - GNPS community spectral libraries (downloaded on 2022/05/12): ALL_GNPS_NO_PROPOGATED.mgf<br> - GNPS bile acids spectral library (downloaded on 2022/05/12): BILELIB19.mgf</p> <p>Analysis output results:<br> - massivekb_peptide_mods.csv: 955,228 peptide MS/MS spectrum pairs from MassIVE-KB<br> - gnps_libraries.csv: 10 million small molecule MS/MS spectrum pairs from the GNPS community spectral libraries<br> - gnps_libraries_metadata.csv: structural (InChI, SMILES) and class information (computed using Classyfire) for 58,165 small molecule spectra from the GNPS community spectral libraries<br> - gnps_bilelib.csv: 340,637 bile acids MS/MS spectrum pairs from the GNPS bile acids spectral library</p> <p>For more information, see: https://github.com/bittremieux/cosine_neutral_loss/</p>
Olfactory receptor alignments for: Ecological constraints on highly evolvable olfactory receptor genes and morphology in neotropical bats
<p>While evolvability of genes and traits may promote specialization during species diversification, how ecology subsequently restricts such variation remains unclear. Chemosensation requires animals to decipher a complex chemical background to locate fitness-related resources, and thus the underlying genomic architecture and morphology must cope with constant exposure to a changing odorant landscape; detecting adaptation amidst extensive chemosensory diversity is an open challenge. In phyllostomid bats, an ecologically diverse clade that evolved plant-visiting from an insectivorous ancestor, the evolution of novel food detection mechanisms is suggested to be a key innovation, as plant-visiting species rely strongly on olfaction, supplementarily using echolocation. If this is true, exceptional variation in underlying olfactory genes and phenotypes may have preceded dietary diversification. We compared olfactory receptor (OR) genes sequenced from olfactory epithelium transcriptomes and olfactory epithelium surface area of bats with differing diets. Surprisingly, although OR evolution rates were quite variable and generally high, they are largely independent of diet. Olfactory epithelial surface area, however, is relatively larger in plant-visiting bats and there is an inverse relationship between OR evolution rates and surface area. Relatively larger surface areas suggest greater reliance on olfactory detection and stronger constraint on maintaining an already diverse OR repertoire. Instead of the typical case in which specialization and elaboration are coupled with rapid diversification of associated genes, here the relevant genes are already evolving so quickly that increased reliance on smell has led to stabilizing selection, presumably to maintain the ability to consistently discriminate among specific odorants — a potential ecological constraint on sensory evolution.</p>
A Framework for High-throughput Sequence Alignment using Real Processing-in-Memory Systems
<p>Sequence alignment is a fundamentally memory bound computation whose performance in modern systems is limited by the memory bandwidth bottleneck. Processing-in-memory architectures alleviate this bottleneck by providing the memory with computing competencies. We propose Alignment-in-Memory (AIM), a framework for high-throughput sequence alignment using processing-in-memory, and evaluate it on UPMEM, the first publicly-available general-purpose programmable processing-in-memory system.</p>
Raw data for "Complete Alignment of a KB-Mirror System guided by Ptychography"
<p>All raw datasets for each figure presented in the article.</p>
Dataset for "Quantum trapping and rotational self-alignment in triangular Casimir microcavities"
<p><strong># Data and plotting code for "Quantum trapping and rotational self-alignment in triangular Casimir microcavities"</strong></p> <p><br><strong>## Contents</strong></p> <ul> <li><code>calculations</code>: results of scuff-em calculations of triangle dimers for different sizes of edge lengths</li> <li><code>experimental</code>: raw data used to plot images in the main text and SI</li> </ul> <p><strong>## Description of the data</strong></p> <p>The data stored in the <code>calculations</code> directory contain the following files for and from scuff-em calculations:</p> <ul> <li><code>calculations/final-data-and-plotting</code>: binary .mat files with all relevant data calculated for the publication using scuff-em and associated .m plotting scripts to visualize the data as presented in the main text and SI.</li> <li><code>calculations/final-data-and-plotting/tr2-<>.mat</code>: data files for plotting Figure 3 in the main text.</li> <li><code>calculations/final-data-and-plotting/plot_triangles_for_different_sizes.m</code>: script to plot above data</li> <li><code>calculations/final-data-and-plotting/fld-v5-tri-L4000nm-gap-0.12-v3.mat</code>: data file to plot Figure S8 and elements of Figure 3.</li> <li><code>calculations/final-data-and-plotting/plot_for_L4000.m</code>: script to plot Figure S8 and elements of Figure 3.</li> <li><code>calculations/input-data-for-calculations</code>: input data to perform scuff-em calculations to obtain above data</li> <li><code>calculations/input-data-for-calculations/tri-*</code>: exemplary input file to calculate triangle dimer with edge length L = 4000 nm, gap of 200 nm, relative shift of top triangle to (x,y)=(0,0), i.e. no shift.</li> <li><code>calculations/input-data-for-calculations/obj-*</code>: exemplary input files to calculate triangle dimers with relative rotation arouns the z-axis for different edge lengths.</li> <li><code>calculations/input-data-for-calculations/mesh-files</code>: mesh files to perform calculations</li> <li><code>experimental</code>: raw and processed data use to plot figures in the text and SI</li> <li><code>experimental/data-for-main-text-plots/Figure-<>.xlsx</code>: excel files with data used in figure in the main text; the ending denotes the figure number and panel within it</li> <li><code>experimental/data-for-SI-figures-spectra/Figure-S<>.xlsx</code>: excel files with data used in the SI figures; the data contain the raw measured spectra and TMM fits</li> <li><code>experimental/SI-Videos/SI-Video<></code>: folders containing videos (<code>figure<>.avi</code>) used to obtain the triangle overlap and angular rotation during diffusion and the data obtain from analysing these videos (<code>data.txt</code>), where the <code><></code> refers to the figure number in text or SI; Videos 1-3 were analysed for surface overlap and the data files contain overlap-vs-time; Videos 4-5 were analysed for angular rotation and the data files contain angle-vs-time</li> </ul> <p><strong>## Reproduction of the data</strong></p> <p>The calculations were performed using scuff-em, see <a href="https://homerreid.github.io/scuff-em-documentation/" target="_blank" rel="noopener">documentation</a> and <a href="https://github.com/homerreid/scuff-em/" target="_blank" rel="noopener">code</a> for practical details.</p> <p>The included input data in the *<code>calculations/input-data-for-calculations</code>* contain all used mesh files and input parameters to run the calculations. The exemplary input files should be modified to change the relative position and alignment of the dimer elements. Note, that for the full-shift-and-rotation calculations we used the *tri* mesh file and associated input file. For the non-shifted calculations for different edge lengths we used the *obj* input and mesh files.</p> <p>Analysis of the videos. Every video is divided into frames using ImageJ. Every image is analysed with a MATLAB code which was written to analyse the angle between two triangle flakes by using contrast in brightness of the edges compared to the background. The extracted data are plotted in the main text figures and SI videos are created by using separate MATLAB code to merged data plotting with the real time video.</p>
Phylogenetic and recombination analysis of adenovirus isolates reveals discordance between serotype and phylogeny: Multiple sequence alignments
<p><strong>Background</strong></p> <p>Human adenovirus (HAdV) infections are caused by seven mastadenovirus species (A-G) and are the source for a variety of pathologies including gastrointestinal, respiratory, neurological, and ocular disease. While HAdV-D is the most common cause of adenovirus ocular infections, human adenoviruses B and E have also been isolated from the eye. </p> <p><strong>Results</strong></p> <p>In the course of classifying three new atypical ocular adenovirus samples, taken from the vitreous humor, we found that all three isolates were HAdV-B species, with isolate BP-AdV1 sorting with the B1 clade, and isolates BP-AdV2 and BP-AdV3 grouping into the B2 clade. The three Bascom Palmer HAdV-B genomes were then combined with over 300 HAdV-B genome sequences, including 9 ocular HAdV-B genome sequences. The whole genome phylogenetic analysis showed that 9 of the 11 ocular sequences grouped into the B1 clade, forming two clusters within B1. Attempts to categorize the penton, hexon and fiber serotypes using phylogeny of the three Bascom Palmer samples were inconclusive due to incongruence between serotype and phylogeny in the dataset. Recombination analysis using a subset of HAdV-B strains to generate a hybridization network detected recombination between non-human primate and human derived strains, recombination between one HAdV-B strain and the HAdV-E outgroup and limited recombination between the B1 and B2 clades. </p> <p><strong>Conclusions</strong></p> <p>The discordance between serotype and phylogeny detected in this study suggests that the current penton/hexon/fiber-based classification mechanism does not accurately describe the natural history and phylogenetic relationships amongst adenoviruses. A new adenovirus strain classification strategy may be beneficial to the field.</p>
Alignment of ITS sequences of Erysiphe quercicola and related Erysiphe species
<p>Internal transcribed spacer sequences (ITS) of the nuc rDNA are given by their scientific name, the specimen and GenBank accession numbers, the host name and country in an alignment composed by MUSCLE implemented in MEGA X.</p>
360 Recording of downtown ULM Germany, with time aligned gps data
Open the record for dataset details and reuse information.
Aligned_sequences
<p>A csv file of aligned sequence for the pre-training of DNABERT-S. Used during the master thesis 'Transformer-based AI models as classification tool for DNA-metabarcoding' by JSchreijer. <span><br></span></p>
Corallium rubrum genome and transcripts alignment
<p><em>Corallium rubrum</em>, the precious red coral, is an octocoral endemic to the western Mediterranean Sea. It plays a key role in the well-known Mediterranean coralligenous ecosystem, a biodiversity hotspot. We sequenced its genome for further investigations into its biology (skeleton formation and color, genetic diversity), ecology, and evolutionary history.</p>
The alignment of 163 plastome haplotypes of East Asian Cerris oaks and 29 plastomes of related oak species
<p>This dataset includes the alignment of 163 plastome haplotypes of East Asian Cerris oaks and 29 plastomes of related oak species. The alignment was generated through four steps: (1) We used PhyloSuite v.1.2.2 to extract protein-coding genes (PCGs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) from the plastomes of 761 East Asian Cerris oak trees and 29 accessions of 22 related oak species. (2) The extracted regions were aligned individually with MAFFT v.7.313 and adjusted manually using BioEdit v.7.2.5. Specifically, inversions and length variations in simple sequence repeats were excluded because of their tendency for homoplasy. Eight ambiguously aligned regions in rps16-trnQ, psbM-trnD, ndhF-rpl32, rpl32-trnL, ndhD-psaC, and psaC-ndhE IGSs, and ndhF and ycf1 PCGs were also discarded to reduce phylogenetic noise. (3) The individual alignments were concatenated according to their respective positions in the plastome to generate a whole-plastome alignment with only one inverted repeat (IR) retained. (4) Unique plastome haplotypes were determined by DnaSP v.5.10.01.</p>
Multiple Sequence Alignments for Octopus bocki
<p><span>Multiple sequence alignments were created in MEGA11: Molecular Evolutionary Genetics Analysis version 11 (<em>Whelan and Goldman, 2001</em>)</span><span> using Muscle default parameters (UPGMA cluster method with -2.9 gap open and 0 gap extension penalties)</span></p> <p><em><span>Whelan, S. and Goldman, N. (2001). A general empirical model of protein evolution derived from multiple protein families using a maximum-likelihood approach. Molecular Biology and Evolution 18:691-699.</span></em></p>
concatenated_alignment_4378bases Description of three widespread new Peronospora species parasitising Caryophyllales
<p>concatenated_alignment_4378bases in Description of three widespread new Peronospora species parasitising Caryophyllales</p>
Multiple sequence alignment of USP Zf-UBD proteins
<p>Using Molsoft's ICM-Pro, a multiple sequence alignment of USP Zf-UBDs was done against HDAC6 Zf-UBD. </p>
Salivary levels of cariogenic bacterial species during orthodontic treatment with thermoplastic aligners or fixed appliances: a prospective cohort study
<p>Dataset for the analyses of the paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.