Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,199
datasets available to search
ShareScore release 0.9.0
Dataset results
1,199 results for “aligners”
Multiple sequence alignments of full-length L1 elements with evidence of retrotransposition activity.
<p>DNA sequences for full-length L1 elements showing evidence of retrotransposition activity were aligned using MUSCLE v3.8 with default number of iterations. Manual inspection of the multiple alignments was performed with Jalview v2.11 in order to remove upstream and downstream spurious sequences.</p> <ul> <li><em>selected_active_L1sOK_aligned_trimmed.fa</em> file includes aligned DNA sequences for 86 full-length L1Hs with medium-high activity and for 1 active L1Pt (outgroup).</li> <li><em> all_active_L1sOK_aligned_trimmed.fa</em> file includes aligned DNA sequences for 143 active L1Hs and for 1 active L1Pt (outgroup).</li> </ul>
Multilingual Knowledge Graph Completion With Joint Relation and Entity Alignment
<p>Code and data accompanying AlignKGC.</p>
Improving phonetic alignment by handling secondary sequence structures
<p>Supplementary material accompanying the paper "Improving phonetic alignment by handling secondary sequence structures".</p> <p>The data consists of 5 files:</p> <ul> <li>gold_standard.psa : the gold standard used in the analysis in PSA format</li> <li> sca-secondary.psa : the output of the algorithm with the secondary extension</li> <li>sca-traditional.psa : the output of the traditional algorithm</li> <li>sca-secondary-diff.psa : the differences of the secondary extension compared to the GS</li> <li>sca-traditional-diff.psa : the differences of the traditional algorithm compared to the GS</li> </ul> <p>For a description of the file-format used in this dataset, please refer to the LingPy tutorial under http://lingpy.org.</p>
Neutralization Data and Aligned ENV Sequences for Predicting Antibody Affinities using Artificial Neural Networks
<p>Sample file with neutralization data (IC<sub>50</sub>) for different antibodies and viral strains, adapted from J. Huang, G. Ofek, L. Laub, M. K. Louder, N. A. Doria-Rose, N. S. Longo, H. Imamichi, R. T. Bailer, B. Chakrabarti, S. K. Sharma, S. M. Alam, T. Wang, Y. Yang, B. Zhang, S. A. Migueles, R. Wyatt, B. F. Haynes, P. D. Kwong, J. R. Mascola, and M. Connors, “Broad and potent neutralization of HIV-1 by a gp41-specific human antibody.,” <em>Nature</em>, vol. 491, no. 7424, pp. 406–12, Nov. 2012.</p> <p> </p> <p>Aligned ENV sequences downloaded from the HIV Sequence Database (www.hiv.lanl.gov/content/sequence/HIV/mainpage.html). There are 4907 sequences and the alignment length is 1369.</p>
Comparison of Friction Properties Among Diverse Conventional-Ligating Lingual Bracket Systems According To Tooth Displacement During Leveling And Alignment: An In Vitro Mechanical Study
<p>The purpose of this study was to evaluate the effects of tooth displacement on frictional force when conventional-ligating lingual brackets (CL-LB), CL-LBs with narrow bracket width and customized CL-LBs were used with leveling/alignment wire. </p> <p>CL-LBs (7<sup>th</sup>Generation), CL-LBs with narrow bracket width (STb) and customized CL-LBs (Incognito) were tested under three conditions of tooth displacement [no displacement (control); 1mm palatal displacement (PD) of the maxillary right lateral incisor (MXLI); and 1mm gingival displacement (GD) of the maxillary right canine (MXC)](9 groups, <em>n</em>=6 per group). Static (SFF) and kinetic frictional forces (KFF) were measured in a stereolithographic typodont system and artificial saliva while drawing a 0.016-inch copper or super-elastic nickel-titanium archwire at a speed of 0.5 mm/min for 5 minutes at 36.5°C.</p>
Results of the paper "Composition Identification in Ottoman-Turkish Makam Music Using Transposition-Invariant Partial Audio-Score Alignment"
<p>This repository contains the composition identification and tonic identification results along with the statistical significance values presented in the paper:</p> <p>Şentürk, S., & Serra X. (2016). <strong>Composition Identification in Ottoman-Turkish Makam Music Using Transposition-Invariant Partial Audio-Score Alignment.</strong> In Proceedings of 13th Sound and Music Computing Conference (SMC 2016), (pp. XX–XX)., Hamburg, Germany.</p> <p>Please cite the publication above in any work using this dataset.</p> <p>For the details of the results, please refer to the paper. For any further information please contact the authors.</p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.</p>
Data alignment and phylogenetic trees from Phylogeny, ecology, morphological evolution, and reclassification of the diatom orders Surirellales and Rhopalodiales
<p>Alignment and tree files from Ruck et al. 2016:</p> <p>Phylogeny, ecology, morphological evolution, and reclassification of the diatom orders Surirellales and Rhopalodiales</p>
Aligned and trimmed 16S and COI DNA sequences of Oceaniidae (Hydrozoa)
<p>Aligned and trimmed 16S and COI sequences of Oceaniidae (Hydrozoa) used for the study "The polyps of <em>Oceania armata</em> identified by DNA barcoding (Cnidaria, Hydrozoa)"</p> <p>Format is Fasta, files are text files</p>
Flash propagation and inferred charge structure relative to radar-observed ice alignment signatures in a small Florida Mesoscale Convective System
<p>Data for paper of above title, submitted to <em>Geophysical Research Letters</em>, June 2017. Manuscript number: 2017GL072767</p>
Data for "Assessment of using field-aligned currents to drive the Global Ionosphere Thermosphere Model: A case study for the 2013 St Patrick's Day geomagnetic storm"
GITM Simulation results for the paper "Assessment of using field-aligned currents to drive the Global Ionosphere Thermosphere Model: A case study for the 2013 St Patrick's Day geomagnetic storm"
« Lend your Money, Lose your Friend? » - Chinese Official Lending and Bilateral Political Alignment: The Case of Africa
<p>This database provides a set of 45 variables related to UNGA voting affinity vis-à-vis China, official lending, and other bilateral economic and political indicators for China and 43 African countries over the period 2000-2020. The indicators are grouped into four categories: voting data, loan and debt, economic indicators, and political indicators. </p><p>This dataset was compiled in order to conduct research and econometric work for a journal article entitled: </p><p><strong>« Lend your Money, Lose your Friend? » - Chinese Official Lending and Bilateral Political Alignment: The Case of Africa</strong></p><p>, written by Clément Durif, Junior Resarch Fellow at the Asia Centre <a href="mailto:clement.durif@sciencespo.fr">clement.durif@sciencespo.fr</a><br><br>The status of this journal article is pending submittal and acceptation from a Journal Publication</p>
Phylogenomics on Klebsormidiophyceae: Assemblies, SuperTranscripts, BUSCO, Transdecoder, Decontamination, Orthofinder, PhyloPyPruner, Prequal, and concatenated Alignment
<p>Files used for the phylogenomic analysis of Klebsormidiophyceae</p>
Educator Perceptions of DevOps Teaching Recommendations and Their Alignment with Common Challenges
<p>DevOps education presents unique pedagogical challenges due to the diversity of tools, rapid technological change, and the multidisciplinary nature of the field. Although previous work has proposed recommendations to address these challenges, it is unclear how educators perceive these recommendations and whether they align with the challenges encountered in practice. In this paper, we present a mixed methods study involving 11 DevOps educators who interacted with Improve, a tool that presents a curated set of educational challenges and recommendations derived from previous literature. Educators indicated which recommendations they already use, which they intend to use, and which challenges they experience and are motivated to address. Our findings show that 22.6% of the recommendations were new to educators and considered potentially useful, while 59.2% were already in use. Additionally, 66.3% of the challenges were considered relevant, with most of them having linked recommendations that educators were already using or willing to adopt. This study provides empirical insights into the perceived usefulness of existing recommendations, identifies gaps in challenge-recommendation mappings, and supports future efforts to design and disseminate more targeted educational guidance for DevOps teaching.</p>
Multiple sequence alignments: Detection and isolation of a new member of Burkholderiaceae‑related endofungal bacteria from Saksenaea boninensis sp. nov., a new thermotolerant fungus in Mucorales
<p><strong>Methods:</strong></p><p>Nucleotide sequences were aligned independently for each region using MAFFT v7.212 (Katoh and Standley, 2013). The obtained alignment blocks were subject to Gblocks 0.91b (Castresana, 2000) to remove poorly aligned positions with the relaxed selection setting described in Talavera & Castresana (2007) using the following parameters (-t = d -b2 = 9 -b3 = 10 -b4 = 5 -b5 = h). After automatically removing gaps, the alignment blocks were viewed using MEGA 6.06 software (Tamura et al., 2013) and poorly aligned positions at either end of the alignments were removed manually. Pairwise distances of the nucleotide sequences (ITS2, ITS1-5.8S-ITS2, LSU, and tef1) of the ex-type strains of seven <i>Saksenaea</i> spp. and the representative isolate <i>S. boninensis</i> Sak4 were calculated by MEGA 6.06 software (Tamura et al. 2013). Multiple sequence alignment of 16S rRNA gene of the family <i>Burkholderiaceae</i> was prepared for the phylogeny of a bacterial endosymbiont. Multiple sequence alignments of ITS, LSU, and tef1 genes of <i>Saksenaea</i> spp. (Mucorales) were separately prepared for the phylogeny of a fungal host. Concatenated dataset of these genes were also prepared. All nucleotide sequences were retrieved from GenBank (See "Sequence_ID.csv" and taxon names of each alignment). </p><p> </p><p><strong>Description of files:</strong></p><p><strong>A. Phylogeny of the family </strong><i><strong>Burkholderiaceae</strong></i><strong> (Bacterial endosymbiont):</strong></p><p>1. Burkholderiaceae_16S_RAW.fasta</p><p>Non-aligned dataset of 16S rRNA gene of the family <i>Burkholderiaceae</i>.</p><p> </p><p>2. Burkholderiaceae_16S_aligned.fasta</p><p>Aligned dataset of 16S rRNA gene of the family <i>Burkholderiaceae</i>.</p><p> </p><p><strong>B. Phylogenies of </strong><i><strong>Saksenaea</strong></i><strong> spp. (Fungal host):</strong></p><p>1. Sequence_ID_v2.csv</p><p>Taxon names, accession numbers, and sequence ID for the concatenated multiple sequence alignment are listed.</p><p> </p><p>2. Saksenaea_ITS_RAW_v2.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. </p><p> </p><p>3. Saksenaea_ITS_aligned_v2.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. Only used for ITS1-5.8S-ITS2 phylogeny.</p><p> </p><p>4. Saksenaea_LSU_RAW_v2.fasta</p><p>Non-aligned dataset of LSU gene region of <i>Saksenaea</i> spp. </p><p> </p><p>5. Saksenaea_LSU_aligned_v2.fasta</p><p>Aligned dataset of LSU gene region of <i>Saksenaea</i> spp. Only used for LSU phylogeny.</p><p> </p><p>6. Saksenaea_tef1_RAW_v2.fasta</p><p>Non-aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. </p><p> </p><p>7. Saksenaea_tef1_aligned_v2.fasta</p><p>Aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. Only used for tef1 phylogeny.</p><p> </p><p><strong><Concatenated dataset 1 (ITS2, LSU, tef1)></strong></p><p>8. Saksenaea_ITS2_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 1.</p><p> </p><p>9. Saksenaea_ITS2_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 1.</p><p> </p><p>10. Saksenaea_LSU_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of LSU gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>11. Saksenaea_LSU_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of LSU gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>12. Saksenaea_tef1_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>13. Saksenaea_tef1_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of tef1 gene region of <i>Saksenaea</i> spp. used for preparation of concatenated datasets 1 and 2.</p><p> </p><p>14. Saksenaea_ITS2_LSU_tef1_concatenated_dataset1.fasta</p><p>Concatenated dataset of three multiple sequence alignments (9, 11, and 13). This concatenated dataset was used for the main phylogeny of <i>Saksenaea</i> spp.</p><p> </p><p><strong><Concatenated dataset 2 (ITS1-5.8S-ITS2, LSU, tef1)></strong></p><p>15. Saksenaea_ITS_for_concatenated_RAW_v2.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 2.</p><p> </p><p>16. Saksenaea_ITS_for_concatenated_aligned_v2.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of <i>Saksenaea</i> spp. used for preparation of a concatenated dataset 2.</p><p>Blank sequences were inserted for five isolates of <i>Saksenaea longicolla</i> after the alignment.</p><p> </p><p>17.Saksenaea_ITS_LSU_tef1_concatenated_dataset2.fasta</p><p>Concatenated dataset of three multiple sequence alignments (15, 11, and 13). This concatenated dataset was used for the main phylogeny of <i>Saksenaea</i> spp.</p><p> </p><p><strong>C. Pairwise distances of the ex-type strains of </strong><i><strong>Saksenaea</strong></i><strong> spp.</strong></p><p>1. Saksenaea_ITS_type_RAW.fasta</p><p>Non-aligned dataset of ITS1-5.8S-ITS2 region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>2. Saksenaea_ITS2_type_aligned.fasta</p><p>Aligned dataset of ITS2 region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>3.Saksenaea_ITS_type_aligned.fasta</p><p>Aligned dataset of ITS1-5.8S-ITS2 region of the ex-type strains of <i>Saksenaea</i> spp. without <i>Saksenaea longicolla</i>.</p><p> </p><p>4. Saksenaea_LSU_type_RAW.fasta</p><p>Non-aligned dataset of LSU gene region of the ex-type strains of <i>Saksenaea </i>spp.</p><p> </p><p>5. Saksenaea_LSU_type_aligned.fasta</p><p>Aligned dataset of LSU gene region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>6. Saksenaea_tef1_type_RAW.fasta</p><p>Non-aligned dataset of tef1 gene region of the ex-type strains of <i>Saksenaea</i> spp.</p><p> </p><p>7. Saksenaea_tef1_type_aligned.fasta</p><p>Aligned dataset of tef1 gene region of the ex-type strains of <i>Saksenaea</i> spp.</p>
Mitogenome alignment of 159 unique haplotypes representing 455 individual killer whales
<p><span>Genome sequences can reveal the extent of inbreeding in small populations. Here we present the first genomic characterization of type D killer whales, a distinctive eco/morphotype with a circumpolar, subantarctic distribution. Effective population size is the lowest estimated from any killer whale genome and indicates a severe population bottleneck. Consequently, type D genomes show among the highest level of inbreeding reported for any mammalian species (F<sub>ROH</sub> </span><span></span><span> 0.65). Detected recombination events of different haplotypes are up to an order of magnitude rarer than in other killer whale genomes studied to date. Comparison of genomic data from a museum specimen of a type D killer whale that stranded in New Zealand in 1955, with three modern genomes from the Cape Horn area, reveals high covariance and identity-by-state of alleles, suggesting these genomic characteristics and demographic history are shared among different social groups within this morphotype. Limitations to the insights gained in this study stem from the </span><span>non-independence of the three closely related modern genomes, the short coalescence time of most variation within the genomes, and the nonequilibrium population history which violates the assumptions of many model-based methods. Long-range linkage disequilibrium and extensive runs of homozygosity found in type D genomes provide the potential basis for coupling of genetic barriers to gene flow with other killer whale populations, and the distinctive morphology. </span></p>
Alignment of 845 loci for phylogenomic analysis of Klebsormidiophyceae
<p>Improved alignment with more informative loci (845)</p>
Stripe-Like Echoes Scattered from Nighttime F-Region Field-Aligned Irregularities at Low-Latitudes
<p>The dataset reports the estimated vertical Total Electron Content (TEC) from 221 GPS receivers came from the Crustal Movement Observation Network of China on 9 September 2017. Every receiver's data is saved in a TXT file, whose time resolution is thirty seconds.</p>
Data for Aligning rTWT with 802.1Qbv. A Network Calculus Approach
<div>This project stores Matlab scripts to solve the equations from <a href="https://dl.acm.org/doi/10.1145/3565287.3617606" target="_blank" rel="nofollow noreferrer noopener">Aligning rTWT with 802.1Qbv: a Network Calculus Approach </a></div> <div>This information can also be found in the next link, please check it for updates: <a href="https://gitlab.netcom.it.uc3m.es/predict-6g/twtscheduler" target="_blank" rel="noopener">PREDICT 6G / Aligning rTWT with 802.1Qbv. A Network Calculus Approach · GitLab (uc3m.es)</a>. </div> <div> </div> <div><em>This work has been partially funded by the European Commission Horizon Europe SNS JU <a href="https://predict-6g.eu/" target="_blank" rel="nofollow noreferrer noopener">PREDICT-6G</a> (GA 101095890) Project and the Spanish Ministry of Economic Affairs and Digital Transformation and the European Union-NextGenerationEU through the UNICO 5G I+D <a href="https://unica6g.it.uc3m.es/6g-edgedt/" target="_blank" rel="nofollow noreferrer noopener">6G-EDGEDT</a> and 6G-<a href="https://unica6g.it.uc3m.es/6g-datadriven/" target="_blank" rel="nofollow noreferrer noopener">DATADRIVEN</a>. </em></div> <div> </div>
Aligned Closed White Buddha
Scanned in ~20 seconds, with the Android Version of the [Astrivis Smartphone 3D SCanner](http://astrivis.com) Source: Objaverse 1.0 / Sketchfab
Civil War Monument CloudCompare (Base Aligned)
The American Civil War monument dedicated to General Lee located near Bradfordville, FL. This is an experiment to test CloudCompare software by comparing the [2016 version](https://sketchfab.com/models/8e753b01f9594236bc2840f53c5d10e1) of this monument to the [2017 version](https://sketchfab.com/models/7a678d7bc8af477489c1ed287d9ac1ad) after the stone slab was knocked over and reset. Since the stone was reset slightly differently than before, two different comparisons were necessary. This version aligns the base, you can find the face-aligned version [here](https://sketchfab.com/models/384dfb5f86a34193ab5e542f72bd4a8c). The model is of the 2016 version, the colors represent how the 2017 version has changed. Blue means very similar, Red means very dissimilar. * Point cloud created using Agixoft PhotoScan. * Comparison created using CloudCompare * Made from 72 Photos taken with a Canon PowerShot A2500. * Photos and modeling by Tristan Harrenstein Source: Objaverse 1.0 / Sketchfab
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.