Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,199

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,199 results for “aligners”

Learn how ShareScore rates datasets ↗
zenodo40/100

CONCATENATING sample files to prepare for multiple sequence alignment in Galaxy

<p>These are a few sample files to practice the correct way to concatenate&nbsp;files with the reference strain at the top, in order to continue with the next step, which is doing a multiple sequence alignment.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Aligning Active Particles Simulation Data

<p>Simulation Data for Vicsek Model and related models, created by aappp software, see: https://github.org/kuersten/aappp/</p> <p>The example simulation data presented here show some of the features of the aappp software</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Text-fig. 27. Scanning electron microscope (SEM) images of stamens and pollen of Valvidistemon globiferus gen. et sp. nov.; Catefica locality, Portugal. a) Stamen in oblique lateral view showing laterally hinged valves, massive connective between the thecae and prominent, globular, apical extension of the connective; b) Stamen in oblique lateral view on the opposite side from (a) showing broken laterally hinged valves and distinct endothecium cells; c) Detail of stamen showing the large, longitudinally aligned cells of the massive connective, broad, poorly defined stamen base, and the laterally hinged valves of one of the thecae; d) Reticulate pollen attached to the inside of the anther wall. Specimen, Catefica 49-S107779 (holotype, a–d). Scale bars = 600 Μm (a, b), 100 Μm (c), 20 Μm (d). in The Early Cretaceous Mesofossil Flora Of Catefica, Portugal: Angiosperms

Text-fig. 27. Scanning electron microscope (SEM) images of stamens and pollen of Valvidistemon globiferus gen. et sp. nov.; Catefica locality, Portugal. a) Stamen in oblique lateral view showing laterally hinged valves, massive connective between the thecae and prominent, globular, apical extension of the connective; b) Stamen in oblique lateral view on the opposite side from (a) showing broken laterally hinged valves and distinct endothecium cells; c) Detail of stamen showing the large, longitudinally aligned cells of the massive connective, broad, poorly defined stamen base, and the laterally hinged valves of one of the thecae; d) Reticulate pollen attached to the inside of the anther wall. Specimen, Catefica 49-S107779 (holotype, a–d). Scale bars = 600 Μm (a, b), 100 Μm (c), 20 Μm (d).

opencc-by-4.0Dec 2022View details →
zenodo40/100

Dataset for the Galaxy Training Network (GTN) Tutorial "Viewing Cancer Alignments in a Genome Browser"

<p>Datasets for the Galaxy Training Network (GTN) Tutorial &quot;Viewing Cancer Alignments in a Genome Browser&quot;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Assemblies and alignment data generated for NAHRwhals manuscript.

<p>This repository contains 56 human assemblies used for a manuscript describing the NAHRwhals SV identifying tool (<a href="https://github.com/WHops/NAHRwhals" target="_new" rel="noreferrer">https://github.com/WHops/NAHRwhals</a>). All underlying raw data as well as half of the assemblies are directly taken from the Human Genome Structural Variation Consortium (HGSVC). Raw HiFi reads underlying the assemblies can be obtained from:&nbsp;<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/" target="_new" rel="noreferrer">http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/</a>.</p> <p>Assemblies were created with in two batches. Batch one (HG00512, HG00513, HG00514, HG00731, HG00732, HG00733, HG02818, HG03125, HG03486, NA12878, NA19238, NA19239, NA19240, NA24385) was created by the HGSVC (Ebert et al. 2021) and used in the NAHRwhals manuscript. Batch two (GM19129, GM19434, HG00171, HG00864, HG02018, HG02282, HG02769, HG02953, HG03452, HG03520, NA12329, NA19036, NA19983, NA20847) is based on HGSVC raw data but was created specifically for the manuscript by Tobias Rausch.</p> <p>For more information, please refer to the NAHRwhals paper "Impact and characterization of serial structural variations across humans and great apes" by H&ouml;ps et al., 2024 for further details on assembly generation and intended usage.</p> <p>Contact: <a rel="noreferrer">wolfram.hoeps@gmail.com</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

pWCP alignment sequences

<p>Alignment sequences of pWCP across locations.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

RVFV data aligning only to M-Fragment to test the PARANOiD pipeline

<p>PARANOiD is a versatile software for fully automated analysis of iCLIP and iCLIP2 data. It contains all steps necessary for preprocessing, the determination of cross-link locations and several additional steps, which can be used to detect specific characteristics, e.g. definite distances between cross-link events or identify binding motifs. The cross-link sites are presented as WIG files that can be easily visualized e.g. using IGV, for which a config file is offered. Additionally, results are offered as statistical plots for a quick overview and as standardized bioinformatics file formats or TSV files, which can be used for further analysis steps.</p> <p>The data provided are used as a test case for PARANOiD.</p> <p>The data was extracted from RVFV MP-12 virions (virion-reads-M-fragment-only.fastq) and BHK cells infected with RVFV (BHK-reads-M-fragment-only.fastq) applying the iCLIP2 method for RVFV N iCLIP. Three independent biological replicates were performed for each sample. Sequencing was performed using the MiSeq Sequencer (Illumina) with MiSeq Reagent Kit v2 Micro (Illumina) for N-iCLIP from virus particles and MiSeq Reagent Kit v3 (Illumina) for N-iCLIP from infected BHK cells.</p> <p>The original reads have been aligned to the RVFV MP-12 reference genome and only reads aligning to the M-fragments were extracted. The whole dataset will be publish at a later date</p> <p>File description:</p> <p>virion-reads-M-fragment-only.fastq - Reads obtained from RVFV virions</p> <p>BHK-reads-M-fragment-only.fastq - Reads obtained from BHK cells infected with RVFV</p> <p>reference_RVFV.fasta - RVFV MP-12 reference genome</p> <p>barcodes-RVFV.tsv - Barcodes for virion-reads-M-fragment-only.fastq</p> <p>barcodes-RVFV-merge-all.tsv - Barcodes for merging all samples of virion-reads-M-fragment-only.fastq</p> <p>barcodes-BHK.tsv - Barcodes for BHK-reads-M-fragment-only.fastq</p> <p>barcodes-BHK-merge-all.tsv - Barcodes for merging all relevant samples of BHK-reads-M-fragment-only.fastq</p>

opencc-by-4.0May 2023View details →
dryad40/100

Machine learning can be as good as maximum likelihood when reconstructing phylogenetic trees and determining the best evolutionary model on four taxon alignments

<p><span>Machine learning can be as good as maximum likelihood when reconstructing phylogenetic topologies and determining the best evolutionary model on four taxon alignments.</span></p> <p><span>Phylogenetic tree reconstruction with molecular data is important in many fields of life science research. The gold standard in this discipline is the Maximum Likelihood tree reconstruction method. Here we show that for quartet trees, Machine Learning using neural networks can be as good as the Maximum Likelihood method to infer the best tree topology and the best model of sequence evolution for nucleotide as well as amino acid sequences. For this purpose we simulated data sets for a wide range of branch lengths, evolutionary models and model parameters and compared the topologies and inferred models obtained with Machine learning with those obtained with the Maximum Likelihood and the Neighbour Joining method. Our results show that neural networks are a promising avenue for determining relatedness between taxa, which is likely to accelerate the construction of phylogenetic trees in the future, while maintaining a high accuracy.</span></p>

opencc-zeroMar 2023View details →
zenodo40/100

Additional annotation, alignment, and results from Ka/Ks analysis for Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus)

<p><strong>Annotation files, alignments, and results summaries from&nbsp;Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus).</strong></p> <p>Pairwise genome alignments contain the .maf suffix</p> <p>FASTA alignments from stitched gene blocks&nbsp;contain the .fasta suffix</p> <p>CSV file containing the Ka/Ks results</p> <p>RepeatMasker .out file</p>

opencc-by-4.0Mar 2023View details →
dryad40/100

Data from: Strong selection is poorly aligned with genetic variation in Ipomoea hederacea

<p><span>The multivariate evolution of populations is the result of the interactions between natural selection, drift, and the underlying genetic structure of the traits involved. Covariances among traits bias responses to selection, and the multivariate axis which describes the greatest genetic variation is expected to be aligned with patterns of divergence across populations. An exception to this expectation is when selection acts on trait combinations lacking genetic variance, which limits evolutionary change. Here we used a common garden field experiment of individuals from 57 populations of <em>Ipomoea hederacea</em> to characterize linear and nonlinear selection on five quantitative traits in the field. We then formally compare patterns of selection to previous estimates of within-population genetic covariance structure (the G-matrix) and population divergence in these traits. We found that selection is poorly aligned with previous estimates of genetic covariance structure and population divergence. In addition, the trait combinations favoured by selection were generally lacking genetic variation, possessing approximately 15-30% as much genetic variation as the most variable combination of traits. Our results suggest that patterns of population divergence are likely the result of the interplay between adaptive responses, correlated response, and selection favoring traits lacking genetic variation.  </span></p>

opencc-zeroApr 2023View details →
zenodo40/100

Sentence-aligned student translations of Crito (Ancient Greek, English, German, Persian)

<p>This dataset is a corpus of five student&#39;s translation of Plato&#39;s Crito aligned at sentence-level with the original Ancient Greek text, one German translation, and two English translations. The Ancient Greek text is the Burnet edition, made available by Perseus Digital Library. For more information, see:&nbsp;<br> https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0169%3Atext%3DCrito%3Asection%3D43a<br> &nbsp;</p> <p>The details of the eight translations (five Persian, two English, and one German translations) are as follows:</p> <ul> <li>German Translation:&nbsp;The German translation of Schleiermmacher available on Project Gutenberg has been aligned at the sentence level.<br> For more information, see: Plato, F. Schleiermacher, Platons Werke, In der Realschulbuchhandlung, 1809<br> Link to the text on Project Gutenberg:<br> https://www.projekt-gutenberg.org/platon/platowr1/kriton.html<br> &nbsp;</li> <li>English Translations: Two different English translations of &quot;Crito&quot;, one by Benjamin Jowett and the other by Harold North Fowler are included in the dataset.<br> For more information on Jowett&#39;s translation, see:<br> Plato, H. N. Fowler, W. Lamb, Plato in Twelve Volumes, Vol. 1 translated by Harold North<br> Fowler; Introduction by W.R.M. Lamb, volume 1, Harvard University Press and Wiliam<br> Heinemann Ltd., Cambridge, MA and London, 1966.<br> Fowler&#39;s translation on Perseus Digital Library:<br> https://www.perseus.tufts.edu/hopper/text?doc=plat.+crito+43a<br> For more information on Jowett&#39;s translation, see:<br> Plato, B. Jowett, Crito, The Internet Classics Archive, Massachusetts Institute of Technology,<br> http://classics.mit.edu/Plato/crito.html.<br> &nbsp;</li> <li>Persian Translations: The dataset consists of five Persian translations by students who have already completed a 30-hour Homeric Greek course. Each translator has translated the text into Persian using treebanks, commentaries, lexicon entries, and English and German translations. The translators themselves aligned the Persian translations to the Greek text at word-level using Ugarit. The alignments are available in their Ugarit profile:<br> Shouresh Assimi: https://ugarit.ialigner.com/userProfile.php?userid=50956<br> Aylar Mahmoudzadeh Sarabi: https://ugarit.ialigner.com/userProfile.php?userid=63464&amp;tgid=9576<br> Nima Mohammadi: https://ugarit.ialigner.com/userProfile.php?userid=52434&amp;tgid=9362<br> Kimia Nikpour: https://ugarit.ialigner.com/userProfile.php?userid=52378<br> Farshid Rahimi: https://ugarit.ialigner.com/userProfile.php?userid=50932&amp;tgid=9727</li> </ul> <p>The group&#39;s initial goal was to produce one finalized translation of Crito to Persian, but due to the intriguing variations in the translations and the text&#39;s intricacy, it was decided to provide three finalized translations rather than one. The finalized translations will be available in Beyond Translation as part of the Perseus Digital Library under a Creative Commons license. For more information on our final versions of Crito, see:&nbsp;http://beyond-translation.perseus.org</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Identifying and Aligning Medical Claims Made on Social Media with Medical Evidence

<p>A synthetically generated dataset of medical claims.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

FIGURE 4 Maximum Clade Credibility Tree inferred using a concatenate COI, 16S, 28S and 18S alignment using BEAST. Node bars are 95 in The role of allopatric speciation and ancient origins of Bathynellidae (Crustacea) in the Pilbara (Western Australia): two new genera from the De Grey River catchment

FIGURE 4 Maximum Clade Credibility Tree inferred using a concatenate COI, 16S, 28S and 18S alignment using BEAST. Node bars are 95% Higher Posterior Density, scale bar is in million years ago (Ma), starting from present 0. Numbers above bars = node age; numbers below bars (bold) = posterior probability of the node.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Data for "On optimizing mass spectral library search and spectral alignment by weighting low-intensity peaks and m/z frequency"

<p>Data for <a href="https://github.com/enveda/weighting-spectral-similarity#on-optimizing-mass-spectral-library-search-and-spectral-alignment-by-weighting-low-intensity-peaks-and-mz-frequency">On optimizing mass spectral library search and spectral alignment by weighting low-intensity peaks and m/z frequency</a>.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Replication data for: The Effect of IS-Innovation Strategy Alignment on Corporate Performance: Investigating the Role of Environmental Uncertainty by Heterogeneity

<p>A base de dados t&eacute;cnico-cient&iacute;fica de 856 empresas brasileiras examinou o alinhamento entre sistemas de informa&ccedil;&atilde;o estrat&eacute;gicos (ISS), na abordagem da estrat&eacute;gia como pr&aacute;tica, &nbsp;e inova&ccedil;&atilde;o de exploration e exploitation, e seu impacto no desempenho corporativo (CP) sobre incerteza ambiental. Os resultados mostram que todos os tipos de alinhamento entre ISS e inova&ccedil;&atilde;o influenciam positivamente o CP. O alinhamento com inova&ccedil;&atilde;o ambidestra teve um impacto 62% maior no CP do que o alinhamento com inova&ccedil;&otilde;es incrementais. Al&eacute;m disso, inova&ccedil;&otilde;es disruptivas tiveram efeitos positivos em ambientes hostis, enquanto inova&ccedil;&otilde;es explorat&oacute;rias e ambidestras tiveram impactos fortes em ambientes altamente din&acirc;micos.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Plant Cassandra retrotransposons: LTR alignment data and Astereaceae full length annotation

<p>Supplemental material to the article:&nbsp;</p><p>&nbsp;</p><p><strong>"Evolving together: Cassandra retrotransposons gradually mirror promoter mutations of the 5S rRNA genes"</strong></p><p><strong>Abstract:</strong></p><p>&nbsp;The 5S rRNA genes are among the most conserved nucleotide sequences across all species. Similar to the 5S preservation we observe the occurrence of 5S-related non-autonomous retrotransposons, so-called Cassandra. Cassandras harbor highly conserved 5S rDNA-related sequences within their long terminal repeats (LTRs), advantageously providing them with the 5S internal promoter. However, the dynamics of Cassandra retrotransposon evolution in the context of 5S rRNA gene sequence information and structural arrangement are still unclear, especially: 1) do we observe repeated or gradual domestication of the highly conserved 5S promoter by Cassandras and 2) do changes in 5S organization such as in the linked 35S-5S rDNA arrangements impact Cassandra evolution? Here, we show evidence for gradual co-evolution of Cassandra sequences with their corresponding 5S rDNAs. To follow the impact of 5S rDNA variability on Cassandra TEs, we investigate the Asteraceae family where highly variable 5S rDNAs, including 5S promoter shifts and both linked and separated 35S-5S rDNA arrangements have been reported. Cassandras within the Asteraceae mirror 5S rDNA promoter mutations of their host genome, likely as an adaptation to the host's specific 5S transcription factors and hence compensating for evolutionary changes in the 5S rDNA sequence. Changes in the 5S rDNA sequence and in Cassandras seem uncorrelated with linked/separated rDNA arrangements. We place all these observations into the context of angiosperm 5S rDNA-Cassandra evolution, discuss Cassandra's origin hypotheses (single or multiple) and Cassandra's possible impact on rDNA and plant genome organization, giving new insights into the interplay of ribosomal genes and transposable elements.</p>

opencc-by-4.0Oct 2023View details →
dryad40/100

Interstitial cortisol measurements aligned by wake time, healthy volunteers

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad40/100

The primate Major Histocompatibility Complex: Sets of posterior trees from BEAST2 for the whole-class multi-gene alignments

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad40/100

Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II

Open the record for dataset details and reuse information.

publicNov 2020View details →
dryad40/100

CASTER: Direct species tree inference from whole-genome alignments

Open the record for dataset details and reuse information.

publicNov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record