Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

50

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

50 results for “genomic alignment”

Learn how ShareScore rates datasets ↗
zenodo24/100

Simulated nucleotide sequences for testing alignment-free genome distance estimates

<p>This repository contains (12&times;500=)6,000 pairs of nucleotide sequences that have been simulated for testing alignment-free genome distance estimates, as described in <a href="https://riojournal.com/article/36178/">Criscuolo (2019)</a>. Given an evolutionary distance <em>d</em> varying from 0.05 to 0.60 (step = 0.05), the program <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> was used to simulate the evolution of 500 nucleotide sequence pairs with <em>d</em> substitution events per character (GTR+&Gamma; evolutionary model).</p> <p>For each of the 12 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 0.60, an XZ-compressed file containing 500 lines is available. Each line contains 18 fields separated by blank spaces:<br> &nbsp; [1] &nbsp; &nbsp; seed value used during simulation,<br> &nbsp; [2] &nbsp; &nbsp; true evolutionary distance <em>d</em> between the two simulated sequences,<br> &nbsp; [3] &nbsp; &nbsp; total number of simulated characters,<br> &nbsp; [4] &nbsp; &nbsp; number of non-indel characters with nucleotide mismatch,<br> &nbsp; [5] &nbsp; &nbsp; number of non-indel characters,<br> &nbsp; [6-9] &nbsp; A, C, G, T frequencies used during simulation,<br> &nbsp; [10-15] &nbsp; GTR parameters used during simulation,<br> &nbsp; [16] &nbsp; &nbsp; &Gamma; distribution parameter used during simulation,<br> &nbsp; [17-18] &nbsp; two simulated sequences with indel events as gaps.</p> <p>Of note, each pair of aligned sequences without gaps can be regenerated using <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> v1.3.4 with parameters from fields [1,3,6-16] and the following two-leaf model tree:</p> <pre>(t1:d,t2:0.000);</pre> <p>where <em>d</em> is given in field [2].</p> <p>___</p> <p>Criscuolo A (2019) <em>A fast alignment-free bioinformatics procedure to infer accurate distance-based phylogenetic trees from genome assemblies</em>. Research Ideas and Outcomes, 5:e36178. doi:<a href="https://doi.org/10.3897/rio.5.e36178">10.3897/rio.5.e36178</a></p>

opencc-by-4.0Sep 2020View details →
dryad24/100

Alignment-free methods for polyploid genomes: quick and reliable genetic distance estimation

<p>Polyploid genomes pose several inherent challenges to population genetic analyses. While alignment-based methods are fundamentally limited in their applicability to polyploids, alignment-free methods bypass most of these limits. We investigated the use of Mash, a k-mer analysis tool that uses the MinHash method to reduce complexity in large genomic datasets, for basic population genetic analyses of polyploid sequences. We measured the degree to which Mash correctly estimated pairwise genetic distance in simulated haploid and polyploid short-read sequences with various levels of missing data. Mash-based estimates of genetic distance were comparable to alignment-based estimates, and were less impacted by missing data. We also used Mash to analyze publicly available short-read data for three polyploid and one diploid species, then compared Mash results to published results. For both simulated and real data, Mash accurately estimated pairwise genetic differences for polyploids as well as diploids as much as 476 times faster than alignment-based methods, though we found that Mash genetic distance estimates could be biased by per-sample read depth. Mash may be a particularly useful addition to the toolkit of polyploid geneticists for rapid confirmation of alignment-based results and for basic population genetics in reference-free systems or those with only poor quality sequence data available.</p>

opencc-zeroJul 2021View details →
dryad24/100

Alignment-free methods for polyploid genomes: quick and reliable genetic distance estimation

Open the record for dataset details and reuse information.

publicJul 2021View details →
geo24/100

Stage-resolved genome architecture maps throughout meiotic prophase link regional variations in chromosome organization with homolog alignment (Hi-C, Cut&Tag, RNA-seq)

GEO Series GSE155967. Mus musculus. 22 samples. Type: Other; Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by high throughput sequencing.

openGEO-OpenAug 2021View details →
geo20/100

Stage-resolved genome architecture maps throughout meiotic prophase link regional variations in chromosome organization with homolog alignment

GEO Series GSE155638. Mus musculus. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenAug 2021View details →
zenodo20/100

Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part2 - genomic alignments (hg19 + hg38)

<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs &quot;181114_M00528_0390_000000000-C7P58&quot; (aka &quot;NC_LIMMS3&quot;) and &quot;190218_M00528_0406_000000000-CB4HR&quot; (aka &quot;NC_LIMMS4&quot;) FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics&nbsp;2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under&nbsp;the following Digital Object Identifier: 10.5281/zenodo.2572390.</p> <p><em><strong>&quot;181114_M00528_0390_000000000-C7P58&quot; (&quot;NC_LIMMS3&quot;):</strong></em></p> <p><strong>sample_name&nbsp;&nbsp; &nbsp;group&nbsp;&nbsp; &nbsp;barcode_sequence &nbsp;&nbsp; index_sequence</strong><br> LIMMS43_04_PETRI_S4D7_rep1&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;ACAGAT&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS44_24_PETRI_S4D7_rep2&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;ATCGTG&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS45_31_PETRI_S4D7_rep3&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;CACGAT&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS46_36_PETRI_S4D14_rep1&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;CACTGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS47_46_PETRI_S4D14_rep2&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;CTGACG&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS48_63_PETRI_S4D14_rep3&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;GAGTGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS49_79_PETRI_CELLARTIS_rep1&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;GTATAC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS50_92_PETRI_CELLARTIS_rep2&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;TCGAGC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS51_09_PETRI_CELLARTIS_rep3&nbsp;&nbsp; &nbsp;iPSC_CLONE_TODAI&nbsp;&nbsp; &nbsp;ACATGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS52_21_PETRI_TODAI_rep1&nbsp;&nbsp; &nbsp;iPSC_CLONE_CELLARTIS&nbsp;&nbsp; &nbsp;ATCATA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS53_33_PETRI_TODAI_rep2&nbsp;&nbsp; &nbsp;iPSC_CLONE_CELLARTIS&nbsp;&nbsp; &nbsp;CACGTG&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS54_45_PETRI_TODAI_rep3&nbsp;&nbsp; &nbsp;iPSC_CLONE_CELLARTIS&nbsp;&nbsp; &nbsp;CGATGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS55_57_iPSC_rep1&nbsp;&nbsp; &nbsp;CONTROL_iPSC&nbsp;&nbsp; &nbsp;GAGATA&nbsp;&nbsp; &nbsp;NNNNNNNN</p> <p><em><strong>&quot;190218_M00528_0406_000000000-CB4HR&quot; (&quot;NC_LIMMS4&quot;):</strong></em></p> <p><strong>sample_name&nbsp;&nbsp; &nbsp;group&nbsp;&nbsp; &nbsp;barcode_sequence &nbsp;&nbsp; index_sequence</strong><br> LIMMS56_04_iPSC_rep4&nbsp;&nbsp; &nbsp;CONTROL_iPSC&nbsp;&nbsp; &nbsp;ACAGAT&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS57_24_LSECS_1_11&nbsp;&nbsp; &nbsp;LSECS_PETRI_MONO&nbsp;&nbsp; &nbsp;ATCGTG&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS58_31_LSECS_2_11&nbsp;&nbsp; &nbsp;LSECS_PETRI_MONO&nbsp;&nbsp; &nbsp;CACGAT&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS59_36_LSECS_3_11&nbsp;&nbsp; &nbsp;LSECS_PETRI_MONO&nbsp;&nbsp; &nbsp;CACTGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS60_46_LSECS_1-06&nbsp;&nbsp; &nbsp;LSECS_PETRI_MONO&nbsp;&nbsp; &nbsp;CTGACG&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS61_63_B3_MONO_11_D3&nbsp;&nbsp; &nbsp;BC_MONO_D3&nbsp;&nbsp; &nbsp;GAGTGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS62_79_B9_CO_10_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;GTATAC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS63_92_B13_CO_11_D3&nbsp;&nbsp; &nbsp;BC_CO_D3&nbsp;&nbsp; &nbsp;TCGAGC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS64_09_P2_10_D14&nbsp;&nbsp; &nbsp;PETRI_MONO&nbsp;&nbsp; &nbsp;ACATGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS65_21_P3_10_D14&nbsp;&nbsp; &nbsp;PETRI_MONO&nbsp;&nbsp; &nbsp;ATCATA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS66_33_P3_11_D14&nbsp;&nbsp; &nbsp;PETRI_MONO&nbsp;&nbsp; &nbsp;CACGTG&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS67_45_B1_MONO_10_D14&nbsp;&nbsp; &nbsp;BC_MONO_D14&nbsp;&nbsp; &nbsp;CGATGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS68_57_B2_MONO_10_D14&nbsp;&nbsp; &nbsp;BC_MONO_D14&nbsp;&nbsp; &nbsp;GAGATA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS69_69_B1_MONO_11_D14&nbsp;&nbsp; &nbsp;BC_MONO_D14&nbsp;&nbsp; &nbsp;GCTCTC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS70_81_B2_MONO_11_D14&nbsp;&nbsp; &nbsp;BC_MONO_D14&nbsp;&nbsp; &nbsp;GTATGA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS71_93_B6_CO_10_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;TCGATA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS72_11_B7_CO_10_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;AGTAGC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS73_23_B8_CO_10_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;ATCGCA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS74_35_B9_CO_11_D3&nbsp;&nbsp; &nbsp;BC_CO_D3&nbsp;&nbsp; &nbsp;CACTCT&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS75_47_B11_CO_11_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;CTGAGC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS76_59_B12_CO_11_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;GAGCGT&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS77_71_B14_CO_11_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;GCTGCA&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS78_83_B15_CO_11_D14&nbsp;&nbsp; &nbsp;BC_CO_D14&nbsp;&nbsp; &nbsp;TATAGC&nbsp;&nbsp; &nbsp;NNNNNNNN<br> LIMMS79_95_iPSC_rep1_4&nbsp;&nbsp; &nbsp;CONTROL_iPSC&nbsp;&nbsp; &nbsp;TCGCGT&nbsp;&nbsp; &nbsp;NNNNNNNN</p>

restrictedFeb 2019View details →
zenodo20/100

Genome in a Bottle Direct-RNA Sequencing: GM24631 Calibration-Strand Aligned Reads

Open the record for dataset details and reuse information.

openMay 2024View details →
geo12/100

Circadian regulation of transcriptome in spinach leaves under varying nitrogen levels (aligned to Sp75 reference genome)

GEO Series GSE275461. Spinacia oleracea. 24 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2025View details →
zenodo12/100

Multiple sequence alignments of Treponema pallidum complete genomes using three different references for mapping NGS reads

<p>Each file corresponds to the multiple sequence alignment of 75 complete Treponema pallidum genome sequences using the genomes of strains Nichols, SS14, and CDC-2 as references for mapping. This is supplemental data to the manuscript &quot;Evolutionary processes in the emergence and recent spread of <em>Treponema pallidum</em>, the causative agent of syphilis&quot; by Marta Pla-D&iacute;az, Leonor S&aacute;nchez-Bus&oacute;, Lorenzo Giacani, David &Scaron;majs, Philipp P. Bosshard, Homayoun C. Bagheri, Verena J. Schuenemann, Kay Nieselt, Natasha Arora and Fernando Gonz&aacute;lez-Candelas</p>

restrictedAug 2021View details →
zenodo4/100

Seeker: Alignment-free identification of bacteriophage genomes by deep learning

<p>Training and testing data.</p>

restrictedSep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record