Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
50
datasets available to search
ShareScore release 0.7.1
Dataset results
50 results for “genomic alignment”
Simulated nucleotide sequences for testing alignment-free genome distance estimates
<p>This repository contains (12×500=)6,000 pairs of nucleotide sequences that have been simulated for testing alignment-free genome distance estimates, as described in <a href="https://riojournal.com/article/36178/">Criscuolo (2019)</a>. Given an evolutionary distance <em>d</em> varying from 0.05 to 0.60 (step = 0.05), the program <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> was used to simulate the evolution of 500 nucleotide sequence pairs with <em>d</em> substitution events per character (GTR+Γ evolutionary model).</p> <p>For each of the 12 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 0.60, an XZ-compressed file containing 500 lines is available. Each line contains 18 fields separated by blank spaces:<br> [1] seed value used during simulation,<br> [2] true evolutionary distance <em>d</em> between the two simulated sequences,<br> [3] total number of simulated characters,<br> [4] number of non-indel characters with nucleotide mismatch,<br> [5] number of non-indel characters,<br> [6-9] A, C, G, T frequencies used during simulation,<br> [10-15] GTR parameters used during simulation,<br> [16] Γ distribution parameter used during simulation,<br> [17-18] two simulated sequences with indel events as gaps.</p> <p>Of note, each pair of aligned sequences without gaps can be regenerated using <a href="http://tree.bio.ed.ac.uk/software/seqgen/">SeqGen</a> v1.3.4 with parameters from fields [1,3,6-16] and the following two-leaf model tree:</p> <pre>(t1:d,t2:0.000);</pre> <p>where <em>d</em> is given in field [2].</p> <p>___</p> <p>Criscuolo A (2019) <em>A fast alignment-free bioinformatics procedure to infer accurate distance-based phylogenetic trees from genome assemblies</em>. Research Ideas and Outcomes, 5:e36178. doi:<a href="https://doi.org/10.3897/rio.5.e36178">10.3897/rio.5.e36178</a></p>
Alignment-free methods for polyploid genomes: quick and reliable genetic distance estimation
<p>Polyploid genomes pose several inherent challenges to population genetic analyses. While alignment-based methods are fundamentally limited in their applicability to polyploids, alignment-free methods bypass most of these limits. We investigated the use of Mash, a k-mer analysis tool that uses the MinHash method to reduce complexity in large genomic datasets, for basic population genetic analyses of polyploid sequences. We measured the degree to which Mash correctly estimated pairwise genetic distance in simulated haploid and polyploid short-read sequences with various levels of missing data. Mash-based estimates of genetic distance were comparable to alignment-based estimates, and were less impacted by missing data. We also used Mash to analyze publicly available short-read data for three polyploid and one diploid species, then compared Mash results to published results. For both simulated and real data, Mash accurately estimated pairwise genetic differences for polyploids as well as diploids as much as 476 times faster than alignment-based methods, though we found that Mash genetic distance estimates could be biased by per-sample read depth. Mash may be a particularly useful addition to the toolkit of polyploid geneticists for rapid confirmation of alignment-based results and for basic population genetics in reference-free systems or those with only poor quality sequence data available.</p>
Alignment-free methods for polyploid genomes: quick and reliable genetic distance estimation
Open the record for dataset details and reuse information.
Stage-resolved genome architecture maps throughout meiotic prophase link regional variations in chromosome organization with homolog alignment (Hi-C, Cut&Tag, RNA-seq)
GEO Series GSE155967. Mus musculus. 22 samples. Type: Other; Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by high throughput sequencing.
Stage-resolved genome architecture maps throughout meiotic prophase link regional variations in chromosome organization with homolog alignment
GEO Series GSE155638. Mus musculus. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part2 - genomic alignments (hg19 + hg38)
<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs "181114_M00528_0390_000000000-C7P58" (aka "NC_LIMMS3") and "190218_M00528_0406_000000000-CB4HR" (aka "NC_LIMMS4") FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics 2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under the following Digital Object Identifier: 10.5281/zenodo.2572390.</p> <p><em><strong>"181114_M00528_0390_000000000-C7P58" ("NC_LIMMS3"):</strong></em></p> <p><strong>sample_name group barcode_sequence index_sequence</strong><br> LIMMS43_04_PETRI_S4D7_rep1 iPSC_CLONE_TODAI ACAGAT NNNNNNNN<br> LIMMS44_24_PETRI_S4D7_rep2 iPSC_CLONE_TODAI ATCGTG NNNNNNNN<br> LIMMS45_31_PETRI_S4D7_rep3 iPSC_CLONE_TODAI CACGAT NNNNNNNN<br> LIMMS46_36_PETRI_S4D14_rep1 iPSC_CLONE_TODAI CACTGA NNNNNNNN<br> LIMMS47_46_PETRI_S4D14_rep2 iPSC_CLONE_TODAI CTGACG NNNNNNNN<br> LIMMS48_63_PETRI_S4D14_rep3 iPSC_CLONE_TODAI GAGTGA NNNNNNNN<br> LIMMS49_79_PETRI_CELLARTIS_rep1 iPSC_CLONE_TODAI GTATAC NNNNNNNN<br> LIMMS50_92_PETRI_CELLARTIS_rep2 iPSC_CLONE_TODAI TCGAGC NNNNNNNN<br> LIMMS51_09_PETRI_CELLARTIS_rep3 iPSC_CLONE_TODAI ACATGA NNNNNNNN<br> LIMMS52_21_PETRI_TODAI_rep1 iPSC_CLONE_CELLARTIS ATCATA NNNNNNNN<br> LIMMS53_33_PETRI_TODAI_rep2 iPSC_CLONE_CELLARTIS CACGTG NNNNNNNN<br> LIMMS54_45_PETRI_TODAI_rep3 iPSC_CLONE_CELLARTIS CGATGA NNNNNNNN<br> LIMMS55_57_iPSC_rep1 CONTROL_iPSC GAGATA NNNNNNNN</p> <p><em><strong>"190218_M00528_0406_000000000-CB4HR" ("NC_LIMMS4"):</strong></em></p> <p><strong>sample_name group barcode_sequence index_sequence</strong><br> LIMMS56_04_iPSC_rep4 CONTROL_iPSC ACAGAT NNNNNNNN<br> LIMMS57_24_LSECS_1_11 LSECS_PETRI_MONO ATCGTG NNNNNNNN<br> LIMMS58_31_LSECS_2_11 LSECS_PETRI_MONO CACGAT NNNNNNNN<br> LIMMS59_36_LSECS_3_11 LSECS_PETRI_MONO CACTGA NNNNNNNN<br> LIMMS60_46_LSECS_1-06 LSECS_PETRI_MONO CTGACG NNNNNNNN<br> LIMMS61_63_B3_MONO_11_D3 BC_MONO_D3 GAGTGA NNNNNNNN<br> LIMMS62_79_B9_CO_10_D14 BC_CO_D14 GTATAC NNNNNNNN<br> LIMMS63_92_B13_CO_11_D3 BC_CO_D3 TCGAGC NNNNNNNN<br> LIMMS64_09_P2_10_D14 PETRI_MONO ACATGA NNNNNNNN<br> LIMMS65_21_P3_10_D14 PETRI_MONO ATCATA NNNNNNNN<br> LIMMS66_33_P3_11_D14 PETRI_MONO CACGTG NNNNNNNN<br> LIMMS67_45_B1_MONO_10_D14 BC_MONO_D14 CGATGA NNNNNNNN<br> LIMMS68_57_B2_MONO_10_D14 BC_MONO_D14 GAGATA NNNNNNNN<br> LIMMS69_69_B1_MONO_11_D14 BC_MONO_D14 GCTCTC NNNNNNNN<br> LIMMS70_81_B2_MONO_11_D14 BC_MONO_D14 GTATGA NNNNNNNN<br> LIMMS71_93_B6_CO_10_D14 BC_CO_D14 TCGATA NNNNNNNN<br> LIMMS72_11_B7_CO_10_D14 BC_CO_D14 AGTAGC NNNNNNNN<br> LIMMS73_23_B8_CO_10_D14 BC_CO_D14 ATCGCA NNNNNNNN<br> LIMMS74_35_B9_CO_11_D3 BC_CO_D3 CACTCT NNNNNNNN<br> LIMMS75_47_B11_CO_11_D14 BC_CO_D14 CTGAGC NNNNNNNN<br> LIMMS76_59_B12_CO_11_D14 BC_CO_D14 GAGCGT NNNNNNNN<br> LIMMS77_71_B14_CO_11_D14 BC_CO_D14 GCTGCA NNNNNNNN<br> LIMMS78_83_B15_CO_11_D14 BC_CO_D14 TATAGC NNNNNNNN<br> LIMMS79_95_iPSC_rep1_4 CONTROL_iPSC TCGCGT NNNNNNNN</p>
Genome in a Bottle Direct-RNA Sequencing: GM24631 Calibration-Strand Aligned Reads
Open the record for dataset details and reuse information.
Circadian regulation of transcriptome in spinach leaves under varying nitrogen levels (aligned to Sp75 reference genome)
GEO Series GSE275461. Spinacia oleracea. 24 samples. Type: Expression profiling by high throughput sequencing.
Multiple sequence alignments of Treponema pallidum complete genomes using three different references for mapping NGS reads
<p>Each file corresponds to the multiple sequence alignment of 75 complete Treponema pallidum genome sequences using the genomes of strains Nichols, SS14, and CDC-2 as references for mapping. This is supplemental data to the manuscript "Evolutionary processes in the emergence and recent spread of <em>Treponema pallidum</em>, the causative agent of syphilis" by Marta Pla-Díaz, Leonor Sánchez-Busó, Lorenzo Giacani, David Šmajs, Philipp P. Bosshard, Homayoun C. Bagheri, Verena J. Schuenemann, Kay Nieselt, Natasha Arora and Fernando González-Candelas</p>
Seeker: Alignment-free identification of bacteriophage genomes by deep learning
<p>Training and testing data.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.