Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

50

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

50 results for “genomic alignment”

Learn how ShareScore rates datasets ↗
dryad32/100

Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)

<p class="BodyA"><span><b>Background:</b> The great diversity in plant genome size and chromosome number is partly due to polyploidization (i.e., genome doubling events). The differences in genome size and chromosome number among diploid plant species can be a window into the intriguing phenomenon of past genome doubling that may be obscured through time by the process of diploidization. The genus <i>Hibiscus </i>L. (Malvaceae) has a wide diversity of chromosome numbers and a complex genomic history. <i>Hibiscus </i>is ideal for exploring past genomic events because although two ancient genome duplication events have been identified, more are likely to be found due to its diversity of chromosome numbers. To reappraise the history of whole genome duplication events, we tested  three alternative scenarios describing different polyploidization events.</span></p> <p class="BodyA"><span><b>Results:</b> Using target sequence capture, we designed a new probe set for <i>Hibiscus </i>and generated 87 orthologous genes from four diploid species. We detected paralogues in &gt;54% putative single-copy genes. 34 of these genes were selected for testing three different genome duplication scenarios using gene counting. All species of <i>Hibiscus</i> sampled shared one genome duplication with <i>H. syriacus</i> and one whole genome duplication occurred along the branch leading to <i>H. syriacus</i>.</span></p> <p class="BodyA"><span><b>Conclusions:</b> Here, we corroborated the independent genome doubling previously found in the lineage leading to <i>H. syriacus </i>and a shared genome doubling of this lineage and the remainder of <i>Hibiscus</i>. Additionally, we found a previously undiscovered genome duplication shared by the /Pavonia and /Malvaviscus clades (both nested within <i>Hibiscus</i>) with the occurrences of two copies in what were otherwise single-copy genes. Our results highlight the complexity of genomic diversity in some plant groups, which makes orthology assessment and accurate phylogenomic inference difficult.</span></p>

opencc-zeroJan 2021View details →
zenodo32/100

Genome alignments from "Split-alignment of genomes finds orthologies more accurately"

<p>Here are the genome alignments resulting from &quot;Split-alignment of genomes finds orthologies more accurately&quot;, Genome Biology 2015, 16:106.</p> <ul> <li>In the terminology of that paper, these are 2-split, post-masked alignments.</li> <li>They are in MAF format (http://genome.ucsc.edu/FAQ/FAQformat.html), with &quot;p&quot; lines (http://last.cbrc.jp/doc/last-split.html).</li> <li>Alignments with high ambiguity have not been removed, so if you want unambiguously 1-to-1 alignments, remove those annotated with mismap &gt; 0.00001 or so.</li> <li>The chimp and orangutan alignments were made using the 500/30 gap costs suggested at the end of the paper.</li> </ul> <p>&nbsp;</p>

opencc-zeroMay 2015View details →
zenodo32/100

Genome Alignment of Cancer Sequencing Data

<p>Part of the GTN Cancer Analysis learning Pathway based on the Bioinformatics.ca Cancer Workshop</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.

<p>List of 200 most abundant VOG descriptions.</p>

opencc-by-4.0Jul 2023View details →
dryad32/100

Alignments from: Gene count from target sequence capture places three whole genome duplication events in Hibiscus L. (Malvaceae)

Open the record for dataset details and reuse information.

publicJan 2021View details →
dryad32/100

Genome reduction is associated with bacterial pathogenicity across different scales of temporal and ecological divergence - between species core gene alignments

Open the record for dataset details and reuse information.

publicNov 2020View details →
dryad32/100

Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)

Open the record for dataset details and reuse information.

publicJun 2020View details →
dryad32/100

Nucleotide alignments of eight meiosis genes under extreme selection following whole genome duplication in Arabidopsis lyrata/A.arenosa.

Open the record for dataset details and reuse information.

publicJun 2020View details →
zenodo28/100

Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods

<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+&Gamma;. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+&Gamma;: six rates and one &Gamma; shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>&nbsp;</p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p>&nbsp;</p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1]&nbsp; &nbsp;</code>&nbsp;&nbsp; integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5]&nbsp;</code> &nbsp; frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10]&nbsp;</code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] &nbsp; </code> &nbsp; &Gamma; shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without &Gamma;) specified to <em>INDELible</em>,</li> <li><code>[12] &nbsp; </code> &nbsp; length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] &nbsp; </code> &nbsp; length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] &nbsp; </code> &nbsp; no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] &nbsp; </code> &nbsp; no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] &nbsp; </code> &nbsp; observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] &nbsp; </code> &nbsp; aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] &nbsp; </code> &nbsp; aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

Alignment used for the phylogenies of "The earliest diverging extant scleractinian corals recovered by mitochondrial genomes"

<p>Alignment of scleractinian corals based on mitochondrial genomes and including species of the family Micrabaciidae. This alignment was used for running Maximum Likelihood and Bayesian Inference phylogenies. &quot;tree02_alignment.phy&quot; refers to the actual alignment while &quot;tree02_alignment.partitions.txt&quot; indicates where each gene partition begins/ends.</p>

opencc-by-4.0Dec 2020View details →
dryad28/100

Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment

Reduced-representation genome sequencing such as RADseq aids the analysis of genomes by reducing the quantity of data, thereby lowering both sequencing costs and computational burdens. RADseq was initially designed for studying genetic variation across genomes at the population level, but has also proved to be suitable for interspecific phylogeny reconstruction. RADseq data pose challenges for standard phylogenomic methods, however, due to incomplete coverage of the genome and large amounts of missing data. Alignment-free methods are both efficient and accurate for phylogenetic reconstructions with whole genomes and are especially practical for non-model organisms; nonetheless, alignment-free methods have not been applied with reduced genome sequencing data. Here, we test a full-genome assembly and alignment-free method, AAF, in application to RADseq data and propose two procedures for reads selection to remove reads from restriction sites that were not found in taxa being compared. We validate these methods using both simulations and real datasets. Reads selection improved the accuracy of phylogenetic construction in every simulated scenario and the two real datasets, making AAF as good or better than a comparable alignment-based method, even though AAF had much lower computational burdens. We also investigated the sources of missing data in RADseq and their effects on phylogeny reconstruction using AAF. The AAF pipeline modified for RADseq or other reduced-representation sequencing data, phyloRAD, is available on github (https://github.com/fanhuan/phyloRAD).

opencc-zeroDec 2017View details →
dryad28/100

MtDNA genomes from Ranis individuals aligned with previously published ancient & modern humans

<p>The Middle to Upper Palaeolithic transition in Europe is associated with the regional disappearance of Neanderthals and the spread of Homo sapiens. Archaeological evidence indicates the presence of technocomplexes at the interface of this transition, complicating our understanding of the period and the association of those with specific hominin groups. One such technocomplex where the maker is unknown is the Lincombian-Ranisian-Jerzmanowician (LRJ), which covers an area in northwestern and central Europe from the UK to Poland. This paper presents the morphological and proteomic species identification, mitochondrial DNA analysis, and direct radiocarbon dating of human remains directly associated to an LRJ assemblage at the cave site of Ilsenhöhle in Ranis (Germany).  </p>

opencc-zeroOct 2023View details →
zenodo28/100

Multiple whole genome alignment of 63 Nymphalidae (HAL file, alignment version 1.0), plus protein-coding annotations for each of the species in the hall file.

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo28/100

Transcriptome profiling of derived-hepatocyte progenitors from human iPSCs with nanoCAGE - part1 - genomic alignments (hg19 + hg38)

<p>This repository contains genomic alignments (BED files) of paired-end nanoCAGE sequencing data (CAGEscan data) collected from Illumina MiSeq run IDs &quot;170630_M00528_0292_000000000-B9JY8&quot; (aka &quot;NC_LIMMS&quot;) and &quot;180221_M00528_0334_000000000-B6PJM&quot; (aka &quot;NC_LIMMS2&quot;). FASTQ files were processed with the MOIRAI pipeline OP-WORKFLOW-CAGEscan-short-reads-v2.1 (Hasegawa et al. BMC Bioinformatics&nbsp;2014 May 16;15:144. doi: 10.1186/1471-2105-15-144.). Filtered pairs of reads were aligned on the human genome assemblies hg19 and hg38. See tables below for a detailed description of the samples contained in each nanoCAGE library, including barcodes and index sequences used for the demultiplexing of sequencing reads. Corresponding raw sequencing data files (FASTQ files) were deposited at Zenodo under&nbsp;the following Digital Object Identifier: 10.5281/zenodo.1014009.</p> <p>&nbsp;</p> <p><em><strong>&quot;170630_M00528_0292_000000000-B9JY8&quot; (&quot;NC_LIMMS&quot;) :</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>1&nbsp;&nbsp; iPSC_control_rep1&nbsp;&nbsp; 4&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>2&nbsp;&nbsp; iPSC_control_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>3&nbsp;&nbsp; iPSC_control_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>4&nbsp;&nbsp; S3P1_OK_rep1&nbsp;&nbsp; 36&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>5&nbsp;&nbsp; S3P1_OK_rep2&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>6&nbsp;&nbsp; S3P1_OK_rep3&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>7&nbsp;&nbsp; S4P1_OK_rep1&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>8&nbsp;&nbsp; S4P1_OK_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>9&nbsp;&nbsp; S4P1_OK_rep3&nbsp;&nbsp; 9&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>10&nbsp;&nbsp; S4P2_OK_rep1&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>11&nbsp;&nbsp; S4P2_OK_rep2&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>12&nbsp;&nbsp; S4P2_OK_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>13&nbsp;&nbsp; S1P1_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>14&nbsp;&nbsp; S1P1_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>15&nbsp;&nbsp; S1P1_rep3&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>16&nbsp;&nbsp; S3P1_FAILED_rep1&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>17&nbsp;&nbsp; S3P1_FAILED_rep2&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>18&nbsp;&nbsp; S3P1_FAILED_rep3&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>19&nbsp;&nbsp; S4P1_FAILED_rep1&nbsp;&nbsp; 35&nbsp;&nbsp; CACTCT&nbsp;&nbsp; NNNNNNNN</p> <p>20&nbsp;&nbsp; S4P1_FAILED_rep2&nbsp;&nbsp; 47&nbsp;&nbsp; CTGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>21&nbsp;&nbsp; S4P1_FAILED_rep3&nbsp;&nbsp; 59&nbsp;&nbsp; GAGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>22&nbsp;&nbsp; S4P2_FAILED_rep1&nbsp;&nbsp; 71&nbsp;&nbsp; GCTGCA&nbsp;&nbsp; NNNNNNNN</p> <p>23&nbsp;&nbsp; S4P2_FAILED_rep2&nbsp;&nbsp; 83&nbsp;&nbsp; TATAGC&nbsp;&nbsp; NNNNNNNN</p> <p>24&nbsp;&nbsp; S4P2_FAILED_rep3&nbsp;&nbsp; 95&nbsp;&nbsp; TCGCGT&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p> <p><em><strong>&quot;180221_M00528_0334_000000000-B6PJM&quot; (&quot;NC_LIMMS2&quot;):</strong></em></p> <p><strong>ID&nbsp;&nbsp; Sample_name&nbsp;&nbsp; Barcode_number&nbsp;&nbsp; Barcode_sequence &nbsp; Index_sequence</strong></p> <p>25&nbsp;&nbsp; PETRI_rep1&nbsp;&nbsp; 04&nbsp;&nbsp; ACAGAT&nbsp;&nbsp; NNNNNNNN</p> <p>26&nbsp;&nbsp; PETRI_rep2&nbsp;&nbsp; 24&nbsp;&nbsp; ATCGTG&nbsp;&nbsp; NNNNNNNN</p> <p>27&nbsp;&nbsp; PETRI_rep3&nbsp;&nbsp; 31&nbsp;&nbsp; CACGAT&nbsp;&nbsp; NNNNNNNN</p> <p>28&nbsp;&nbsp; BIOCHIP_E_rep1&nbsp;&nbsp; 6&nbsp;&nbsp; CACTGA&nbsp;&nbsp; NNNNNNNN</p> <p>29&nbsp;&nbsp; BIOCHIP_M_rep1&nbsp;&nbsp; 46&nbsp;&nbsp; CTGACG&nbsp;&nbsp; NNNNNNNN</p> <p>30&nbsp;&nbsp; BIOCHIP_S_rep1&nbsp;&nbsp; 63&nbsp;&nbsp; GAGTGA&nbsp;&nbsp; NNNNNNNN</p> <p>31&nbsp;&nbsp; BIOCHIP_E_rep2&nbsp;&nbsp; 79&nbsp;&nbsp; GTATAC&nbsp;&nbsp; NNNNNNNN</p> <p>32&nbsp;&nbsp; BIOCHIP_M_rep2&nbsp;&nbsp; 92&nbsp;&nbsp; TCGAGC&nbsp;&nbsp; NNNNNNNN</p> <p>33&nbsp;&nbsp; BIOCHIP_S_rep2&nbsp;&nbsp; 09&nbsp;&nbsp; ACATGA&nbsp;&nbsp; NNNNNNNN</p> <p>34&nbsp;&nbsp; BIOCHIP_E_rep3&nbsp;&nbsp; 21&nbsp;&nbsp; ATCATA&nbsp;&nbsp; NNNNNNNN</p> <p>35&nbsp;&nbsp; BIOCHIP_M_rep3&nbsp;&nbsp; 33&nbsp;&nbsp; CACGTG&nbsp;&nbsp; NNNNNNNN</p> <p>36&nbsp;&nbsp; BIOCHIP_S_rep3&nbsp;&nbsp; 45&nbsp;&nbsp; CGATGA&nbsp;&nbsp; NNNNNNNN</p> <p>37&nbsp;&nbsp; HEPATOCYTES_rep1&nbsp;&nbsp; 57&nbsp;&nbsp; GAGATA&nbsp;&nbsp; NNNNNNNN</p> <p>38&nbsp;&nbsp; HEPATOCYTES_rep2&nbsp;&nbsp; 69&nbsp;&nbsp; GCTCTC&nbsp;&nbsp; NNNNNNNN</p> <p>39&nbsp;&nbsp; iPSC_control_rep1-2&nbsp;&nbsp; 81&nbsp;&nbsp; GTATGA&nbsp;&nbsp; NNNNNNNN</p> <p>40&nbsp;&nbsp; BIOCHIP_E_rep2-2&nbsp;&nbsp; 93&nbsp;&nbsp; TCGATA&nbsp;&nbsp; NNNNNNNN</p> <p>41&nbsp;&nbsp; BIOCHIP_M_rep1-2&nbsp;&nbsp;&nbsp; 11&nbsp;&nbsp; AGTAGC&nbsp;&nbsp; NNNNNNNN</p> <p>42&nbsp;&nbsp; BIOCHIP_S_rep2-2&nbsp;&nbsp; 23&nbsp;&nbsp; ATCGCA&nbsp;&nbsp; NNNNNNNN</p> <p>&nbsp;</p>

openOct 2017View details →
zenodo28/100

genome alignment of 88 speceis

<p>genome alignment of 88 speceis</p>

opencc-by-4.0Mar 2023View details →
dryad28/100

VCF of structural variant calls of Nanopore data aligned to dm6 reference genome

<p>Heterozygous chromosome inversions suppress meiotic crossover (CO) formation within an inversion, potentially because they lead to gross chromosome rearrangements that produce inviable gametes. Interestingly, COs are also severely reduced in regions nearby but outside of inversion breakpoints even though COs in these regions do not result in rearrangements. Our mechanistic understanding of why COs are suppressed outside of inversion breakpoints is limited by a lack of data on the frequency of noncrossover gene conversions (NCOGCs) in these regions. To address this critical gap, we mapped the location and frequency of rare CO and NCOGC events that occurred outside of the <em>dl</em>-<em>49</em> <em>chrX</em> inversion in <em>D</em>. <em>melanogaster</em>. We created full-sibling wildtype and inversion stocks and recovered COs and NCOGCs in the syntenic regions of both stocks, allowing us to directly compare rates and distributions of recombination events. We show that COs are completely suppressed within 500 kb of inversion breakpoints, are severely reduced within 2 Mb of an inversion breakpoint, and increase above wildtype levels 2–4 Mb from the breakpoint. We find that NCOGCs occur evenly throughout the chromosome and, importantly, occur at wild-type levels near inversion breakpoints. We propose a model in which COs are suppressed by inversion breakpoints in a distance-dependent manner through mechanisms that influence DNA double-strand break repair outcome but not double-strand break location or frequency. We suggest that subtle changes in the synaptonemal complex and chromosome pairing might lead to unstable interhomolog interactions during recombination that permits NCOGC formation but not CO formation.</p>

opencc-zeroMar 2023View details →
dryad28/100

MtDNA genomes from Ranis individuals aligned with previously published ancient & modern humans

Open the record for dataset details and reuse information.

publicOct 2023View details →
dryad28/100

Data from: Reconstructing phylogeny from reduced-representation genome sequencing data without assembly or alignment

Open the record for dataset details and reuse information.

publicMay 2018View details →
dryad28/100

VCF of structural variant calls of Nanopore data aligned to dm6 reference genome

Open the record for dataset details and reuse information.

publicMar 2023View details →
geo24/100

RNA-seq alignment to individualized genomes

GEO Series GSE45684. Mus musculus. 1086 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record