Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

199

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

199 results for “genome reference”

Learn how ShareScore rates datasets ↗
dryad32/100

Data for: Three amphioxus reference genomes reveal gene and chromosome evolution of chordates

Open the record for dataset details and reuse information.

publicJan 2023View details →
dryad32/100

Data from: Genomics of Compositae crops: reference transcriptome assemblies, and evidence of hybridization with wild relatives

Open the record for dataset details and reuse information.

publicAug 2013View details →
dryad32/100

Data from: Geographic patterns of genetic variation in three genomes of North American diploid strawberries with special reference to Fragaria vesca subsp. bracteata

Open the record for dataset details and reuse information.

publicJul 2015View details →
dryad32/100

Construction of a chromosome-scale long-read reference genome assembly for potato

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad32/100

Data from: Two low coverage bird genomes and a comparison of reference-guided versus de novo genome assemblies

Open the record for dataset details and reuse information.

publicAug 2015View details →
dryad32/100

Data from: A genomic reference panel for Drosophila serrata

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad28/100

Data from: Leishmania naiffi and Leishmania guyanensis reference genomes highlight genome structure and gene evolution in the Viannia subgenus

The unicellular protozoan parasite Leishmania causes the neglected tropical disease leishmaniasis, affecting 12 million people in 98 countries. In South America where the Viannia subgenus predominates, so far only L. (Viannia) braziliensis and L. (V.) panamensis have been sequenced, assembled and annotated as reference genomes. Addressing this deficit in molecular information can inform species typing, epidemiological monitoring and clinical treatment. Here, L. (V.) naiffi and L. (V.) guyanensis genomic DNA was sequenced to assemble these two genomes as draft references from short sequence reads. The methods used were tested using short sequence reads for L. braziliensis M2904 against its published reference as a comparison. This assembly and annotation pipeline identified 70 additional genes not annotated on the original M2904 reference. Phylogenetic and evolutionary comparisons of L. guyanensis and L. naiffi with ten other Viannia genomes revealed four traits common to all Viannia: aneuploidy, 22 orthologous groups of genes absent in other Leishmania subgenera, elevated TATE transposon copies, and a high NADH-dependent fumarate reductase gene copy number. Within the Viannia, there were limited structural changes in genome architecture specific to individual species: a 45 Kb amplification on chromosome 34 was present in all bar L. lainsoni, L. naiffi had a higher copy number of the virulence factor leishmanolysin, and laboratory isolate L. shawi M8408 had a possible minichromosome derived from the 3' end of chromosome 34. This combination of genome assembly, phylogenetics and comparative analysis across an extended panel of diverse Viannia has uncovered new insights into the origin and evolution of this subgenus and can help improve diagnostics for leishmaniasis surveillance.

opencc-zeroDec 2017View details →
dryad28/100

Data from: GIbPSs: a toolkit for fast and accurate analyses of genotyping-by-sequencing data without a reference genome

Genotyping-by-sequencing (GBS) and related methods are increasingly used for studies of non-model organisms from population genetic to phylogenetic scales. We present GIbPSs, a new genotyping toolkit for the analysis of data from various protocols such as RAD, double-digest RAD, GBS, and two-enzyme GBS without a reference genome. GIbPSs can handle paired-end GBS data and is able to assign reads from both strands of a restriction fragment to the same locus. GIbPSs is most suitable for population genetic and phylogeographic analyses. It avoids genotyping errors due to indel variation by identifying and discarding affected loci. GIbPSs creates a genotype database that offers rich functionality for data filtering and export in numerous formats. We performed comparative analyses of simulated and real GBS data with GIbPSs and another program, pyRAD. This program accounts for indel variation by aligning homologous sequences. GIbPSs performed better than pyRAD in several aspects. It required much less computation time and displayed higher genotyping accuracy. GIbPSs retained smaller numbers of loci overall in analyses of real GBS data. It nevertheless delivered more complete genotype matrices with greater locus overlap between individuals and greater numbers of loci sampled in all individuals.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Linkage disequilibrium and inversion-typing of the Drosophila melanogaster Genome Reference Panel

We calculated the linkage disequilibrium between all pairs of variants in the Drosophila Genome Reference Panel with minor allele count ≥5. We used r2 ≥ 0.5 as the cutoff for a highly correlated SNP. We make available the list of all highly correlated SNPs for use in association studies. Seventy-six percent of variant SNPs are highly correlated with at least one other SNP, and the mean number of highly correlated SNPs per variant over the whole genome is 83.9. Disequilibrium between distant SNPs is also common when minor allele frequency (MAF) is low: 37% of SNPs with MAF < 0.1 are highly correlated with SNPs more than 100 kb distant. Although SNPs within regions with polymorphic inversions are highly correlated with somewhat larger numbers of SNPs, and these correlated SNPs are on average farther away, the probability that a SNP in such regions is highly correlated with at least one other SNP is very similar to SNPs outside inversions. Previous karyotyping of the DGRP lines has been inconsistent, and we used LD and genotype to investigate these discrepancies. When previous studies agreed on inversion karyotype, our analysis was almost perfectly concordant with those assignments. In discordant cases, and for inversion heterozygotes, our results suggest errors in two previous analyses or discordance between genotype and karyotype. Heterozygosities of chromosome arms are, in many cases, surprisingly highly correlated, suggesting strong epsistatic selection during the inbreeding and maintenance of the DGRP lines.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Chromosomal level reference genome of Tachypleus tridentatus provides insights into evolution and adaptation of horseshoe crabs

Horseshoe crabs including Tachypleus tridentatus are a group of marine arthropods living fossil species which have existed on the earth for 500 million years. However, the genetic mechanisms underlying their unique adaptive ability are still unclear. Here, we assembled the first chromosome-level T. tridentatus genome, and proofed that this genome is of very high quality with contig N50 1.69 Mb. By comparison with other arthropods, some gene families of T. tridentatus experienced significant expansion, whichare related to several signaling pathways, endonuclease activity, and metabolic processes. Based on the comparative analysis of genomics and 27 transcriptomes from 9 tissues, we found that the expanded Dscam genes usually locate at the key hub positions of immune network. Furthermore, the Dscam genes showed higher levels of expression in the yellow connective tissue, the birthplace of blood cells with strong differentiation capability, than the other 8 tissues. Besides, Dscam genes are positively correlated with the expression of the core immunity gene, clotting factor B, which is implicated in the coagulation cascade reaction. The effective and unusual immune ability endowed by the expansion and expression of Dscam genes in horseshoe crabs may be a factor that makes horseshoe crabs having a strong environmental adaptability with ~500 million years. The high-quality chromosome-level genome of a horseshoe crab and unique genomic features reported in this study provide important data resources for future studies on the evolutionary history of marine ecological systems.

opencc-zeroDec 2018View details →
dryad28/100

Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C

The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.

opencc-zeroJun 2019View details →
zenodo28/100

NSG-adapted reference genome - TRACERx PDX study

<p>NSG-adapted mouse reference genome and scripts to reproduce the results from whole-genome sequencing data derived from the tail of a male NOD SCID gamma (NSG) mouse (Mus musculus). The mouse was purchased from Charles River and had carried a subcutaneous human non-small cell lung cancer xenograft.</p> <p>The raw whole-genome sequencing data are available on ENA under accession number PRJEB65917 (https://www.ebi.ac.uk/ena/browser/view/PRJEB65917).</p> <p>The original manuscript describing this sample is available here.</p>

opencc-by-nc-4.0Dec 2023View details →
zenodo28/100

Phased T2T reference genome and pangenome reveal expanded resistance gene analogs in apple domestication

<p>Two haplotypes of the T2T genome of Golden Delicious apple and their annotation files</p>

opencc-by-4.0Mar 2024View details →
zenodo28/100

High-quality, chromosome-level reference genomes of the viviparous Caribbean skinks Spondylurus nitidus and S. culebrae

<p>Output files from the assembly of 2 reference genomes detailed in the publication Rivera et al. 2024. High-quality, chromosome-level reference genomes of the viviparous Caribbean skinks&nbsp;<em>Spondylurus nitidus</em>&nbsp;and&nbsp;<em>S. culebrae. Genome Biology and Evolution</em>, evae079.</p>

opencc-by-4.0Apr 2024View details →
zenodo28/100

Three reference genomes for freshwater diatom ecology and evolution

<p>This repository contains the genome assemblies and gene models provided for "Three reference genomes for freshwater diatom ecology and evolution"</p> <p>Authors:</p> <p>Wade R. Roberts (email: wader [at] uark [dot] edu)</p> <p>Andrew J. Alverson (email: aja [at] uark [dot] edu)</p> <p>&nbsp;</p> <p>The Whole Genome Shotgun (WGS) projects are available from NCBI GenBank under accession JALLPB020000000 (C. tholiformis), JALLBG020000000 (D. pseudostelligera), and JALLAZ020000000 (P. triporus).<strong></strong></p> <p><br>The following files are included:</p> <p>Cyclostephanos tholiformis strain AJA228-03</p> <p>&nbsp; &nbsp; aja228-03.consensus.fasta</p> <p>&nbsp; &nbsp; aja228-03.consensus.gff3<br>&nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; aja228-03.consensus.proteins.fasta</p> <p>&nbsp; &nbsp; aja228-03.consensus.combined_uniprot_annotation.csv</p> <p>&nbsp; &nbsp; aja228-03.consensus.panther_annotation.csv</p> <p>&nbsp; &nbsp; aja228-03.consensus.pfam_annotation.csv</p> <p>&nbsp;</p> <p>Discostella pseudostelligera strain AJA232-27</p> <p>&nbsp; &nbsp; aja232-27.consensus.fasta</p> <p>&nbsp; &nbsp; aja232-27.consensus.gff3<br>&nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; aja232-27.consensus.proteins.fasta</p> <p>&nbsp; &nbsp; aja232-27.consensus.combined_uniprot_annotation.csv</p> <p>&nbsp; &nbsp; aja232-27.consensus.panther_annotation.csv</p> <p>&nbsp; &nbsp; aja232-27.consensus.pfam_annotation.csv</p> <p>&nbsp;</p> <p>Praestephanos triporus strain AJA276-08</p> <p>&nbsp; &nbsp; aja276-08.consensus.fasta</p> <p>&nbsp; &nbsp; aja276-08.consensus.gff3<br>&nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; aja276-08.consensus.proteins.fasta</p> <p>&nbsp; &nbsp; aja276-08.consensus.combined_uniprot_annotation.csv</p> <p>&nbsp; &nbsp; aja276-08.consensus.panther_annotation.csv</p> <p>&nbsp; &nbsp; aja276-08.consensus.pfam_annotation.csv</p> <p>&nbsp;</p> <p>R code and phylogenetic tree to reproduce Figure 1 in the manuscript</p> <p>&nbsp; &nbsp; plot-figure-1.R</p> <p>&nbsp; &nbsp; busc.prot.concat.partition.rooted.tree</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo28/100

Genomic reference data for calculating conditional & conjuncational false discovery rates with cfdr.pleio

<p><strong>Background</strong></p> <p>This data set provides genomic reference information extracted from the 1000 Genomes Project for calculating conditional and conjunctional false discovery rates for genetic variants that are pleiotropic for pairs of phenotypes.</p> <ol> <li>genref.zip is preprocessed information that can be directly used with the R package cfdr.pleio, available from github<br> &nbsp;</li> <li>genref_rawdata.zip contains raw data downloaded from the 1000 Genome projects. You only need to download this if you want to re-create or modify the preprocessed reference data from above. The original code for creating the preprocessed data is available from github.</li> </ol> <p><strong>Usage </strong></p> <p>See the instructions at the two named github repositories for details.</p> <p><strong>Credit</strong></p> <p>The code for generating the preprocessed reference data as well as some auxiliary information were forked from the MATLAB package <a href="https://github.com/precimed/pleiofdr">pleiofdr</a></p> <p>The original concept of the conditional and conjunctional FDR were described in <a href="https://doi.org/10.1371/journal.pgen.1003455">Andreassen et al. (2013)</a>. For a recent overview of methods and applications, see <a href="https://doi.org/10.1007/s00439-019-02060-2">Smeland et al. (2019).</a></p> <p>&nbsp;</p>

openapgl-v3Dec 2021View details →
dryad28/100

Data from: Reference genome of Lumpfish Cyclopterus lumpus Linnaeus provides evidence of male heterogametic sex determination through the AMH pathway

<p>Teleosts exhibit extensive diversity of sex determination (SD) systems and mechanisms, providing the opportunity to study the evolution of sex determination and sex chromosomes. Here we sequenced the genome of the Common Lumpfish (<i>Cyclopterus lumpus</i> Linnaeus), a species of increasing importance to aquaculture, and identified the SD region and master SD locus using a 70K SNP array and tissue-specific expression data. The chromosome-level assembly identified 25 diploid chromosomes with a total size of 572.89 Mb, a scaffold N50 of 23.86 Mb, and genome annotation predicted 21,480 protein-coding genes. Genome wide association analysis located a highly sex-associated region on chromosome 13, suggesting that anti-Müllerian hormone (AMH) is the putative SD factor. Linkage disequilibrium and heterozygosity across chromosome 13 support a proto-XX/XY system, with an absence of widespread chromosome divergence between sexes. We identified three copies of <i>AMH</i> in the Lumpfish primary and alternate haplotype assemblies localized in the SD region. Comparison to sequences from other teleosts suggested a monophyletic relationship and conservation within the Cottioidei. One <i>AMH</i> copy showed similarity to <i>AMH/AMHY</i> in a related species and was also the only copy with expression in testis tissue, suggesting this copy may be the functional copy of <i>AMH</i> in Lumpfish. The two other copies arranged in tandem inverted duplication were highly similar, suggesting a recent duplication event. This study provides a resource for the study of early sex chromosome evolution and novel genomic resources that benefits Lumpfish conservation management and aquaculture.</p>

opencc-zeroDec 2021View details →
zenodo28/100

Results of mining Gordonia reference genomes by the antiSMASH tool

<p>Data table.</p>

opencc-by-4.0Aug 2022View details →
zenodo28/100

Reference genome

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad28/100

VCF of structural variant calls of Nanopore data aligned to dm6 reference genome

<p>Heterozygous chromosome inversions suppress meiotic crossover (CO) formation within an inversion, potentially because they lead to gross chromosome rearrangements that produce inviable gametes. Interestingly, COs are also severely reduced in regions nearby but outside of inversion breakpoints even though COs in these regions do not result in rearrangements. Our mechanistic understanding of why COs are suppressed outside of inversion breakpoints is limited by a lack of data on the frequency of noncrossover gene conversions (NCOGCs) in these regions. To address this critical gap, we mapped the location and frequency of rare CO and NCOGC events that occurred outside of the <em>dl</em>-<em>49</em> <em>chrX</em> inversion in <em>D</em>. <em>melanogaster</em>. We created full-sibling wildtype and inversion stocks and recovered COs and NCOGCs in the syntenic regions of both stocks, allowing us to directly compare rates and distributions of recombination events. We show that COs are completely suppressed within 500 kb of inversion breakpoints, are severely reduced within 2 Mb of an inversion breakpoint, and increase above wildtype levels 2–4 Mb from the breakpoint. We find that NCOGCs occur evenly throughout the chromosome and, importantly, occur at wild-type levels near inversion breakpoints. We propose a model in which COs are suppressed by inversion breakpoints in a distance-dependent manner through mechanisms that influence DNA double-strand break repair outcome but not double-strand break location or frequency. We suggest that subtle changes in the synaptonemal complex and chromosome pairing might lead to unstable interhomolog interactions during recombination that permits NCOGC formation but not CO formation.</p>

opencc-zeroMar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record