Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

68

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

68 results for “De novo genome assembly”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: "De novo assembly transcriptome for the rostrum dace (Leuciscus burdigalensis, Cyprinidae: fish) naturally infected by a copepod ectoparasite" in Genomic Resources Notes accepted 1 December 2014 to 31 January 2015

Open the record for dataset details and reuse information.

publicFeb 2015View details →
dryad32/100

Data from: "De novo assembled transcriptome of organs involved in reproduction in an endangered endemic Iberian cyprinid fish (Squalius pyrenaicus)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015

Open the record for dataset details and reuse information.

publicAug 2015View details →
dryad32/100

De novo genome assembly of Leptodactylus fuscus

Open the record for dataset details and reuse information.

publicApr 2021View details →
dryad32/100

De novo genome assembly for Eulemur rufifrons

Open the record for dataset details and reuse information.

publicAug 2024View details →
dryad32/100

De novo genome assembly of Tectona grandis (Teak) with 2993 scaffolds

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad32/100

Data from: Two low coverage bird genomes and a comparison of reference-guided versus de novo genome assemblies

Open the record for dataset details and reuse information.

publicAug 2015View details →
dryad32/100

Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies

Open the record for dataset details and reuse information.

publicOct 2013View details →
dryad32/100

Data from: The de novo genome assembly and annotation of a female domestic dromedary of North African origin

Open the record for dataset details and reuse information.

publicJul 2015View details →
dryad28/100

Data from: De novo genome assembly and annotation of rice sheath rot fungus Sarocladium oryzae reveals genes involved in Helvolic acid and Cerulenin biosynthesis pathways

Background: Sheath rot disease caused by Sarocladium oryzae is an emerging threat for rice cultivation at global level. However, limited information with respect to genomic resources and pathogenesis is a major setback to develop disease management strategies. Considering this fact, we sequenced the whole genome of highly virulent Sarocladium oryzae field isolate, Saro-13 with 82x sequence depth. Results: The genome size of S. oryzae was 32.78 Mb with contig N50 18.07 Kb and 10526 protein coding genes. The functional annotation of protein coding genes revealed that S. oryzae genome has evolved with many expanded gene families of major super family, proteinases, zinc finger proteins, sugar transporters, dehydrogenases/reductases, cytochrome P450, WD domain G-beta repeat and FAD-binding proteins. Gene orthology analysis showed that around 79.80 % of S. oryzae genes were orthologous to other Ascomycetes fungi. The polyketide synthase dehydratase, ATP-binding cassette (ABC) transporters, amine oxidases, and aldehyde dehydrogenase family proteins were duplicated in larger proportion specifying the adaptive gene duplications to varying environmental conditions. Thirty-nine secondary metabolite gene clusters encoded for polyketide synthases, nonribosomal peptide synthase, and terpene cyclases. Protein homology based analysis indicated that nine putative candidate genes were found to be involved in helvolic acid biosynthesis pathway. The genes were arranged in cluster and structural organization of gene cluster was similar to helvolic acid biosynthesis cluster in Metarhizium anisophilae. Around 9.37 % of S. oryzae genes were identified as pathogenicity genes, which are experimentally proven in other phytopathogenic fungi and enlisted in pathogen-host interaction database. In addition, we also report 13212 simple sequences repeats (SSRs) which can be deployed in pathogen identification and population dynamic studies in near future. Conclusions: Large set of pathogenicity determinants and putative genes involved in helvolic acid and cerulenin biosynthesis will have broader implications with respect to Sarocladium disease biology. This is the first genome sequencing report globally and the genomic resources developed from this study will have wider impact worldwide to understand Rice-Sarocladium interaction.

opencc-zeroDec 2015View details →
dryad28/100

Data from: De novo genome assembly of Geosmithia morbida, the causal agent of thousand cankers disease

Geosmithia morbida is a filamentous ascomycete that causes thousand cankers disease in the eastern black walnut tree. This pathogen is commonly found in the western U.S.; however, recently the disease was also detected in several eastern states where the black walnut lumber industry is concentrated. G. morbida is one of two known phytopathogens within the genus Geosmithia, and it is vectored into the host tree via the walnut twig beetle. We present the first de novo draft genome of G. morbida. It is 26.5 Mbp in length and contains less than 1% repetitive elements. The genome possesses an estimated 6,273 genes, 277 of which are predicted to encode proteins with unknown functions. Approximately 31.5% of the proteins in G. morbida are homologous to proteins involved in pathogenicity, and 5.6% of the proteins contain signal peptides that indicate these proteins are secreted. Several studies have investigated the evolution of pathogenicity in pathogens of agricultural crops; forest fungal pathogens are often neglected because research efforts are focused on food crops. G. morbida is one of the few tree phytopathogens to be sequenced, assembled and annotated. The first draft genome of G. morbida serves as a valuable tool for comprehending the underlying molecular and evolutionary mechanisms behind pathogenesis within the Geosmithia genus.

opencc-zeroDec 2015View details →
dryad28/100

Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C

The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.

opencc-zeroJun 2019View details →
zenodo28/100

Assemblies generated in the paper "phasebook: haplotype-aware de novo assembly of diploid genomes from long reads"

<p>Assemblies generated in the paper &quot;phasebook: haplotype-aware de novo assembly of diploid genomes from long reads&quot;</p>

opencc-by-4.0Sep 2021View details →
dryad28/100

A chromosome-scale de novo genome assembly of the dwarf tomato variety Micro-Tom

<p>The cultivated tomato (<em>Solanum lycopersicum</em>) is an important crop and model species for genetics and plant molecular biology research. The dwarf tomato variety Micro-Tom is used extensively in research because it is rapid flowering, easy to grow in high volumes in minimal space, and is amenable to genetic transformation. Here we provide a de novo chromosome-scale genome assembly of Micro-Tom that was generated using PacBio HiFi reads and scaffolded using chromosome confirmation capture data. The HiFi data was assembled using the Hifiasm assembler and OmniC data was used for scaffolding using Salsa and several rounds of manual curation and validation.</p>

opencc-zeroOct 2023View details →
dryad28/100

Data from: De novo sequencing, assembly, and annotation of four threespine stickleback genomes based on microfluidic partitioned DNA libraries

Open the record for dataset details and reuse information.

publicJun 2019View details →
dryad28/100

Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C

Open the record for dataset details and reuse information.

publicJun 2019View details →
dryad28/100

Data from: De novo genome assembly and annotation of rice sheath rot fungus Sarocladium oryzae reveals genes involved in Helvolic acid and Cerulenin biosynthesis pathways

Open the record for dataset details and reuse information.

publicMar 2017View details →
dryad28/100

Data from: De novo genome assembly of Geosmithia morbida, the causal agent of thousand cankers disease

Open the record for dataset details and reuse information.

publicApr 2017View details →
dryad28/100

A chromosome-scale de novo genome assembly of the dwarf tomato variety Micro-Tom

Open the record for dataset details and reuse information.

publicOct 2023View details →
dryad28/100

Data from: De novo assembly of genomes from long sequence reads reveals uncharted territories of Propionibacterium freudenreichii

Open the record for dataset details and reuse information.

publicOct 2018View details →
geo24/100

De Novo Genome Assembly of Guinea Grass Exposed to Elevated CO2 and Temperature

GEO Series GSE122194. Megathyrsus maximus. 33 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record