Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

124

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

124 results for “de novo genome”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: The de novo genome assembly and annotation of a female domestic dromedary of North African origin

Open the record for dataset details and reuse information.

publicJul 2015View details →
dryad32/100

Data from: Genome-wide prediction models that incorporate de novo GWAS are a powerful new tool for tropical rice improvement

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad28/100

Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.

Next-generation sequencing (NGS) technologies are increasingly applied in many organisms, including non-model organisms that are important for ecological and conservation purposes. Illumina and 454 sequencing are among the most used NGS technologies and have been shown to produce optimal results at reasonable costs when used together. Here, we describe the combined application of these two NGS technologies to characterize the transcriptome of a plant species of ecological and conservation relevance for which no genomic resource is available, Scabiosa columbaria. We obtained 528,557 reads from a 454 GS-FLX run and a total of 28,993,627 reads from two lanes of an Illumina GAII single run. After reads trimming, the de novo assembly of both types of reads produced 109,630 contigs. Both the contigs and the >75 bp remaining singletons were blasted against Uniprot/Swissprot database, resulting in 29,676 and 10,515 significant hits, respectively. Based on sequence similarity with known gene products, these sequences represent at least 12,516 unique genes, most of which are well covered by contig sequences. In addition, we identified 4,320 microsatellite loci, of which 856 had flanking sequences suitable for PCR primer design. We also identified 75,054 putative SNPs. This annotated sequence collection and the relative molecular markers represent a main genomic resource for S. columbaria which should contribute to future research in conservation and population biology studies. Our results demonstrate the utility of NGS technologies as starting point for the development of genomic tools in nonmodel but ecologically important species.

opencc-zeroDec 2009View details →
dryad28/100

Data from: De novo genome assembly and annotation of rice sheath rot fungus Sarocladium oryzae reveals genes involved in Helvolic acid and Cerulenin biosynthesis pathways

Background: Sheath rot disease caused by Sarocladium oryzae is an emerging threat for rice cultivation at global level. However, limited information with respect to genomic resources and pathogenesis is a major setback to develop disease management strategies. Considering this fact, we sequenced the whole genome of highly virulent Sarocladium oryzae field isolate, Saro-13 with 82x sequence depth. Results: The genome size of S. oryzae was 32.78 Mb with contig N50 18.07 Kb and 10526 protein coding genes. The functional annotation of protein coding genes revealed that S. oryzae genome has evolved with many expanded gene families of major super family, proteinases, zinc finger proteins, sugar transporters, dehydrogenases/reductases, cytochrome P450, WD domain G-beta repeat and FAD-binding proteins. Gene orthology analysis showed that around 79.80 % of S. oryzae genes were orthologous to other Ascomycetes fungi. The polyketide synthase dehydratase, ATP-binding cassette (ABC) transporters, amine oxidases, and aldehyde dehydrogenase family proteins were duplicated in larger proportion specifying the adaptive gene duplications to varying environmental conditions. Thirty-nine secondary metabolite gene clusters encoded for polyketide synthases, nonribosomal peptide synthase, and terpene cyclases. Protein homology based analysis indicated that nine putative candidate genes were found to be involved in helvolic acid biosynthesis pathway. The genes were arranged in cluster and structural organization of gene cluster was similar to helvolic acid biosynthesis cluster in Metarhizium anisophilae. Around 9.37 % of S. oryzae genes were identified as pathogenicity genes, which are experimentally proven in other phytopathogenic fungi and enlisted in pathogen-host interaction database. In addition, we also report 13212 simple sequences repeats (SSRs) which can be deployed in pathogen identification and population dynamic studies in near future. Conclusions: Large set of pathogenicity determinants and putative genes involved in helvolic acid and cerulenin biosynthesis will have broader implications with respect to Sarocladium disease biology. This is the first genome sequencing report globally and the genomic resources developed from this study will have wider impact worldwide to understand Rice-Sarocladium interaction.

opencc-zeroDec 2015View details →
dryad28/100

Data from: De novo genome assembly of Geosmithia morbida, the causal agent of thousand cankers disease

Geosmithia morbida is a filamentous ascomycete that causes thousand cankers disease in the eastern black walnut tree. This pathogen is commonly found in the western U.S.; however, recently the disease was also detected in several eastern states where the black walnut lumber industry is concentrated. G. morbida is one of two known phytopathogens within the genus Geosmithia, and it is vectored into the host tree via the walnut twig beetle. We present the first de novo draft genome of G. morbida. It is 26.5 Mbp in length and contains less than 1% repetitive elements. The genome possesses an estimated 6,273 genes, 277 of which are predicted to encode proteins with unknown functions. Approximately 31.5% of the proteins in G. morbida are homologous to proteins involved in pathogenicity, and 5.6% of the proteins contain signal peptides that indicate these proteins are secreted. Several studies have investigated the evolution of pathogenicity in pathogens of agricultural crops; forest fungal pathogens are often neglected because research efforts are focused on food crops. G. morbida is one of the few tree phytopathogens to be sequenced, assembled and annotated. The first draft genome of G. morbida serves as a valuable tool for comprehending the underlying molecular and evolutionary mechanisms behind pathogenesis within the Geosmithia genus.

opencc-zeroDec 2015View details →
dryad28/100

Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C

The red spotted grouper Epinephelus akaara (E. akaara) is one of the most economically important marine fish in China, Japan and Southeast Asia, and is a threatened species. The species is also considered a good model for studies of sex-inversion, development, genetic diversity and immunity. Despite its importance, molecular resources for E. akaara remain limited and no reference genome has been published to date. In this study, we constructed a chromosome-level reference genome of E. akaara by taking advantage of long-read single molecule sequencing and de novo assembly by Oxford Nanopore Technologies (ONT) and Hi-C. A red-spotted grouper genome of 1.135 Gb was assembled from a total of 106.29 Gb polished Nanopore sequence (GridION, ONT), equivalent to 96-fold genome coverage. The assembled genome represents 96.8% completeness (BUSCO) with a contig N50 length of 5.25 Mb and a longest contig of 25.75 Mb. The contigs were clustered and ordered onto 24 pseudo-chromosomes covering approximately 95.55% of the genome assembly with Hi-C data, with a scaffold N50 length of 46.03 Mb. The genome contained 43.02% repeat sequences and 5,480 non-coding RNAs. Furthermore, after mining several RNA-seq datasets, 23,809 (99.5%) genes were functionally annotated from a total of 23,924 predicted protein-coding sequences. The high-quality chromosome-level reference genome of E. akaara was assembled for the first time and will be a valuable resource for molecular breeding and functional genomics studies of red-spotted grouper in the future.

opencc-zeroJun 2019View details →
zenodo28/100

Assemblies generated in the paper "phasebook: haplotype-aware de novo assembly of diploid genomes from long reads"

<p>Assemblies generated in the paper &quot;phasebook: haplotype-aware de novo assembly of diploid genomes from long reads&quot;</p>

opencc-by-4.0Sep 2021View details →
dryad28/100

A chromosome-scale de novo genome assembly of the dwarf tomato variety Micro-Tom

<p>The cultivated tomato (<em>Solanum lycopersicum</em>) is an important crop and model species for genetics and plant molecular biology research. The dwarf tomato variety Micro-Tom is used extensively in research because it is rapid flowering, easy to grow in high volumes in minimal space, and is amenable to genetic transformation. Here we provide a de novo chromosome-scale genome assembly of Micro-Tom that was generated using PacBio HiFi reads and scaffolded using chromosome confirmation capture data. The HiFi data was assembled using the Hifiasm assembler and OmniC data was used for scaffolding using Salsa and several rounds of manual curation and validation.</p>

opencc-zeroOct 2023View details →
dryad28/100

Data from: De novo sequencing, assembly, and annotation of four threespine stickleback genomes based on microfluidic partitioned DNA libraries

Open the record for dataset details and reuse information.

publicJun 2019View details →
dryad28/100

Data from: De novo assembly of a chromosome-level reference genome of red spotted grouper (Epinephelus akaara) using nanopore sequencing and Hi-C

Open the record for dataset details and reuse information.

publicJun 2019View details →
dryad28/100

Data from: De novo genome assembly and annotation of rice sheath rot fungus Sarocladium oryzae reveals genes involved in Helvolic acid and Cerulenin biosynthesis pathways

Open the record for dataset details and reuse information.

publicMar 2017View details →
dryad28/100

Data from: De novo transcriptome characterization and development of genomic tools for Scabiosa columbaria L. using next-generation sequencing techniques.

Open the record for dataset details and reuse information.

publicDec 2010View details →
dryad28/100

Data from: De novo genome assembly of Geosmithia morbida, the causal agent of thousand cankers disease

Open the record for dataset details and reuse information.

publicApr 2017View details →
dryad28/100

A chromosome-scale de novo genome assembly of the dwarf tomato variety Micro-Tom

Open the record for dataset details and reuse information.

publicOct 2023View details →
dryad28/100

Data from: De novo assembly of genomes from long sequence reads reveals uncharted territories of Propionibacterium freudenreichii

Open the record for dataset details and reuse information.

publicOct 2018View details →
geo24/100

Integrated genomic analyses of de novo pathways underlying atypical meningiomas [ChIP-Seq]

GEO Series GSE91372. Homo sapiens. 40 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenDec 2016View details →
geo24/100

DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Drosophila genome-wide UMI-STARR-seq]

GEO Series GSE183936. Drosophila melanogaster; synthetic construct. 6 samples. Type: Other.

openGEO-OpenFeb 2022View details →
geo24/100

De novo methylaton of the unmethylated K. phaffii genome by human DNA methyltransferases (DNMTs)

GEO Series GSE139063. Komagataella phaffii. 65 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing.

openGEO-OpenFeb 2020View details →
geo24/100

Genomic analysis of hESC pedigrees enables identification of de novo genetic alterations and determination of the timing and origin of mutational events

GEO Series GSE49452. Homo sapiens. 21 samples. Type: SNP genotyping by SNP array; Genome variation profiling by SNP array.

openGEO-OpenApr 2014View details →
geo24/100

De novo mutations in the genome organizer CTCF cause Intellectual Disability

GEO Series GSE46833. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenJun 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record