Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

89

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

89 results for “draft genome”

Learn how ShareScore rates datasets ↗
zenodo32/100

Draft Genome Sequences of Aquitalea sp strain MWU-14.2238

<p>Annotated genome of&nbsp;Aquitalea sp MWU-14.2238.&nbsp;</p>

opencc-by-3.0-usMar 2022View details →
zenodo32/100

Draft Genome Sequence of Pseudomonas sp. strain MWU13.2105, isolated from Wild Cranberry Bog in Truro, Massachusetts

<p>Annotated genome of Pseudomonas sp. strain MWU13.2105.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Draft Genome Sequence of Pseudomonas sp. strain MWU13.2100, isolated from Wild Cranberry Bog in Truro, Massachusetts

<p>Annotated genome&nbsp;of Pseudomonas sp. strain MWU13.2100.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Draft Genome Manuscript for Curtobacterium sp. Isolated from Berries Surfaced in Commercial Cranberry Bogs in Massachusetts, USA

<p>Annotated genome of&nbsp;Curtobacterium sp. MWU13.2055</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Draft Genome Manuscript for Pseudomonas sp. Strain MWU13.3659 Isolated from Berries Surfaced in Commercial Cranberry Bogs in Massachusetts, USA

<p>Annotated genome of Pseudomonas sp. MWU13.3659</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

Draft genome sequences of Arabidopsis thaliana-associated micro-organisms from Reijerscamp soil, the Netherlands

<p><strong>Methodological summary and relevant references</strong></p> <p>Compressed tar archive containing 447 draft bacterial genomes and their annotations used in several studies including Fourie&nbsp;<em>et al</em>. (2024; in review) and Selten et al. (2024; in prep). Genome sequences are obtained by Illumina-only sequencing of microbial cultures. Illumina reads were demultiplexed and cleaned with cutadapt (version 2.8) (Martin, 2011) and assembled into genomes using A5 (A5-miseq version 20160825) (Coil et al., 2014). Genome contamination and heterogeneity was checked with CheckM (version 1.1.3) (Parks et al., 2015) and any genomes with multiple single copy gene occurrences were subjected to MaxBin (version 2.2.7) (Wu et al., 2014) to separate the genomes from contaminated bacterial cultures. Any non-bacterial contigs in the genome assemblies were removed using MMSeqs2 (version 13.45111) (Steineigger &amp; Sch&ouml;ding, 2017). Open reading frames were found and annotated by PROKKA (version 1.14.6) (Seemann, 2014) and EggNOG (version 2.1.4-2) (Cantalapiedra et al., 2021) respectively.&nbsp;Microbial cultures were derived from&nbsp;<em>Arabidopsis thaliana</em> roots grown in Reijerscamp soil, described in Stringlis <em>et al</em>., 2018 https://doi.org/10.1073/pnas.1722335115.</p> <p><strong>The uploaded files are</strong></p> <ol> <li>Genome assemblies</li> <li>Prokka gene predictions in GFF3 format</li> <li>Predicted transcripts from genes in (2)</li> <li>Predicted proteins from genes in (2), and</li> <li>EggNOG annotations for the proteins in (4)</li> </ol> <p><strong>Genomes and annotations pending upload om NCBI GenBank (April 2024)</strong></p>

opencc-by-nc-nd-4.0Apr 2024View details →
dryad32/100

Draft genome assembly of the alpine carnation Dianthus sylvestris (Caryophyllaceae)

<p>A draft reference sequence for the alpine carnation <em>Dianthus sylvestris</em> was obtained using a combination of paired-end and mate-pair Illumina libraries implemented in the Allpath-LG assembler. The initial assembly was improved through nucleotide correction, gap filling and the assessment of misassemblies, and subsequently annotated with the support of RNA-seq data through the Maker2 pipeline. The final assembly consists of ~440 MB in ~21 K scaffolds (N50 ~61 Kb) covering 86% of Busco eukaryotic genes and ~73% of the estimated genome size (600 Mb). Structural annotation includes ~22 K predicted genes (N50 ~1.8 Kb). </p>

opencc-zeroFeb 2023View details →
zenodo32/100

Jaminaea angkorensis draft genome assembly

<p>A draft assembly of the genome of yeast&nbsp;Jaminaea angkorensis produced by SPAdes assembler from Illumina reads, The main purpose is to serve as a reference for analyzing DNA methylation in nanopore sequencing data produced with in vitro methylation. The nanopore data can be downloaded from ENA bioproject&nbsp;<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB64246">PRJEB64246</a>.</p>

opencc-by-4.0Jul 2023View details →
dryad32/100

Data from: Development of genomic tools in a widespread tropical tree, Symphonia globulifera L.f.: a new low-coverage draft genome, SNP and SSR markers

Open the record for dataset details and reuse information.

publicOct 2016View details →
dryad32/100

Data from: DISCOMARK: nuclear marker discovery from orthologous sequences using draft genome data

Open the record for dataset details and reuse information.

publicJul 2016View details →
dryad32/100

The draft genome of the blood pheasant (Ithaginis cruentus): phylogeny and high-altitude adaptation

Open the record for dataset details and reuse information.

publicFeb 2021View details →
dryad32/100

Draft genomes of two Atlantic bay scallop subspecies Argopecten irradians irradians and A. i. concentricus

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad32/100

Data from: Draft genome of the American eel (Anguilla rostrata)

Open the record for dataset details and reuse information.

publicOct 2016View details →
dryad32/100

Draft genome assembly of alpine carnations Dianthus sylvestris and D. carthusianorum (Caryophyllaceae)

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad28/100

Data from: Draft assembly of elite inbred line PH207 provides insights into genomic and transcriptome diversity in maize

Intense artificial selection over the last 100 years has produced elite maize (Zea mays) inbred lines that combine to produce high-yielding hybrids. To further our understanding of how genome and transcriptome variation contribute to the production of high-yielding hybrids, we generated a draft genome assembly of the inbred line PH207 to complement and compare with the existing B73 reference sequence. B73 is a founder of the Stiff Stalk germplasm pool, while PH207 is a founder of Iodent germplasm, both of which have contributed substantially to the production of temperate commercial maize and are combined to make heterotic hybrids. Comparison of these two assemblies revealed over 2,500 genes present in only one of the two genotypes and 136 gene families that have undergone extensive expansion or contraction. Transcriptome profiling revealed extensive expression variation, with as many as 10,564 differentially expressed transcripts and 7,128 transcripts expressed in only one of the two genotypes in a single tissue. Genotype-specific genes were more likely to have tissue/condition-specific expression and lower transcript abundance. The availability of a high-quality genome assembly for the elite maize inbred PH207 expands our knowledge of the breadth of natural genome and transcriptome variation in elite maize inbred lines across heterotic pools.

opencc-zeroDec 2015View details →
dryad28/100

Data from: A draft fur seal genome provides insights into factors affecting SNP validation and how to mitigate them

Custom genotyping arrays provide a flexible and accurate means of genotyping single nucleotide polymorphisms (SNPs) in a large number of individuals of essentially any organism. However, validation rates, defined as the proportion of putative SNPs that are verified to be polymorphic in a population, are often very low. A number of potential causes of assay failure have been identified, but none have been explored systematically. In particular, as SNPs are often developed from transcriptomes, parameters relating to the genomic context are rarely taken into account. Here, we assembled a draft Antarctic fur seal (Arctocephalus gazella) genome (assembly size: 2.41Gb; scaffold/contig N50: 3.1Mb/27.5kb). We then used this resource to map the probe sequences of 144 putative SNPs genotyped in 480 individuals. The number of probe-to-genome mappings and alignment length together explained almost a third of the variation in validation success, indicating that sequence uniqueness and proximity to intron-exon boundaries play an important role. The same pattern was found after mapping the probe sequences to the Walrus and Weddell seal genomes, suggesting that the genomes of species divergent by as much as 23 million years can hold information relevant to SNP validation outcomes. Additionally, re-analysis of genotyping data from seven previous studies found the same two variables to be significantly associated with SNP validation success across a variety of taxa. Finally, our study reveals considerable scope for validation rates to be improved, either by simply filtering for SNPs whose flanking sequences align uniquely and completely to a reference genome, or through predictive modeling.

opencc-zeroDec 2015View details →
zenodo28/100

Draft genome sequence of Xylaria bambusicola isolate GMP-LS, the root and basal stem rot pathogen of sugarcane in Indonesia

<p>Supplementary Figure 1</p><p>Supplementary Table 1</p><p>Supplementary Table 2</p>

opencc-by-4.0Dec 2023View details →
zenodo28/100

Draft Genome Sequences of Pseudomonas sp. Strain MWU13-2924, Isolated from a Wild Cranberry Bog in Truro, MA

<p>Annotated genome of Pseudomonas sp. MWU13.2924.</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

Figure 1 in A draft genome of a field-collected Steinernema feltiae strain NW

Figure 1: Genome alignment of strains SN and NW. Only top 100 longest scaffolds from SN (laid across the x-axis) and top 100 longest contigs from NW (y-axis) were shown here to minimize noise. Each contig/scaffold is shown between two lines (vertical for SN and horizontal for NW) along the axes. A colored dot is plotted wherever the two sequences agree; the forward matches are shown in purple, while the reverse matches are shown in blue. If the two genomes were perfectly identical, a series of purple dots would be drawn diagonally.

opencc-by-4.0Mar 2020View details →
zenodo28/100

Test datasets for GenoVi: draft and complete genomes

<p>Test dataset for GenoVi (Genome Visualizer).<br> All the genomes available in this repository were used to create all the analysis done by Cumsille et al., 2022 for the publication of GenoVi.</p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record