Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
89
datasets available to search
ShareScore release 0.9.0
Dataset results
89 results for “Draft genome”
Draft Genome Sequences of Aquitalea sp strain MWU-14.2238
<p>Annotated genome of Aquitalea sp MWU-14.2238. </p>
Draft Genome Sequence of Pseudomonas sp. strain MWU13.2105, isolated from Wild Cranberry Bog in Truro, Massachusetts
<p>Annotated genome of Pseudomonas sp. strain MWU13.2105.</p>
Draft Genome Sequence of Pseudomonas sp. strain MWU13.2100, isolated from Wild Cranberry Bog in Truro, Massachusetts
<p>Annotated genome of Pseudomonas sp. strain MWU13.2100.</p>
Draft Genome Manuscript for Curtobacterium sp. Isolated from Berries Surfaced in Commercial Cranberry Bogs in Massachusetts, USA
<p>Annotated genome of Curtobacterium sp. MWU13.2055</p>
Draft Genome Manuscript for Pseudomonas sp. Strain MWU13.3659 Isolated from Berries Surfaced in Commercial Cranberry Bogs in Massachusetts, USA
<p>Annotated genome of Pseudomonas sp. MWU13.3659</p>
Draft genome sequences of Arabidopsis thaliana-associated micro-organisms from Reijerscamp soil, the Netherlands
<p><strong>Methodological summary and relevant references</strong></p> <p>Compressed tar archive containing 447 draft bacterial genomes and their annotations used in several studies including Fourie <em>et al</em>. (2024; in review) and Selten et al. (2024; in prep). Genome sequences are obtained by Illumina-only sequencing of microbial cultures. Illumina reads were demultiplexed and cleaned with cutadapt (version 2.8) (Martin, 2011) and assembled into genomes using A5 (A5-miseq version 20160825) (Coil et al., 2014). Genome contamination and heterogeneity was checked with CheckM (version 1.1.3) (Parks et al., 2015) and any genomes with multiple single copy gene occurrences were subjected to MaxBin (version 2.2.7) (Wu et al., 2014) to separate the genomes from contaminated bacterial cultures. Any non-bacterial contigs in the genome assemblies were removed using MMSeqs2 (version 13.45111) (Steineigger & Schöding, 2017). Open reading frames were found and annotated by PROKKA (version 1.14.6) (Seemann, 2014) and EggNOG (version 2.1.4-2) (Cantalapiedra et al., 2021) respectively. Microbial cultures were derived from <em>Arabidopsis thaliana</em> roots grown in Reijerscamp soil, described in Stringlis <em>et al</em>., 2018 https://doi.org/10.1073/pnas.1722335115.</p> <p><strong>The uploaded files are</strong></p> <ol> <li>Genome assemblies</li> <li>Prokka gene predictions in GFF3 format</li> <li>Predicted transcripts from genes in (2)</li> <li>Predicted proteins from genes in (2), and</li> <li>EggNOG annotations for the proteins in (4)</li> </ol> <p><strong>Genomes and annotations pending upload om NCBI GenBank (April 2024)</strong></p>
Draft genome assembly of the alpine carnation Dianthus sylvestris (Caryophyllaceae)
<p>A draft reference sequence for the alpine carnation <em>Dianthus sylvestris</em> was obtained using a combination of paired-end and mate-pair Illumina libraries implemented in the Allpath-LG assembler. The initial assembly was improved through nucleotide correction, gap filling and the assessment of misassemblies, and subsequently annotated with the support of RNA-seq data through the Maker2 pipeline. The final assembly consists of ~440 MB in ~21 K scaffolds (N50 ~61 Kb) covering 86% of Busco eukaryotic genes and ~73% of the estimated genome size (600 Mb). Structural annotation includes ~22 K predicted genes (N50 ~1.8 Kb). </p>
Jaminaea angkorensis draft genome assembly
<p>A draft assembly of the genome of yeast Jaminaea angkorensis produced by SPAdes assembler from Illumina reads, The main purpose is to serve as a reference for analyzing DNA methylation in nanopore sequencing data produced with in vitro methylation. The nanopore data can be downloaded from ENA bioproject <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB64246">PRJEB64246</a>.</p>
Data from: Development of genomic tools in a widespread tropical tree, Symphonia globulifera L.f.: a new low-coverage draft genome, SNP and SSR markers
Open the record for dataset details and reuse information.
Data from: DISCOMARK: nuclear marker discovery from orthologous sequences using draft genome data
Open the record for dataset details and reuse information.
The draft genome of the blood pheasant (Ithaginis cruentus): phylogeny and high-altitude adaptation
Open the record for dataset details and reuse information.
Draft genomes of two Atlantic bay scallop subspecies Argopecten irradians irradians and A. i. concentricus
Open the record for dataset details and reuse information.
Data from: Draft genome of the American eel (Anguilla rostrata)
Open the record for dataset details and reuse information.
Draft genome assembly of alpine carnations Dianthus sylvestris and D. carthusianorum (Caryophyllaceae)
Open the record for dataset details and reuse information.
Data from: Draft assembly of elite inbred line PH207 provides insights into genomic and transcriptome diversity in maize
Intense artificial selection over the last 100 years has produced elite maize (Zea mays) inbred lines that combine to produce high-yielding hybrids. To further our understanding of how genome and transcriptome variation contribute to the production of high-yielding hybrids, we generated a draft genome assembly of the inbred line PH207 to complement and compare with the existing B73 reference sequence. B73 is a founder of the Stiff Stalk germplasm pool, while PH207 is a founder of Iodent germplasm, both of which have contributed substantially to the production of temperate commercial maize and are combined to make heterotic hybrids. Comparison of these two assemblies revealed over 2,500 genes present in only one of the two genotypes and 136 gene families that have undergone extensive expansion or contraction. Transcriptome profiling revealed extensive expression variation, with as many as 10,564 differentially expressed transcripts and 7,128 transcripts expressed in only one of the two genotypes in a single tissue. Genotype-specific genes were more likely to have tissue/condition-specific expression and lower transcript abundance. The availability of a high-quality genome assembly for the elite maize inbred PH207 expands our knowledge of the breadth of natural genome and transcriptome variation in elite maize inbred lines across heterotic pools.
Data from: A draft fur seal genome provides insights into factors affecting SNP validation and how to mitigate them
Custom genotyping arrays provide a flexible and accurate means of genotyping single nucleotide polymorphisms (SNPs) in a large number of individuals of essentially any organism. However, validation rates, defined as the proportion of putative SNPs that are verified to be polymorphic in a population, are often very low. A number of potential causes of assay failure have been identified, but none have been explored systematically. In particular, as SNPs are often developed from transcriptomes, parameters relating to the genomic context are rarely taken into account. Here, we assembled a draft Antarctic fur seal (Arctocephalus gazella) genome (assembly size: 2.41Gb; scaffold/contig N50: 3.1Mb/27.5kb). We then used this resource to map the probe sequences of 144 putative SNPs genotyped in 480 individuals. The number of probe-to-genome mappings and alignment length together explained almost a third of the variation in validation success, indicating that sequence uniqueness and proximity to intron-exon boundaries play an important role. The same pattern was found after mapping the probe sequences to the Walrus and Weddell seal genomes, suggesting that the genomes of species divergent by as much as 23 million years can hold information relevant to SNP validation outcomes. Additionally, re-analysis of genotyping data from seven previous studies found the same two variables to be significantly associated with SNP validation success across a variety of taxa. Finally, our study reveals considerable scope for validation rates to be improved, either by simply filtering for SNPs whose flanking sequences align uniquely and completely to a reference genome, or through predictive modeling.
Draft genome sequence of Xylaria bambusicola isolate GMP-LS, the root and basal stem rot pathogen of sugarcane in Indonesia
<p>Supplementary Figure 1</p><p>Supplementary Table 1</p><p>Supplementary Table 2</p>
Draft Genome Sequences of Pseudomonas sp. Strain MWU13-2924, Isolated from a Wild Cranberry Bog in Truro, MA
<p>Annotated genome of Pseudomonas sp. MWU13.2924.</p>
Figure 1 in A draft genome of a field-collected Steinernema feltiae strain NW
Figure 1: Genome alignment of strains SN and NW. Only top 100 longest scaffolds from SN (laid across the x-axis) and top 100 longest contigs from NW (y-axis) were shown here to minimize noise. Each contig/scaffold is shown between two lines (vertical for SN and horizontal for NW) along the axes. A colored dot is plotted wherever the two sequences agree; the forward matches are shown in purple, while the reverse matches are shown in blue. If the two genomes were perfectly identical, a series of purple dots would be drawn diagonally.
Test datasets for GenoVi: draft and complete genomes
<p>Test dataset for GenoVi (Genome Visualizer).<br> All the genomes available in this repository were used to create all the analysis done by Cumsille et al., 2022 for the publication of GenoVi.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.