Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

89

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

89 results for “Draft Genomes”

Learn how ShareScore rates datasets ↗
zenodo36/100

Draft genomes of 972 carbapenem-resistant Pseudomonas aeruginosa isolates (shovill assemblies of Reyes et al 2023 dataset raw reads)

<p>This dataset contains shovill assemblies of raw reads released under&nbsp;NCBI BioProject PRJNA824880 generated in Reyes J et al (The Lancet Microbe. Volume 4 Issue 3 Pages e159-e170 (March 2023); DOI: 10.1016/S2666-5247(22)00329-9)</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
dryad36/100

Long-read-based draft genome sequence of Indian black gram IPU-94-1 'Uttara': Insights into disease resistance and seed storage protein genes

<p>Black gram [Vigna mungo (L.) Hepper var. <em>mung<a>o</a></em>] [LAV1] is a warm-season legume highly prized for its protein content along with significant folate and iron proportions. To expedite the genetic enhancement of black gram, a high-quality draft genome from the center of origin of the crop is indispensable. Here, we established a draft genome sequence of an Indian black gram cultivar, 'Uttara' (IPU 94-1), known for its high resistance to mungbean yellow mosaic virus. Pacific Biosciences of California, Inc. (PacBio) single-molecule real-time (SMRT) and Illumina sequencing assembled a draft reference-guided assembly with a cumulative size of ~454.4 Mb, of which, 444.4 Mb was anchored on 11 pseudomolecules corresponding to 11 chromosomes. Uttara assembly denotes features of a high-quality draft genome illustrated through high N50 value (42.88 Mb), gene completeness (benchmarking universal single-copy ortholog [BUSCO] score 94.17%), and low levels of ambiguous nucleotides (N) percent (0.0005%). Gene discovery using transcript evidence predicted 28,881 protein-coding genes, from which, ~95% were functionally annotated. A global survey of genes associated with disease resistance revealed 119 nucleotide binding site–leucine rich repeat (NBS-LRR) proteins, while 23 genes encoding seed storage proteins (SSPs) were discovered in black gram. A large set of microsatellite loci were discovered for marker development in the crop. Our draft genome of an Indian black gram provides the foundational genomic resources for the improvement of important agronomic traits and ultimately will help in accelerating black gram breeding programs.</p>

opencc-zeroJul 2023View details →
zenodo36/100

Draft Genome sequencing of Nocardia sp. strain WB46 Isolated from Salix purpurea Growing in a Site Chronically Contaminated by Petroleum Hydrocarbons

<p>Assembled contigs of the genome of <em>Nocardia</em> sp. strain WB46 isolated from <em>the rhizosphere of Salix purpurea </em>growing in a site chronically contaminated by petroleum hydrocarbons located at Varennes, Qubec, Canada.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Draft genome and annotation of Raphidonema monicae

<p>Microalgae synthesize diverse lipids that perform a wide variety of structural, metabolic, and signalling roles in the cell. However, we know relatively little about lipid metabolism across different protist groups, and especially those adapted to extreme environments. As part of a project to investigate the lipid metabolism of <em>Raphidonema monicae</em> strain SAG 2030, a cold-adapted microalga isolated from Antarctica, a draft genome was assembled and annotated.</p> <p>The genome was assembled from a single DNA library sequenced with Illumina 150 bp paired-end reads and assembled by NovoGene (Hong Kong) with SOAPdenovo2. Novogene also perfomed curation and structural annotation of CDS and peptide sequences. The data were subsequently functionally annotated in our lab using BlastP, InterProScan and emapper.py, and the results were curated with OmicsBox software.</p> <p>The assembly sequence and annotation are provided as data files supporting the lipidomic investigation of the effects of abiotic stress and include the following:</p> <p>File 1. &lsquo;<strong>KAD1.seq</strong>&rsquo; is the assembled genome sequence contigs that have been curated by NovoGene, .fasta format.</p> <p>File 2. &lsquo;<strong>KAD1.gff</strong>&rsquo; is the annotation file .gff format for the gene CDS, .gff format.</p> <p>File 3. &lsquo;<strong>KAD1.cds</strong>&rsquo; contains the CDS nucleotide sequences, .fasta format.</p> <p>File 4. &lsquo;<strong>KAD1.pep</strong>&rsquo; The corresponding peptide sequences, .fasta format.</p> <p>File 5. &lsquo;<strong>kad.annotation.table.xls</strong>&rsquo; is the combined functional annotation of the peptide sequences using BlastP, InterProScan and Emapper.py in an excel-readable table format.</p>

opencc-by-4.0Jun 2024View details →
dryad36/100

Draft genome of a dung beetle, Phelotrupes auratus

<p>Knowledge of population divergence history is key to understanding organism diversification mechanisms. The geotrupid dung beetle, <em>Phelotrupes auratus</em>, which inhabits montane forests and exhibits three color forms (red, green, and indigo), diverged into five local populations (west/red, south/green, south/indigo, south/red, and east/red) in the Kinki District of Honshu, Japan, based on the combined interpretation of genetic cluster and color-form data. Here, we estimated the demographic histories of these local populations using the newly assembled draft genome sequence of <em>P. auratus</em> and whole-genome resequencing data obtained from each local population. Using coalescent simulation analysis, we estimated <em>P. auratus</em> population divergences at ca. 3,800, 2,100, 600, and 200 years ago, with no substantial gene flow between diverged populations, implying the existence of persistent barriers to gene flow. Notably, the last two divergence events led to three local populations with different color forms. The initial divergence may have been affected by climatic cooling around that time, and the last three divergence events may have been associated with the increasing impact of human activities. Both climatic cooling and increasing human activity may have caused habitat fragmentation and a reduction in the numbers of large mammals supplying food (dung) for <em>P. auratus</em>, thereby promoting the decline, segregation, and divergence of local populations. Our research demonstrates that geographic population divergence in an insect with conspicuous differences in traits such as body color may have occurred rapidly under the influence of human activity.</p>

opencc-zeroJan 2023View details →
zenodo36/100

Draft Genome Sequences of Two Bacteriocin-Producing Enterococcus faecium Strains Isolated from Nonfermented Animal Foods in Spain

<p>Raw sequences of two bacteriocin-producing Enterococcus faecium.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
dryad36/100

PacBio whole-genome sequencing and draft assembly of Stentor coeruleus

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Long-read-based draft genome sequence of Indian black gram IPU-94-1 ‘Uttara’: Insights into disease resistance and seed storage protein genes

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad36/100

Data from: An annotated draft genome of the mountain hare (Lepus timidus)

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad36/100

Draft genome of a dung beetle, Phelotrupes auratus

Open the record for dataset details and reuse information.

publicJan 2023View details →
zenodo32/100

Supplementary data for Kvist et. al. 2020 "Draft genome of the European medicinal leech Hirudo medicinalis (Annelida: Clitellata: Hirudiniformes) with emphasis on anticoagulants"

<p>anticoagulant_prots.tar.gz: Proteins sequence, alignment, exon structure, and tree files for identified anticoagulant proteins plus putative anticoagulant proteins from the&nbsp;Hirudo medicinalis&nbsp;ROM11733 v1 assembly</p> <p>Himedicinalis_ROM11733_annotation.tar.gz: InterProScan, UniProtKB, MAKER, RepeatModeller, Rfam, and tRNAscan-SE annotation files&nbsp;for&nbsp;the&nbsp;Hirudo medicinalis&nbsp;ROM11733 v1 assembly</p> <p>PhymmBL_tax_assign.tar.gz: PhymmBL files of the tree rounds of taxnonomic assignment</p>

opencc-by-nc-4.0Nov 2019View details →
zenodo32/100

Draft genome assembly of Aipysurus laevis

<p>Draft genome assemblies of Aipysurus laevis, assembled by SOAPdenovo2. Version 1.0 is the final draft, with version 0.1 being a first pass assembly. The annotation file (GFF3) was generated for version 1.0 using a range of homology and predictive annotation methods.</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

bslucas98/Dsaccharalis_genomeassembly: A first draft genome of the sugarcane borer, Diatraea saccharalis

<p>This repository contains: a) Additional&nbsp;Tables S1 and S2; b) The script used to select a subset of raw reads; c) The script used to select protein gene models having &gt;=50% similarity to the B. mori model.</p>

openother-pdOct 2020View details →
dryad32/100

Data from: Development of genomic tools in a widespread tropical tree, Symphonia globulifera L.f.: a new low-coverage draft genome, SNP and SSR markers

Population genetic studies in tropical plants are often challenging because of limited information on taxonomy, phylogenetic relationships and distribution ranges, scarce genomic information and logistic challenges in sampling. We describe a strategy to develop robust and widely applicable genetic markers based on a modest development of genomic resources in the ancient tropical tree species Symphonia globulifera L.f. (Clusiaceae), a keystone species in African and Neotropical rainforests. We provide the first low-coverage (11X) fragmented draft genome sequenced on an individual from Cameroon, covering 1.027 Gbp or 67.5% of the estimated genome size. Annotation of 565 scaffolds (7.57 Mbp) resulted in the prediction of 1046 putative genes (231 of them containing a complete open reading frame) and 1523 exact simple sequence repeats (SSRs, microsatellites). Aligning a published transcriptome of a French Guiana population against this draft genome produced 923 high-quality single nucleotide polymorphisms. We also preselected genic SSRs in silico that were conserved and polymorphic across a wide geographical range, thus reducing marker development tests on rare DNA samples. Of 23 SSRs tested, 19 amplified and 18 were successfully genotyped in four S. globulifera populations from South America (Brazil and French Guiana) and Africa (Cameroon and São Tomé island, FST = 0.34). Most loci showed only population-specific deviations from Hardy–Weinberg proportions, pointing to local population effects (e.g. null alleles). The described genomic resources are valuable for evolutionary studies in Symphonia and for comparative studies in plants. The methods are especially interesting for widespread tropical or endangered taxa with limited DNA availability.

opencc-zeroDec 2015View details →
zenodo32/100

Draft Genome Sequences of 38 Aspergillus parasiticus isolates

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

Draft genome assembly of Macrophomina pseudophaseolina strain WAC 2767, and ex-epitype strain of M. phaseolina.

<p>This genome assembly is published in&nbsp;&quot;Draft genome assemblies of <em>Fusarium marasasianum</em>, <em>Huntiella abstrusa</em>, two <em>Immersiporthe knoxdaviesiana</em> isolates, <em>Macrophomina pseudophaseolina</em>, <em>Macrophomina phaseolina</em>, <em>Naganishia randhawae</em>, and <em>Pseudocercospora cruenta</em>, Wingfield, B.D., De Vos, L., Wilson, A.M.&nbsp;<em>et al.</em>&nbsp;IMA Genome - F16.&nbsp;<em>IMA Fungus</em>&nbsp;<strong>13,&nbsp;</strong>3 (2022)&quot;. 10.1186/s43008-022-00089-z</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Draft Genome Sequence of Pseudomonas sp. Strain MWU13-2922, Isolated from a Wild Cranberry Bog in Truro, Massachusetts

<p>Annotated genome of Pseudomonas sp. MWU13-2922</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Draft Genome Sequences of Aquitalea sp strain MWU-14.2238

<p>Annotated genome of Aquitalea sp MWU-14.2238.&nbsp;</p>

opencc-by-3.0-usMar 2022View details →
zenodo32/100

Draft Genome Sequences of Pseudomonas sp. Strain MWU12-2233, Isolated from a Wild Cranberry Bog in Provincetown, Massachusetts

<p>Annotated genome of Pseudomonas sp. MWU12-2233.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Draft Genome Sequences of Pseudomonas sp strain MWU-15.20650

<p>Annotated genome of Pseudomonas sp. MWU 15-20650.</p>

opencc-by-3.0-usMar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record