Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

89

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

89 results for “draft genome”

Learn how ShareScore rates datasets ↗
zenodo52/100

Draft genome assembly version 1 of the meadow spittlebug Philaenus spumarius (Linnaeus, 1758) (Hemiptera, Aphrophoridae)

<p>We sequenced the genome of the meadow spittlebug, <em>Philaenus spumarius </em>(Linnaeus, 1758), the main insect vector of <em>Xylella fastidiosa </em>Wells et al. 1987 in Europe (Saponari et al., 2014), using 10x Chromium linked-reads. A single <em>P. spumarius</em> adult female from Portugal (Fontanelas, Sintra; GPS location: 38&deg;50&#39;15.75&quot;N; 9&deg;25&#39;20.77&quot;W), collected in September of 2018, was selected for genome sequencing. This population was initially surveyed for colour polymorphism in 1988 (Quartau &amp; Borges, 1997) and was later included in phylogeographic and population genomic studies of this species (Rodrigues et al., 2014; Seabra et al., unpublished). It is also geographically close to the population from which the individual used for the first partial genome assembly was collected (Rodrigues et al., 2016). The availability of this previous genetic information contributed to the choice of this population as the source of genomic material for whole genome sequencing. A subset of males from the same collection date were analysed for genitalia morphology to confirm species identification, as the best diagnostic characters are the appendages of the aedeagus (Drosopoulos &amp; Quartau, 2002).</p> <p>The genomic DNA of the <em>P. spumarius</em> adult from Sintra was extracted using Illustra Nucleon Phytopure kit according to the manufacturer&rsquo;s instructions (GE Healthcare). We assessed the quality and concentration of the DNA using Femto fragment analyser (Agilent). 10x Chromium library preparation and Illumina genome sequencing (HiSeq X, 150bp paired-end) were performed by Novogene Bioinformatics Technology Co, Beijing, China, in accordance with standard protocols.</p> <p>To create the <em>de novo</em> 10x Chromium assembly we ran Supernova 2.1.1 (Weisenfeld et al., 2017) on the 10x Chromium linked-read data with default parameters, using 1.0 billion reads corresponding to 56X coverage. To improve the initial supernova assembly, we performed iterative scaffolding using all of the 10x raw data (2.3 billion of reads). We ran two rounds of Scaff10x (https://github.com/wtsi-hpag/Scaff10X), followed by mis-assembly detection and correction with Tigmint (Jackman et al., 2018). This was followed by a final round of scaffolding with ARCS (Yeo et al., 2018). The assembly was checked for contamination using the BlobTools pipeline (version 0.9.19; Laetsch and Blaxter 2017;&nbsp;Kumar et al., 2013) and k-mer content was analysed with the KAT comp tool (Mapleson et al., 2017). In order to perform these analyses, it was necessary to remove the 10x linked barcodes from the reads with the script process_10xReads.py (https://github.com/ucdavis-bioinformatics/proc10xG).&nbsp;We assessed the quality of our draft genome assembly by searching for conserved, single copy, arthropod genes (n=1,066) with Benchmarking Universal Single-Copy Orthologs (BUSCO) v3.0 (Waterhouse et al., 2018).</p> <p>With the above assembly procedure, we obtained a final assembly of 2.7 Gb, having a scaffold N50 length of 116 Kb (contig N50 = 18 Kb) and the longest scaffold was 3.7 Mb. The length of the assembly was consistent with the genome size estimated by flow cytometry (Rodrigues et al., 2016). The k-mer distribution indicated high heterozygosity, estimated at 2.3%. BlobTools analyses revealed the presence of contigs assigned to <em>Sodalis </em>spp. (Enterobacteriaceae), a symbiont in members of tribe Philaenini (Koga et al., 2013). These contigs were filtered from the final assembly. Gene completeness assessment shows that 956 (89.6%) among 1,066 BUSCOs were &nbsp;found as complete copies, with only 26 (2.4%) missing. Of the BUSCOs that were detected, 878 (82.4%) were complete and single-copy, 78 (7.3%) were complete and duplicated and 84 (7.9%) were fragmented.</p> <p>In conclusion, due in part to high (2.3%) heterozygosity levels, the <em>P. spumarius</em> version 1 genome assembly is highly fragmented. Nonetheless, the assembly is considered complete and is likely to contain the majority of the gene content of <em>P. spumarius.</em></p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Supplementary data for draft genome of a member of the ascomycotal fungal genus Pseudopithomyces (family Didymosphaeriaceae)

<p><span lang="EN-US">We update our previous draft genome of a member of genus <em>Pseudopithomyces</em> (previously annotated as <em>Pseudopithomyces maydicus</em> strain SBW1, now reannotated as <em>Pseudopithomyces sp</em>. strain SBW1. The new draft genome is based on a hybrid assembly utilising both ONT and Illumina data. The draft genome is comprised of 43 contigs with a total length of 39.65Mbp. We predict 13,669 protein coding gene models, of which 4241 (31%) were annotated to KEGG Orthology. Taxonomic assignment to <em>Pseudopithomyces sp.</em> was supported by comparative analysis of extracted ITS regions, mitochondrial DNA sequence and whole genome comparisons using <em>k</em>-mer sketches. </span></p> <p>&nbsp;</p> <p><span lang="EN-US">The following items of Additional Data Files are made available in this repository:</span></p> <p><span lang="EN-US">Additional Data File 1: contigs.fasta</span></p> <p><span lang="EN-US">FASTA file of entire assembly.&nbsp;</span><span lang="EN-US">&nbsp;</span></p> <p>&nbsp;</p> <p><span lang="EN-US">Additional Data File 2: draft_genome.fasta</span></p> <p><span lang="EN-US">FASTA file of draft whole genome sequence.</span></p> <p>&nbsp;</p> <p><span lang="EN-US">Additional Data File 3: ITS_full.fasta</span></p> <p><span lang="EN-US">FASTA file containing full length ITS sequences from contig 23 and contig 42.</span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 4: 2NJ47W4U013-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for ITS region contained on contig 23</span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 5: 2NJXH2GE013-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for ITS region contained on contig 42</span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 6: 2NM79EUG016-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for the mitochondrial genome from contig 40 </span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 6: sourmash_bc10_hy_pm1_3.txt</span></p> <p><span lang="EN-US">Text file containing the MASH similarities of the draft genome compared to 18,883 fungal genomes.</span></p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

Genome drafts of Lotmaria passim strains C2 and C3 isolated from honeybees in Spain

<p>Lotmaria passim is a highly prevalent parasite of honeybees. Herein is reported the draft&nbsp;genome sequences of L. passim C2 and C3 strains of 27.15 Mbp and 26.94 Mbp, respectively.&nbsp;The genomes were sequenced using Illumina MiSeq platform and will allow for further&nbsp;comparative and functional genomics studies.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Support files for draft genome of the Asian buffalo leech, Hirudinaria manillensis

<p>The assembly fasta file, TE and gene gff files, and annotations for the&nbsp;draft genome of the Asian buffalo leech, Hirudinaria manillensis.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Supplementary dataset to "Draft genome assembly of the biofuel grass crop Miscanthus sacchariflorus"

<p><em>Miscanthus sacchariflorus</em> (Maxim.) Hack. is a C4 perennial rhizomatous biofuel grass crop. <em>M. sacchariflorus</em> is among the most widely distributed species within the genus, particularly at cold northern latitudes, and one of the progenitor species of the main biomass commercial crop <em>M.&nbsp;&times;&nbsp;giganteus</em>. We generated a 2.54 Gbps whole-genome assembly of the diploid <em>M. sacchariflorus</em> &ldquo;Robustus 297&rdquo; genotype, which represented ~59% of the expected genome size. We later anchored this assembly in the chromosomal-scale <em>M. sinensis</em> genome to improve its contiguity. We annotated 86,767 and 69,049 protein-coding genes in the unanchored and anchored, respectively. We estimated our assemblies include ~85% of the <em>M. sacchariflorus</em> genes based on homology, core markers and RNA-seq alignments stats. Raw data and further metadata are available under Bioproject PRJNA435476.</p> <ul> <li>Msac_v2.fasta: Unanchored whole-genome assembly (WGA) of M. sacchariflorus in FASTA format.</li> <li>Msac_v3.fasta: The previous WGA re-scaffolded with the M. sinensis public reference.</li> <li>Msac_v3.agp: Chromosomal position in the M. sinensis reference of the previous scaffolds in Msac_v3.fasta</li> <li>Msac_v2.gff3: Gene annotation of the unanchored WGA in GFF3 format, which contains 86,767 coding genes</li> <li>Msac_v3.gff3: Gene annotation of the anchored WGA in GFF3 format, which contains 69,049 coding genes</li> <li>Msac_v2.func_annot.tsv: Text table containing the functional annotation of the 86,767 coding genes in Msac_v2.gff3</li> <li>Msac_v2.repeats_annotation.gff3: Repeats annotation (Repeatmasker) of the unanchored reference.</li> <li>Msac_v2.masked.fasta.gz: Repeats-masked version (Repeatmasker) of Msac_v2.fasta</li> <li>all.satsuma.blocks_Msac_v2-vs-Msin.gz: Every alignment from scaffolds in Msac_v3.fasta into M. sinensis reference</li> <li>Msac_v2.orthology_Msin.tsv: Ortologous between Msac_v2 and M. sinensis</li> <li>Msac_v3-vs-Msin.tsv: Ortologous between Msac_v3 and M. sinensis</li> </ul>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Predicted genes from the Amblyomma americanum draft genome assembly

<p>Data for pub "Predicted genes from the&nbsp;<em>Amblyomma americanum&nbsp;</em>draft genome assembly."</p> <ul> <li>Amblyomma_americanum_filtered_assembly.fasta: Decontaminated A. americanum genome with bacterial contigs removed</li> <li>Amblyomma_americanum_bacterial_contigs_info.tsv: Information about contigs classified as bacteria that were removed</li> <li>Amblyomma_americanum_annotation_data.tar.gz: Directory of annotation data produced by EvidenceModeler as part of the nf-core/genomeannotator workflow. Includes files in FASTA format (predicted genes and proteins), set of proteins clustered at 99% identity in FASTA format, and annotations in both GFF3 and GTF formats. GTF file produced from the GFF3 file with AGAT.</li> <li>Amblyomma_americanum_transcriptome_assembly_data.tar.gz: Directory of data generated for the transcriptome assembly that was used for gene prediction</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Draft de novo genome assemblies of a male and female Amphibolurus muricatus (jacky dragon)

<p>Four de novo nuclear genome assemblies of <em>Amphibolurus muricatus</em></p> <p><strong>Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> &bull; AmpMurF_1.0.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_1.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 1.1: Further scaffolding of assembly 1.0 using RNA-seq data</strong><br> &bull; AmpMurF_1.1.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_1.1.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 2.0: Further scaffolding of assembly 1.0 using SLR-superscaffolder</strong><br> &bull; AmpMurF_2.0.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_2.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing assembly</strong><br> &bull; AmpMurF_3.0.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_3.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Methods<br> Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> Male and female <em>A. muricatus</em> genome sequencing libraries were constructed on the Chromium system (10x Genomics, Pleasanton, CA, USA) by the Ramaciotti Centre for Genomics (Sydney, Australia). The Chromium instrument enables unique barcoding of long stretches of DNA on gel beads. The barcodes allow later reconstruction of long DNA fragments from a series of short DNA fragments with the same barcode (i.e., linked-reads). After barcoding, DNA was sheared into smaller fragments and sequenced on the NovaSeq 6000 platform (Illumina, CA, USA) to generate 151 bp paired-end (PE) reads. A total of 904.9 M raw 10x Genomics Chromium linked-reads were generated. Raw 10x data were assembled with Supernova v2.1.1 (Weisenfeld et al., 2017) and a FASTA file was generated using the &lsquo;pseudohap style&rsquo; option in Supernova mkoutput. All female (~450 M) and male (~550 M) read pairs were utilised (female sequencing depth ca 50.3&times;; male, ca 47.8&times;). The resulting assemblies was further scaffolded with ARKS v1.0.3 (Coombe et al., 2018), reusing the 10x reads, and the companion LINKS program (v1.8.7) (Warren et al., 2015). ARKS employs a <em>k</em>-mer approach to map linked barcodes to the contigs in the initial Supernova assembly to generate a scaffold graph with estimated distances for LINKS input. These assemblies were denoted AmpMurF_1.0 (female) and AmpMurM_1.0 (male). We used GapCloser v1.12 (part of SOAPdenovo2) (Luo et al., 2012) to fill gaps in the assembly. GapCloser was run using the parameter -l 150) and clean&nbsp;10x Genomics reads PE reads. &nbsp;</p> <p><strong>Assembly 1.1: Further scaffolding using RNA-seq data</strong><br> We attempted to improve the v1.0 genome assemblies&rsquo; contiguity using RNA-sequencing reads. RNA-seq reads (from brain, ovary, and testis; see below) were filtered (i.e., cleaned) to remove adapters and low-quality reads using Flexbar v3.4.0 and used to further re-scaffold the v1.0 assemblies (FASTA files before gapclosing) with P_RNA_scaffolder (Zhu et al., 2018). The default Flexbar settings discards all reads with any uncalled bases. A final round of scaffolding was performed on the resulting assemblies using L_RNA_scaffolder (Xue et al., 2013). These assemblies were denoted AmpMurF_1.1 (female) and AmpMurM_1.1 (male). As before, GapCloser and clean&nbsp;10x Genomics reads were used to fill gaps. &nbsp;&nbsp; &nbsp;</p> <p><strong>Assembly 2.0: Further scaffolding using SLR-superscaffolder</strong><br> As an alternative approach, we attempted to improve the v1.0 genome assemblies&rsquo; contiguity using SLR-superscaffolder (Guo et al., 2021). Briefly, SLR-superscaffolder employs single tube long fragment read (stLFR) sequencing (Wang et al., 2019) reads (see section below) to generate hybrid genome assemblies. The software was run with default parameters except for PE_SEED_MIN=300 (minimum contig size to fill; default 1000). These assemblies were denoted AmpMurF_2.0 (female) and AmpMurM_2.0 (male). GapCloser and clean&nbsp;stLFR reads (with the barcode removed using https://github.com/BGI-Qingdao/stLFR_barcode_split) were used to fill gaps. &nbsp;&nbsp; &nbsp;</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing and supernova assembly</strong><br> We also generated independent assemblies for the individuals sequenced on the 10x Genomics Chromium system using single tube long fragment read (stLFR) sequencing (Wang et al., 2019). BGI (Brisbane, Australia) generated ~100&times;-coverage 100-bp paired-end reads (plus a 42-bp stLFR barcode on the right/_2 read). Low-quality reads, PCR duplicates, and adaptors were removed using SOAPnuke v1.5&nbsp;(Chen et al. 2018). The stLFRdenovo pipeline (<a href="https://github.com/BGI-biotools/stLFRdenovo">https://github.com/BGI-biotools/stLFRdenovo</a>), which is based on Supernova and customized for stLFR data, was used to generate a&nbsp;<em>de novo</em>&nbsp;genome assembly. The stLFRdenovo tool &lsquo;FillGaps&rsquo; was used to fill gaps.</p> <p><strong>References</strong><br> Chen, Y., Chen, Y., Shi, C., Huang, Z., Zhang, Y., Li, S., Li, Y., Ye, J., Yu, C., Li, Z., et al. (2018). SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 7, 1-6.<br> Coombe, L., Zhang, J., Vandervalk, B.P., Chu, J., Jackman, S.D., Birol, I., and Warren, R.L. (2018). ARKS: chromosome-scale scaffolding of human genome drafts with linked read kmers. BMC Bioinformatics 19, 234.<br> Guo, L., Xu, M., Wang, W., Gu, S., Zhao, X., Chen, F., Wang, O., Xu, X., Seim, I., Fan, G., et al. (2021). SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme. BMC Bioinformatics 22, 158.<br> Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18.<br> Wang, O., Chin, R., Cheng, X., Wu, M.K.Y., Mao, Q., Tang, J., Sun, Y., Anderson, E., Lam, H.K., Chen, D., et al. (2019). Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly. Genome Res 29, 798-808.<br> Warren, R.L., Yang, C., Vandervalk, B.P., Behsaz, B., Lagman, A., Jones, S.J., and Birol, I. (2015). LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads. Gigascience 4, 35.<br> Weisenfeld, N.I., Kumar, V., Shah, P., Church, D.M., and Jaffe, D.B. (2017). Direct determination of diploid genome sequences. Genome Res 27, 757-767.<br> Xue, W., Li, J.T., Zhu, Y.P., Hou, G.Y., Kong, X.F., Kuang, Y.Y., and Sun, X.W. (2013). L_RNA_scaffolder: scaffolding genomes with transcripts. BMC Genomics 14, 604.<br> Zhu, B.H., Xiao, J., Xue, W., Xu, G.C., Sun, M.Y., and Li, J.T. (2018). P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads. BMC Genomics 19, 175.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Draft genome assembly of a Japanese Oikopleura dioica male individual (O3), using Nanopore long reads.

<p>This draft assembly was used to validate the chrY scaffolds of the OSKA2016 reference genome in the publication&nbsp; &ldquo;A genome database for a Japanese population of the larvacean Oikopleura dioica&rdquo;, Development Growth and Differentiation, Wang and coll., 2020 (in press). It is provided as supplemental data for the reproducibility of this work; please note that no further polishing has been done to correct sequencing errors.</p> <p>Genome sequence reads were produced on a MinION sequencer (Oxford Nanopore Technologies) using high-molecular weight DNA from a male individual of the Oikopleura dioica species of zooplankton. The individual was related to the laboratory strain established from a western Japanese population that was used to produce the OSKA2016 reference genome. The raw reads were basecalled with the Guppy software version 3.3.0 using its dna_r9.4.1_450bps algorithm, and deposited in the European Nucleotide Archive (Study ID: PRJEB38559). The draft assembly was made with the Flye software version 2.7 with the options --genome-size 65m and --min-overlap 3000.</p>

opencc-zeroJun 2020View details →
zenodo40/100

Supplementary Data for: ONT-based draft genome for Alternaria atra

<p>Species of <em>Alternaria</em> (phylum <em>Ascomycota</em>, family <em>Pleosporaceae</em>) are known as serious plant pathogens, causing major losses on a wide range of crops. <em>Alternaria atra</em><em> (Preuss) Woudenb. &amp; Crous </em>(previously known as <em>Ulocladium atrum</em><em>) </em>can grow as a saprophyte on many hosts and causes<em> </em>Ulocladium blight on potato. It has been reported that it can also be used as a biocontrol agent against a.o. <em>Botrytis cinerea.</em></p> <p>Here we present a scaffold-level reference genome assembly for<em> A. atra.</em> The assembly contains 43 scaffolds with a total length of 39.62 Mbp, with scaffold N50 of 3,893,166 bp , L50 of 4 and the longest 10 scaffolds containing 89.9% of the assembled data. RNA Seq-guided, gene prediction using BRAKER resulted in 12,173 protein-coding genes with their functional annotation.</p>

opencc-by-4.0Jan 2021View details →
dryad40/100

Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi

<p>The Puma lineage within the family Felidae consists of three species that last shared a common ancestor around 4.9 million years ago. Whole-genome sequences of two species from the lineage were previously reported: the cheetah (<em>Acinonyx jubatus</em>) and the mountain lion (<em>Puma concolor</em>). The present report describes a whole-genome assembly of the remaining species, the jaguarundi (<em>Puma yagouaroundi</em>). We sequenced the genome of a male jaguarundi with 10X Genomics linked reads and assembled the whole-genome sequence. The assembled genome contains a series of scaffolds that reach the length of chromosome arms and is similar in scaffold contiguity to the genome assemblies of cheetah and puma, with a contig N50 = 100.2 kbp and a scaffold N50 = 49.27 Mbp. We assessed the assembled sequence of the jaguarundi genome using BUSCO, aligned reads of the sequenced individual and another published female jaguarundi to the assembled genome, annotated protein-coding genes, repeats, genomic variants and their effects with respect to the protein-coding genes, and analyzed differences of the two jaguarundis from the reference mitochondrial genome. The jaguarundi genome assembly and its annotation were compared in quality, variants and features to the previously reported genome assemblies of puma and cheetah. Computational analyzes used in the study were implemented in transparent and reproducible way to allow their further reuse and modification.</p>

opencc-zeroJun 2021View details →
zenodo40/100

Genome annotation file for a draft genome assembly for Nucella lapillus

<p><span>A male specimen of wild <em>Nucella lapillus</em>, measuring approximately 1.5&ndash;3 cm in length, was collected <span>from a rocky shore (mid-upper shore) at low tide from near the quay at <span>Portnahaven, Isle of Islay, Argyll and Bute, Scotland (National Grid Reference NR 16614 51966)</span> on <span>26th June 2023. Genomic DNA was extracted from the non-shell tissue of the specimen, and sequenced using PacBio HIFI and Oxford Nanopore technolgies (ONT) long read sequencing platforms. </span></span>The genome assembly was derived from 40.6 Gb of PacBio HiFi reads (read <span>N50, 11291; N90, 9246</span>), and 61.1 Gb of ONT data (read <span>N50, </span>3643<span>; N90, </span>1546).&nbsp; </span>Annotation of protein-coding genes in the cleaned and masked genome assembly of <em>Nucella lapillus</em> was performed using GALBA v1.0.11, an automated pipeline that uses proteins from a closely related species to assist in the training of gene prediction using AUGUSTUS. Proteins from <em>Rapana venosa</em> were provided for this purpose, and the miniprot option was used to perform the protein-to-genome alignments. Functional annotation of predicted protein-coding genes was performed using eggNOG-mapper&nbsp;v2.1.12, and additionally annotated with best hit BLAST results (v2.16.0) against the proteomes of the following marine gastropod species: <em>Rapana venosa</em>, <em>Littorina. saxatilis</em>,&nbsp; <em>Pomocea canaliculata, Stramonita haemastoma</em> and <em>Haliotis rufescens</em>.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

TWO DRAFT GENOMES OF FUNGAL leaf endophytes from tropical gymnosperms

<p>Two ascomycetes, <em>Neofusiccocum sp. </em>Z2&nbsp;and<em> Xylaria sp. </em>Z50<em>,&nbsp;</em>were isolated from healthy leaves of the tropical gymnosperms <em>Zamia pseudoparasitica</em> (Z2) and <em>Z. nana</em> (Z50) from Panama. The two draft genomes possess a broad repertoire of predicted carbohydrate degrading CAZymes, peptidases, secondary metabolites with more secondary metabolite clusters in the&nbsp;<em>Xylaria</em> isolate.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Draft genomes of vinasse bacteria

<p>Twenty-one draft bacterial genomes and their annotations described in the publication Cassman, NA. et. al. 2018, <em>Biotech for Biofuels</em>. The goal of the study was to characterize the microbial assemblage present in sugarcane vinasse, which is the&nbsp;major waste of bioethanol production from sugarcane and generally used as an organic and/or K fertilizer.&nbsp;Briefly, the draft genomes were binned (maxbin2) and&nbsp;manually refined (anvi&#39;o v.2.3.2) from a&nbsp;cross-assembly (metaSPADES v3.8.2)&nbsp;of 18 metagenomes sequenced&nbsp;from six batches of sugarcane vinasse with Illumina Miseq technology. Annotations (prokka v1.12) were carried out against prokka databases (UniProtKB Bacterial and the HAMAP HMM database) along with the dbCAN HMM database (dbCAN-fam-HMMs.v6). The uploaded files are from 1) prokka output: bin nucleotide sequences (fa), protein sequences (faa), annotations (gff and tsv) and annotation info (txt) and 2) anvi&#39;o output: percent recruitment of the bins, bin summary and samples summary.</p>

opencc-by-4.0Feb 2018View details →
zenodo40/100

Draft genome of a solitary bee, Osmia bicornis

<p>GFF3 formatted feature description&nbsp;file for the <em>Osmia </em><em>bicornis</em>&nbsp;genome (Genbank WGS accession number: MPJT00000000).&nbsp;</p>

opencc-by-4.0Jun 2018View details →
zenodo40/100

Draft genome assemblies of killifish from the Fundulus genus with ONT and Illumina sequencing platforms

<p>Four species from the genus Fundulus were selected for genome sequencing to study the physiological and genetic mechanisms that diverge between euryhaline and stenohaline freshwater species within this cyprinodontiform order of ray-finned fishes.</p>

opencc-by-4.0Jun 2019View details →
dryad40/100

Dataset for: The redlegged earth mite draft genome provides new insights into pesticide resistance evolution and demography in its invasive Australian range

<p>Data and analyses for Thia et al. "The redlegged earth mite draft genome provides new insights into pesticide resistance evolution and demography in its invasive Australian range" submitted to <em>Journal of Evolutionary Biology</em>.</p> <p>This repository comprises data and scripts used to replicate the analyses in this paper.</p> <p>The goals of this study were to: (1) assemble a draft reference genome for <em>Halotydeus destructor</em>; (2) perform a comparative analysis of acetylcholinesterase genes among different agricultural arthropod pests; (3) characterise the population genetic patterns among Australian <em>H. destructor</em> populations; and (4) perform demographic modelling to understand the evolutionary relationships between eastern and western populations of <em>H. destructor</em> in Australia.</p>

opencc-zeroMar 2023View details →
dryad40/100

Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi

Open the record for dataset details and reuse information.

publicJun 2021View details →
dryad40/100

Dataset for: The redlegged earth mite draft genome provides new insights into pesticide resistance evolution and demography in its invasive Australian range

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad36/100

Data from: An annotated draft genome of the mountain hare (Lepus timidus)

<p>Hares (genus Lepus) provide clear examples of repeated and often massive introgressive hybridization and striking local adaptations. Genomic studies on this group have so far relied on comparisons to the European rabbit (Oryctolagus cuniculus) reference genome. Here, we report the first de novo draft reference genome for a hare species, the mountain hare (Lepus timidus), and evaluate the efficacy of whole-genome re-sequencing analyses using the new reference versus using the rabbit reference genome. The genome was assembled using the ALLPATHS-LG protocol with a combination of overlapping pair and mate-pair Illumina sequencing (77x coverage). The assembly contained 32,294 scaffolds with a total length of 2.7 Gb and a scaffold N50 of 3.4 Mb. Re-scaffolding based on the rabbit reference reduced the total number of scaffolds to 4,205 with a scaffold N50 of 194 Mb. A correspondence was found between 22 of these hare scaffolds and the rabbit chromosomes, based on gene content and direct alignment. We annotated 24,578 protein coding genes by combining ab-initio predictions, homology search, and transcriptome data, of which 683 were solely derived from hare-specific transcriptome data. The hare reference genome is therefore a new resource to discover and investigate hare-specific variation. Similar estimates of heterozygosity and inferred demographic history profiles were obtained when mapping hare whole-genome re-sequencing data to the new hare draft genome or to alternative references based on the rabbit genome. Our results validate previous reference-based strategies and suggest that the chromosome-scale hare draft genome should enable chromosome-wide analyses and genome scans on hares.</p>

opencc-zeroOct 2020View details →
zenodo36/100

1263 Salmonella enterica draft genomes assembled from Bioproject PRJEB31846

<p>We assembled 1263&nbsp;Salmonella enterica draft genomes (raw data available from PRJEB31846).</p> <p>&nbsp;</p> <ul> <li>The dataset comprises diverse Salmonella enterica serovars collected between the years 1999 and 2019 and sequenced by the National Reference Laboratory for Salmonella on Illumina MiSeq and NextSeq technology. The data was described in more detail in &nbsp;10.1128/AEM.02265-19.</li> </ul> <ul> <li>Data were trimmed (with fastp, version 0.19.5) and assembled (with shovil-spades, version 1.1.0) using the AQUAMIS pipeline (https://gitlab.com/bfr_bioinformatics/AQUAMIS, version v1.2.0). All samples passed basic quality checks, such as sufficient base quality, coverage depth, genome length and contig number. Furthermore, no evidence for sample contamination was detected.</li> <li>The assemblies are input to a validation of chewieSnake (https://gitlab.com/bfr_bioinformatics/chewieSnake).</li> <li>The cgMLST analysis is available in https://bfr_bioinformatics.gitlab.io/chewiesnake_publicationdata/</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record