Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

181

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

181 results for “de novo assembly”

Learn how ShareScore rates datasets ↗
zenodo48/100

DATASET: De novo assembly and functional annotation of the heart + hemolymph transcriptome in the Caribbean spiny lobster Panulirus argus

<p>The spiny lobster <em>Panulirus argus</em> is an ecologically relevant species in shallow water coral reefs and target of the most lucrative fishery in the greater Caribbean region. This study reports, for the first time, the heart + hemolymph transcriptome of the Caribbean spiny lobster<em> Panulirus argus</em> assembled from short Illumina 150&thinsp;bp PE raw reads. A total 80,152,094 raw reads were assembled using the Oyster River Protocol pipeline that aspires to become the standard protocol for <em>de novo</em> transcriptome assembly. The assembly resulted in a total of 254,773 transcripts. Functional gene annotation was conducted using the software package &#39;dammit&#39; that also aspires to become the standard protocol for <em>de novo</em> transcriptome annotation. Lastly, gene enrichment analyses were conducted using the Gene Ontology (GO), KEGG pathway analyses (Kaas), and KOG (WebMGA) databases. This resource will be of utmost importance in future research aiming at exploring the effect of local and regional anthropogenic disturbances as well as global climate change on the molecular physiology of this overexploited species.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly

<p>Data and conda software environment file for the chapter &#39;<em>de novo</em> Genome Assembly&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Data From: The Oyster River Protocol: A multi assembler and kmer approach for de novo transcriptome assembly.

<p>Characterizing transcriptomes in non-model organisms has resulted in a massive increase in our understanding of biological phenomena. This boon, largely made possible via high-throughput sequencing, means that studies of functional, evolutionary and population genomics are now being done by hundreds or even thousands of labs around the world. For many, these studies begin with a <em>de novo</em> transcriptome assembly, which is a technically complicated process involving several discrete steps. The Oyster River Protocol (ORP), described here, implements a standardized and benchmarked set of bioinformatic processes, resulting in an assembly with enhanced qualities over other standard assembly methods. Specifically, ORP produced assemblies have higher Detonate and TransRate scores and mapping rates, which is largely a product of the fact that it leverages a multi-assembler and kmer assembly process, thereby bypassing the shortcomings of any one approach. These improvements are important, as previously unassembled transcripts are included in ORP assemblies, resulting in a significant enhancement of the power of downstream analysis. Further, as part of this study, I show that assembly quality is unrelated with the number of reads generated, above 30 million reads. Code Availability: The version controlled open-source code is available at <a href="https://github.com/macmanes-lab/Oyster_River_Protocol">https://github.com/macmanes-lab/Oyster_River_Protocol</a>. Instructions for software installation and use, and other details are available at <a href="http://oyster-river-protocol.rtfd.org/">http://oyster-river-protocol.rtfd.org/</a>.</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

Plant regeneration in leaf culture of Centaurium erythraea Rafn. Part 3: de novo transcriptome assembly and validation of housekeeping genes for studies of in vitro morphogenesis

<p>Six centaury transcriptomes (embryogenic calli, globular somatic embryos, cotyledonary somatic embryos, adventitious buds, leaves and roots of <em>in vitro</em> grown plants) were sequenced and <em>de novo</em> assembled using <a href="https://github.com/trinityrnaseq/trinityrnaseq/wiki">Trinity</a> .</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly.tar.gz">CE_Assembly.tar.gz</a>&nbsp;- Centaury referent transcriptome comprises of 160.839 Trinity transcripts grouped in 105.726 Trinity genes.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly_fpkm.tar.gz">CE_Assembly_fpkm.tar.gz</a>&nbsp;- fpkm normalized read counts of the&nbsp;assembled transcripts in the six sequenced centaury tissues.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/nt.db_CE_assembly.tar.gz">nt.db_CE_assembly.tar.gz</a>&nbsp;-&nbsp;annotation of assembled transcripts by mapping them against NCBI nucleotide (NT) database&nbsp;using BLASTn . The obtained results were filtered with E-value E &le; 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">swissprot.db_CE_assembly.tar.gz</a>&nbsp;-&nbsp;annotation of assembled transcripts by mapping them against NCBI nucleotide (<a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">s</a>wissprot) database&nbsp;using BLASTx . The obtained results were filtered with E-value E &le; 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/pfam30.db_CE_assembly.tar.gz">pfam30.db_CE_assembly.tar.gz</a>&nbsp;-&nbsp;annotation of assembled transcripts by mapping them against Pfam30 domain database&nbsp;using hmmer3. The obtained results were filtered with independent E-value E &le; 10<sup>-3</sup>.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

De novo transcriptome assembly of hyperaccumulating Noccaea praecox

<p>Trinity de novo assembly for hyperaccumulating plant species Noccaea praecox (syn. Thlaspi praecox), Brassicaceae. The dataset&nbsp;includes&nbsp;annotations from SwissProt, Pfam, Rfam databases and information on transmembrane regions and signal peptide cleavage sites (annotated using BLAST, HMMER, infernal, tmHMM and signalP through Trinotate). Detailed information on the preprocessing, assembly, post-processing and annotations are described in the&nbsp;ReadMe file.</p> <p>Supplementary material for Bočaj, V., Pongrac, P., Fischer, S. <em>et al.</em> <em>De novo</em> transcriptome assembly of hyperaccumulating <em>Noccaea praecox</em> for gene discovery. <em>Sci Data</em> <strong>10</strong>, 856 (2023). <a href="https://doi.org/10.1038/s41597-023-02776-x">https://doi.org/10.1038/s41597-023-02776-x</a></p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

De novo genome assembly of the meadow brown butterfly, Maniola jurtina

<p>1. Whole-genome GFF file (raw and filtered for min. gene length) [<em>Maniola.jurtina.gff3</em>, <em>Maniola_jurtina_filtered.gff3</em>]</p> <p>2. List of <em>M. jurtina</em> proteins [<em>Mjurtina_proteins.fa</em>].</p> <p>3. Results of spot pattern genes BLAST&nbsp;against <em>M. jurtina</em> proteome [<em>Lepidoptera_MJ_protein_matches.xlsx</em>].&nbsp;</p> <p>4. Annotations [blast2go_export.txt]</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Draft de novo genome assemblies of a male and female Amphibolurus muricatus (jacky dragon)

<p>Four de novo nuclear genome assemblies of <em>Amphibolurus muricatus</em></p> <p><strong>Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> &bull; AmpMurF_1.0.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_1.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 1.1: Further scaffolding of assembly 1.0 using RNA-seq data</strong><br> &bull; AmpMurF_1.1.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_1.1.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 2.0: Further scaffolding of assembly 1.0 using SLR-superscaffolder</strong><br> &bull; AmpMurF_2.0.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_2.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing assembly</strong><br> &bull; AmpMurF_3.0.fa.tar.gz (female <em>A. muricatus</em>)<br> &bull; AmpMurM_3.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Methods<br> Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> Male and female <em>A. muricatus</em> genome sequencing libraries were constructed on the Chromium system (10x Genomics, Pleasanton, CA, USA) by the Ramaciotti Centre for Genomics (Sydney, Australia). The Chromium instrument enables unique barcoding of long stretches of DNA on gel beads. The barcodes allow later reconstruction of long DNA fragments from a series of short DNA fragments with the same barcode (i.e., linked-reads). After barcoding, DNA was sheared into smaller fragments and sequenced on the NovaSeq 6000 platform (Illumina, CA, USA) to generate 151 bp paired-end (PE) reads. A total of 904.9 M raw 10x Genomics Chromium linked-reads were generated. Raw 10x data were assembled with Supernova v2.1.1 (Weisenfeld et al., 2017) and a FASTA file was generated using the &lsquo;pseudohap style&rsquo; option in Supernova mkoutput. All female (~450 M) and male (~550 M) read pairs were utilised (female sequencing depth ca 50.3&times;; male, ca 47.8&times;). The resulting assemblies was further scaffolded with ARKS v1.0.3 (Coombe et al., 2018), reusing the 10x reads, and the companion LINKS program (v1.8.7) (Warren et al., 2015). ARKS employs a <em>k</em>-mer approach to map linked barcodes to the contigs in the initial Supernova assembly to generate a scaffold graph with estimated distances for LINKS input. These assemblies were denoted AmpMurF_1.0 (female) and AmpMurM_1.0 (male). We used GapCloser v1.12 (part of SOAPdenovo2) (Luo et al., 2012) to fill gaps in the assembly. GapCloser was run using the parameter -l 150) and clean&nbsp;10x Genomics reads PE reads. &nbsp;</p> <p><strong>Assembly 1.1: Further scaffolding using RNA-seq data</strong><br> We attempted to improve the v1.0 genome assemblies&rsquo; contiguity using RNA-sequencing reads. RNA-seq reads (from brain, ovary, and testis; see below) were filtered (i.e., cleaned) to remove adapters and low-quality reads using Flexbar v3.4.0 and used to further re-scaffold the v1.0 assemblies (FASTA files before gapclosing) with P_RNA_scaffolder (Zhu et al., 2018). The default Flexbar settings discards all reads with any uncalled bases. A final round of scaffolding was performed on the resulting assemblies using L_RNA_scaffolder (Xue et al., 2013). These assemblies were denoted AmpMurF_1.1 (female) and AmpMurM_1.1 (male). As before, GapCloser and clean&nbsp;10x Genomics reads were used to fill gaps. &nbsp;&nbsp; &nbsp;</p> <p><strong>Assembly 2.0: Further scaffolding using SLR-superscaffolder</strong><br> As an alternative approach, we attempted to improve the v1.0 genome assemblies&rsquo; contiguity using SLR-superscaffolder (Guo et al., 2021). Briefly, SLR-superscaffolder employs single tube long fragment read (stLFR) sequencing (Wang et al., 2019) reads (see section below) to generate hybrid genome assemblies. The software was run with default parameters except for PE_SEED_MIN=300 (minimum contig size to fill; default 1000). These assemblies were denoted AmpMurF_2.0 (female) and AmpMurM_2.0 (male). GapCloser and clean&nbsp;stLFR reads (with the barcode removed using https://github.com/BGI-Qingdao/stLFR_barcode_split) were used to fill gaps. &nbsp;&nbsp; &nbsp;</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing and supernova assembly</strong><br> We also generated independent assemblies for the individuals sequenced on the 10x Genomics Chromium system using single tube long fragment read (stLFR) sequencing (Wang et al., 2019). BGI (Brisbane, Australia) generated ~100&times;-coverage 100-bp paired-end reads (plus a 42-bp stLFR barcode on the right/_2 read). Low-quality reads, PCR duplicates, and adaptors were removed using SOAPnuke v1.5&nbsp;(Chen et al. 2018). The stLFRdenovo pipeline (<a href="https://github.com/BGI-biotools/stLFRdenovo">https://github.com/BGI-biotools/stLFRdenovo</a>), which is based on Supernova and customized for stLFR data, was used to generate a&nbsp;<em>de novo</em>&nbsp;genome assembly. The stLFRdenovo tool &lsquo;FillGaps&rsquo; was used to fill gaps.</p> <p><strong>References</strong><br> Chen, Y., Chen, Y., Shi, C., Huang, Z., Zhang, Y., Li, S., Li, Y., Ye, J., Yu, C., Li, Z., et al. (2018). SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 7, 1-6.<br> Coombe, L., Zhang, J., Vandervalk, B.P., Chu, J., Jackman, S.D., Birol, I., and Warren, R.L. (2018). ARKS: chromosome-scale scaffolding of human genome drafts with linked read kmers. BMC Bioinformatics 19, 234.<br> Guo, L., Xu, M., Wang, W., Gu, S., Zhao, X., Chen, F., Wang, O., Xu, X., Seim, I., Fan, G., et al. (2021). SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme. BMC Bioinformatics 22, 158.<br> Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18.<br> Wang, O., Chin, R., Cheng, X., Wu, M.K.Y., Mao, Q., Tang, J., Sun, Y., Anderson, E., Lam, H.K., Chen, D., et al. (2019). Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly. Genome Res 29, 798-808.<br> Warren, R.L., Yang, C., Vandervalk, B.P., Behsaz, B., Lagman, A., Jones, S.J., and Birol, I. (2015). LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads. Gigascience 4, 35.<br> Weisenfeld, N.I., Kumar, V., Shah, P., Church, D.M., and Jaffe, D.B. (2017). Direct determination of diploid genome sequences. Genome Res 27, 757-767.<br> Xue, W., Li, J.T., Zhu, Y.P., Hou, G.Y., Kong, X.F., Kuang, Y.Y., and Sun, X.W. (2013). L_RNA_scaffolder: scaffolding genomes with transcripts. BMC Genomics 14, 604.<br> Zhu, B.H., Xiao, J., Xue, W., Xu, G.C., Sun, M.Y., and Li, J.T. (2018). P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads. BMC Genomics 19, 175.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""

<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. rufipogon</em>&nbsp; DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. glaberrima</em>&nbsp; IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. barthii</em>&nbsp; W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. nivara</em>&nbsp; W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;&nbsp;W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;&nbsp;W2014_O.nivara_scaffolded_anchored.fa.</p>

opencc-by-4.0Feb 2020View details →
dryad40/100

Draft de novo genome assembly of the elusive jaguarundi, Puma yagouaroundi

<p>The Puma lineage within the family Felidae consists of three species that last shared a common ancestor around 4.9 million years ago. Whole-genome sequences of two species from the lineage were previously reported: the cheetah (<em>Acinonyx jubatus</em>) and the mountain lion (<em>Puma concolor</em>). The present report describes a whole-genome assembly of the remaining species, the jaguarundi (<em>Puma yagouaroundi</em>). We sequenced the genome of a male jaguarundi with 10X Genomics linked reads and assembled the whole-genome sequence. The assembled genome contains a series of scaffolds that reach the length of chromosome arms and is similar in scaffold contiguity to the genome assemblies of cheetah and puma, with a contig N50 = 100.2 kbp and a scaffold N50 = 49.27 Mbp. We assessed the assembled sequence of the jaguarundi genome using BUSCO, aligned reads of the sequenced individual and another published female jaguarundi to the assembled genome, annotated protein-coding genes, repeats, genomic variants and their effects with respect to the protein-coding genes, and analyzed differences of the two jaguarundis from the reference mitochondrial genome. The jaguarundi genome assembly and its annotation were compared in quality, variants and features to the previously reported genome assemblies of puma and cheetah. Computational analyzes used in the study were implemented in transparent and reproducible way to allow their further reuse and modification.</p>

opencc-zeroJun 2021View details →
zenodo40/100

Data from "Corset: enabling differential gene expression analysis for de novo assembled transcriptomes"

<p>This dataset contains de novo transcriptome assemblies&nbsp;for three publicly available RNA-seq dataset&nbsp;(SRA055442,&nbsp;SRR453566-SRR453571 and&nbsp;GSE37704&nbsp;). For each assembly we also provide a table with the&nbsp;read counts&nbsp;per&nbsp;contig, the output&nbsp;from corset (clusters and counts), and the results from&nbsp;a genome-based analysis. This dataset was used to assess the performance of the corset software. More detail is provided in the paper: Nadia M Davidson&nbsp;and&nbsp;Alicia Oshlack,<strong>&nbsp;</strong>Corset: enabling differential gene expression analysis for de novo assembled transcriptomes, <em>Genome&nbsp;Biology</em>&nbsp;2014,&nbsp;<strong>15</strong>:410.&nbsp;http://genomebiology.com/2014/15/7/410/abstract</p>

opencc-zeroAug 2014View details →
zenodo40/100

De novo transcriptome assemblies for the spiny mouse (Acomys cahirinus)

<p>Transcriptome assemblies generated per https://dx.doi.org/10.1101/076067 (preprint) / https://dx.doi.org/10.17504/protocols.io.ghebt3e (protocol). Manuscript available at Scientific Reports.</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

PacBio HiFi de-novo assembled genome and mitochondrial genome for Orbicella faveolata

<p>Final assembly using Funannotate of <i>Orbicella faveolata</i> from PacBio HiFi reads. For full methods please see the publication.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
dryad40/100

De novo transcriptome assembly and discovery of drought-responsive genes in eastern white spruce (Picea glauca)

<p>Forests face an escalating threat from the increasing frequency of extreme drought events driven by climate change. To address this challenge, it is crucial to understand how widely distributed species of economic or ecological importance may respond to drought stress. Here, we used RNA-sequencing to investigate transcriptome responses at increasing levels of water stress in white spruce (<em>Picea glauca</em> (Moench) Voss), distributed across North America. We began by generating an expanded transcriptome assembly emphasizing short-term drought stress at different developmental stages. We also analyzed differential gene expression at four time points over 22 days in a controlled drought stress experiment involving 2-year-old plants and three genetically unrelated clones. De novo transcriptome assembly and gene expression analysis revealed a total of 33,287 transcripts (18,934 annotated unique genes), with 4,425 unique drought-responsive genes. Many transcripts that had predicted functions associated with photosynthesis, cell wall organization, and water transport were down-regulated under drought conditions, while transcripts linked to abscisic acid response and defense response were up-regulated. Our study highlights a previously uncharacterized effect of drought stress on lipid metabolism genes in conifers and significant changes in the expression of several transcription factors, suggesting a regulatory response potentially linked to drought response or acclimation. Our research represents a fundamental step in unraveling the molecular mechanisms underlying short-term drought responses in white spruce seedlings. In addition, it provides a valuable source of new genetic data that could contribute to genetic selection strategies aimed at enhancing the drought resistance and resilience of white spruce to changing climates.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Simulation data for benchmarking de novo long read transcriptome assembly software

<p>Method of simulation of differentially expressed biological replicates</p> <p>We first obtained a subset of transcripts that are widely expressed in the GTEx v9 dataset (92 samples) using Gencode comprehensive annotation (v44). We kept transcripts with more than 5 reads in at least 15 samples after Salmon quantification (18145 genes, 40509 transcripts), and stored their mean count per million (CPM) values as the control group&rsquo;s baseline expression. We then generated a perturbed set of CPM values where transcript expression was changed by: (1) randomly selecting 1000 genes and changing all transcripts belonging to that gene concordantly (500 genes 2 fold up and 500 genes 2 fold down), (2) selected another 1000 genes randomly, and then select 2 random transcripts from the gene and swap their expression, (3) selected another 1000 genes randomly, and then select 1 random transcript to change its expression (500 transcripts 2 fold up and 500 transcripts 2 fold down). The updated CPM were stored as the perturbed group baseline expression. We then generated a count matrix and CPM matrix for 3 control replicates and 3 perturbed replicates with gamma distribution, followed by a Poisson distribution <a href="https://www.zotero.org/google-docs/?cUP4ui">(Baldoni et al., 2024)</a>. Both long-read and short-read FASTQ files were simulated using SQANTI-SIM with default settings and ONT R9.4 cDNA error profile (v 0.2.1) <a href="https://www.zotero.org/google-docs/?Qyopst">(Mestre-Tom&aacute;s et al., 2023)</a>. The long read data contained 6 million reads in total, and an average read length of 1085 bp, and short read data was 100 bp paired-end. We then subsampled the short-read data to match the total number of base pairs in the long read data (6.5 billion bases). The simulated data was non-stranded, and contains 2000 DE genes, 2000 genes with DTU, 5927 transcripts with DTU and 6933 DE transcripts.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo40/100

de novo transcriptome assemblies for Peperomia dahlstedtii and Peperomia pellucida

<p><em>De novo</em> transcriptome assemblies generated using Trinity using young leaf tissue from <em>Peperomia dahlstedtii</em> and <em>Peperomia pellucida</em>.&nbsp;</p>

opencc-by-3.0-usDec 2021View details →
zenodo40/100

De novo transcriptome assembly from the killifish, Fundulus rathbuni (gill epithelium)

<p>De novo transcriptome assembly from the killifish, Fundulus rathbuni. Fish were acclimated to either brackish or fresh water then exposed to an acute brackish water challenge. Transcriptome data from gill epithelium tissue were collected. A reference transcriptome assembly was&nbsp;generated from all individuals then used to analyze transcriptional responses to salinity.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

De novo whole genome assembly of the giant tiger prawn (Penaeus monodon) from Vietnam

<p>Basecalled Nanopore FastQ files for the Vietnamese giant tiger prawn and its genome assemblies.</p> <p>FastQ files (LSK109 sample prep sequenced on a MinION device for 48 hours). Read stats are in *_stat.txt:</p> <p>TP_A.fastq : gDNA was extracted using Zymo quick DNA minikit from ethanol-preserved muscle tissue</p> <p>TP_B.fastq : Same as TP_A.fastq</p> <p>TP_C.fastq : gDNA was extracted using conventional salting out method (longer read length but reduced yield)</p> <p>Assemblies:</p> <p>v1_MaSuRCA.fasta: Assembly using poly-G trimmed Illumina reads</p> <p>v2_NanoporeScaf.fasta : Scaffolding with Nanopore long reads</p> <p>v3_RNA_NanoporeScaf.fasta: Scaffolding of v2 with RNA reads</p> <p>v4_NCBI_Filt.fasta: post NCBI contaminant and carry-over adapter removal (final version)</p> <p>Annotation:</p> <p>Braker2_annotation.gff3.gz: Inititial Braker2 gff3 output</p> <p>Braker2_CDS.fna.gz: Initial Braker2 predicted genes</p> <p>Braker2_prot.faa.gz: Protein translation of Braker2_CDS.fna</p> <p>CAZy.tar.gz: CAZy annotation for four crustacean species</p> <p>Filtered_Gene.tar.gz: List of genes with functional annotation and/or orthologs</p> <p>Interproscan_result.tsv: Raw InterProScan output</p> <p>OrthoFinder2.tar.gz: OrthoFinder2 output. Proteins used to infer orthologs are included as &quot;.faa&quot;.</p>

opencc-by-4.0Dec 2018View details →
dryad40/100

Data from: A de novo chromosome-level genome assembly of Coregonus sp. "Balchen": one representative of the Swiss Alpine whitefish radiation

<p>Salmonids are of particular interest to evolutionary biologists due to their incredible diversity of life-history strategies and the speed at which many salmonid species have diversified. In Switzerland alone, over 30 species of Alpine whitefish from the subfamily Coregoninae have evolved since the last glacial maximum, with species exhibiting a diverse range of morphological and behavioural phenotypes. This, combined with the whole genome duplication which occurred in the ancestor of all salmonids, makes the Alpine whitefish radiation a particularly interesting system in which to study the genetic basis of adaptation and speciation and the impacts of ploidy changes and subsequent rediploidization on genome evolution. Although well curated genome assemblies exist for many species within Salmonidae, genomic resources for the subfamily Coregoninae are lacking. To assemble a whitefish reference genome, we carried out PacBio sequencing from one wild-caught <i>Coregonus sp. "Balchen" </i>from Lake Thun to ~90x coverage. PacBio reads were assembled independently using three different assemblers, Falcon, Canu and wtdbg2 and subsequently scaffolded with additional Hi-C data. All three assemblies were highly contiguous, had strong synteny to a previously published <i>Coregonus</i>linkage map, and when mapping additional short-read data to each of the assemblies, coverage was fairly even across most chromosome-scale scaffolds. Here, we present the first <i>de novo</i>genome assembly for the Salmonid subfamily Coregoninae. The final 2.2 Gb wtdbg2 assembly included 40 scaffolds, an N50 of 51.9 Mb, and was 93.3% complete for BUSCOs. The assembly consisted of ~52% TEs and contained 44,525 genes.</p>

opencc-zeroMay 2020View details →
zenodo40/100

Baseline assemblies for "ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads" protocol

<p>ntLink is a flexible&nbsp;<em>de novo</em> genome scaffolding toolkit which can be run in various modes depending on the desired user output, with multiple new functionalities recently introduced. Here, we provide the baseline assembly datasets used in the ntLink protocol paper &quot;ntLink: a toolkit for <em>de novo </em>genome assembly scaffolding and mapping using long reads&quot;. The provided assemblies are ABySS (short-read) and Flye (long-read) assemblies of&nbsp;<em>Caenorhabditis elegans&nbsp;</em>genome sequencing data. The ABySS (v2.1.4) assembly utilized paired-end short reads (accession&nbsp;DRR008444), and was run with the following parameters:&nbsp;k=64&nbsp;l=40 s=1000&nbsp;q=15 B=10G j=8&nbsp;kc=3&nbsp;H=4 S=1000-10000 N=9.The&nbsp;<em>C. elegans</em>&nbsp;Flye (v2.5) assembly was run using Oxford Nanopore long reads (accession SRR10028109) and the following parameters:&nbsp;--nano-raw SRR10028109.fastq&nbsp;-g100m -t48.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

De novo assembly of a long-read Amblyomma americanum genome

<p>Genome assemblies of Amblyomma americanum generated from PacBio HiFi sequencing of 50 individual female ticks. This repository contains the unphased diploid assembly generated by the Flye assembler (Arcadia_Amblyomma_americanum_asm001.fasta). In addition, there are two associated fasta files containing sequences generated by submitting the unphased diploid assembly to separation by the Purge_Dups pipeline&nbsp;(purged pseudo-haploid assembly and haplotig assembly).</p> <p>Flye assembler:&nbsp;https://github.com/fenderglass/Flye</p> <p>Purge_Dups pipeline:&nbsp;https://github.com/dfguan/purge_dups</p> <p>NCBI Bioproject:&nbsp;PRJNA932813</p>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record