Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

167

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

167 results for “transcriptome assembly”

Learn how ShareScore rates datasets ↗
zenodo48/100

DATASET: De novo assembly and functional annotation of the heart + hemolymph transcriptome in the Caribbean spiny lobster Panulirus argus

<p>The spiny lobster <em>Panulirus argus</em> is an ecologically relevant species in shallow water coral reefs and target of the most lucrative fishery in the greater Caribbean region. This study reports, for the first time, the heart + hemolymph transcriptome of the Caribbean spiny lobster<em> Panulirus argus</em> assembled from short Illumina 150&thinsp;bp PE raw reads. A total 80,152,094 raw reads were assembled using the Oyster River Protocol pipeline that aspires to become the standard protocol for <em>de novo</em> transcriptome assembly. The assembly resulted in a total of 254,773 transcripts. Functional gene annotation was conducted using the software package &#39;dammit&#39; that also aspires to become the standard protocol for <em>de novo</em> transcriptome annotation. Lastly, gene enrichment analyses were conducted using the Gene Ontology (GO), KEGG pathway analyses (Kaas), and KOG (WebMGA) databases. This resource will be of utmost importance in future research aiming at exploring the effect of local and regional anthropogenic disturbances as well as global climate change on the molecular physiology of this overexploited species.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Data From: The Oyster River Protocol: A multi assembler and kmer approach for de novo transcriptome assembly.

<p>Characterizing transcriptomes in non-model organisms has resulted in a massive increase in our understanding of biological phenomena. This boon, largely made possible via high-throughput sequencing, means that studies of functional, evolutionary and population genomics are now being done by hundreds or even thousands of labs around the world. For many, these studies begin with a <em>de novo</em> transcriptome assembly, which is a technically complicated process involving several discrete steps. The Oyster River Protocol (ORP), described here, implements a standardized and benchmarked set of bioinformatic processes, resulting in an assembly with enhanced qualities over other standard assembly methods. Specifically, ORP produced assemblies have higher Detonate and TransRate scores and mapping rates, which is largely a product of the fact that it leverages a multi-assembler and kmer assembly process, thereby bypassing the shortcomings of any one approach. These improvements are important, as previously unassembled transcripts are included in ORP assemblies, resulting in a significant enhancement of the power of downstream analysis. Further, as part of this study, I show that assembly quality is unrelated with the number of reads generated, above 30 million reads. Code Availability: The version controlled open-source code is available at <a href="https://github.com/macmanes-lab/Oyster_River_Protocol">https://github.com/macmanes-lab/Oyster_River_Protocol</a>. Instructions for software installation and use, and other details are available at <a href="http://oyster-river-protocol.rtfd.org/">http://oyster-river-protocol.rtfd.org/</a>.</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

Plant regeneration in leaf culture of Centaurium erythraea Rafn. Part 3: de novo transcriptome assembly and validation of housekeeping genes for studies of in vitro morphogenesis

<p>Six centaury transcriptomes (embryogenic calli, globular somatic embryos, cotyledonary somatic embryos, adventitious buds, leaves and roots of <em>in vitro</em> grown plants) were sequenced and <em>de novo</em> assembled using <a href="https://github.com/trinityrnaseq/trinityrnaseq/wiki">Trinity</a> .</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly.tar.gz">CE_Assembly.tar.gz</a>&nbsp;- Centaury referent transcriptome comprises of 160.839 Trinity transcripts grouped in 105.726 Trinity genes.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly_fpkm.tar.gz">CE_Assembly_fpkm.tar.gz</a>&nbsp;- fpkm normalized read counts of the&nbsp;assembled transcripts in the six sequenced centaury tissues.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/nt.db_CE_assembly.tar.gz">nt.db_CE_assembly.tar.gz</a>&nbsp;-&nbsp;annotation of assembled transcripts by mapping them against NCBI nucleotide (NT) database&nbsp;using BLASTn . The obtained results were filtered with E-value E &le; 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">swissprot.db_CE_assembly.tar.gz</a>&nbsp;-&nbsp;annotation of assembled transcripts by mapping them against NCBI nucleotide (<a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">s</a>wissprot) database&nbsp;using BLASTx . The obtained results were filtered with E-value E &le; 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/pfam30.db_CE_assembly.tar.gz">pfam30.db_CE_assembly.tar.gz</a>&nbsp;-&nbsp;annotation of assembled transcripts by mapping them against Pfam30 domain database&nbsp;using hmmer3. The obtained results were filtered with independent E-value E &le; 10<sup>-3</sup>.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

De novo transcriptome assembly of hyperaccumulating Noccaea praecox

<p>Trinity de novo assembly for hyperaccumulating plant species Noccaea praecox (syn. Thlaspi praecox), Brassicaceae. The dataset&nbsp;includes&nbsp;annotations from SwissProt, Pfam, Rfam databases and information on transmembrane regions and signal peptide cleavage sites (annotated using BLAST, HMMER, infernal, tmHMM and signalP through Trinotate). Detailed information on the preprocessing, assembly, post-processing and annotations are described in the&nbsp;ReadMe file.</p> <p>Supplementary material for Bočaj, V., Pongrac, P., Fischer, S. <em>et al.</em> <em>De novo</em> transcriptome assembly of hyperaccumulating <em>Noccaea praecox</em> for gene discovery. <em>Sci Data</em> <strong>10</strong>, 856 (2023). <a href="https://doi.org/10.1038/s41597-023-02776-x">https://doi.org/10.1038/s41597-023-02776-x</a></p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Assembled transcriptomes of ovary, testis, and brain (male and female) of Amphibolurus muricatus (jacky dragon) generated using Trinity v2.11.0

<p><strong><em>A. muricatus</em> transcriptome assemblies generated using Trinity v2.11.0&nbsp;(Haas et al. 2013; Grabherr et al. 2011; Henschel et al. 2012)</strong><br> &bull; Amphibolurus-muricatus_brain.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> brain (male and female).<br> &bull; Amphibolurus-muricatus_combined.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> ovary, testis, and brain (male and female).<br> &bull; Amphibolurus-muricatus_female_brain.fa.tar.gz: Trinity assembly of female <em>A. muricatus</em> brain.<br> &bull; Amphibolurus-muricatus_male_brain.fa.tar.gz: Trinity assembly of male <em>A. muricatus</em> brain.<br> &bull; Amphibolurus-muricatus_ovary.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> ovary.<br> &bull; Amphibolurus-muricatus_testis.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> testis.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <ul> <li>Grabherr, M.G., B.J. Haas, M. Yassour, J.Z. Levin, D.A. Thompson et al., 2011 Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29 (7):644-652.</li> <li>Haas, B.J., A. Papanicolaou, M. Yassour, M. Grabherr, P.D. Blood et al., 2013 De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat Protoc 8 (8):1494-1512.</li> <li>Henschel, R., M. Lieber, L.-S. Wu, P.M. Nista, B.J. Haas et al., 2012 Trinity RNA-Seq assembler performance optimization, pp. 45 in Proceedings of the 1st Conference of the Extreme Science and Engineering Discovery Environment: Bridging from the eXtreme to the campus and beyond. Association for Computing Machinery, Chicag, IL, USA.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Transcriptome assemblies of three diatom and three prymnesiophyte isolates from Station ALOHA and Kaneohe Bay

<p><strong>Culture ID/name</strong></p> <p>AT125A &ndash; Pseudo-nitzschia sp.</p> <p>AT125C &ndash; Pseudo-nitzschia sp.</p> <p>ATCH2 &ndash; Chaetoceros sp.</p> <p>Pn B2 &ndash; Pseudo-nitzschia sp.</p> <p>KB-HA01 &ndash; Chrysochromulina sp. (also called&nbsp;&nbsp;</p> <p>AL-TEMP-12 &ndash; Chrysochromulina sp. (also called&nbsp;</p> <p>NF-H275 &ndash; Chrysochromulina sp.<br> <br> &nbsp;</p> <p><strong>Growth Conditions</strong></p> <p>All cultures were grown at 27&deg;C, 12:12 light:dark cycle, and with a light intensity of 100 &micro;mol photons m<sup>-2</sup>&nbsp;sec<sup>-1</sup>. AT125C, AT125A, and Pn B2 were grown with Aquil media. ATCH2 was grown with F/20 media with the phosphate concentration modified to a final concentration of 0.5&micro;M. KB-HA01 was grown with F/2 media and AL-TEMP-12 and NF-H275 were grown with K media. None of the cultures were axenic. All cultures were filtered in &ldquo;light&rdquo; and &ldquo;dark&rdquo; conditions and were in exponential phase when filtered. (Filter types and volumes filtered listed below.) After filtration, all filters were placed into 2mL screwcap tubes, flash frozen with liquid nitrogen, and stored at -80&deg;C.</p> <p><strong>Growth Conditions</strong></p> <p>All cultures were grown at 27&deg;C, 12:12 light:dark cycle, and with a light intensity of 100 &micro;mol photons m<sup>-2</sup>&nbsp;sec<sup>-1</sup>. AT125C, AT125A, and Pn B2 were grown with Aquil media. ATCH2 was grown with F/20 media with the phosphate concentration modified to a final concentration of 0.5&micro;M. KB-HA01 was grown with F/2 media and AL-TEMP-12 and NF-H275 were grown with K media. None of the cultures were axenic. All cultures were filtered in &ldquo;light&rdquo; and &ldquo;dark&rdquo; conditions and were in exponential phase when filtered. (Filter types and volumes filtered listed below.) After filtration, all filters were placed into 2mL screwcap tubes, flash frozen with liquid nitrogen, and stored at -80&deg;C.<br> <br> [TRANSCRIPTOME SEQUENCING]<br> <br> [QC AND ASSEMBLY]<br> <br> [POST-ASSEMBLY PROCESSING]<br> Diamond v2.0.5.143 was used to blast (e-value: 1e-5) to a cross-kingdom reference sequence database (as described in Coesel et al., 2021).&nbsp;Diamond v2.0.5.143 was used to find the least common ancestor of each contig based upon the&nbsp;blast results. Contigs that were identified as bacteria, archaea, or viruses were excluded&nbsp;from the assemblies.<br> <br> &nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Data from "Corset: enabling differential gene expression analysis for de novo assembled transcriptomes"

<p>This dataset contains de novo transcriptome assemblies&nbsp;for three publicly available RNA-seq dataset&nbsp;(SRA055442,&nbsp;SRR453566-SRR453571 and&nbsp;GSE37704&nbsp;). For each assembly we also provide a table with the&nbsp;read counts&nbsp;per&nbsp;contig, the output&nbsp;from corset (clusters and counts), and the results from&nbsp;a genome-based analysis. This dataset was used to assess the performance of the corset software. More detail is provided in the paper: Nadia M Davidson&nbsp;and&nbsp;Alicia Oshlack,<strong>&nbsp;</strong>Corset: enabling differential gene expression analysis for de novo assembled transcriptomes, <em>Genome&nbsp;Biology</em>&nbsp;2014,&nbsp;<strong>15</strong>:410.&nbsp;http://genomebiology.com/2014/15/7/410/abstract</p>

opencc-zeroAug 2014View details →
zenodo40/100

Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue for single nucleotide polymorphisms discovery (data set)

<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of <em>Festuca pratensis</em> and ten samples (corresponding to six genotypes) of <em>Lolium multiflorum</em> and fourteen samples of<em> Lolium perenne</em> (corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant <em>F. pratensis</em> transcripts were assembled by combining the contigs of all six genotypes based on orthology with the <em>Brachypodium distachyon </em>proteome. Similarly, <em>19,036</em> non-redundant<em> L. multiflorum</em> transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em> L. perenne</em> transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>

opencc-zeroFeb 2016View details →
zenodo40/100

Orthology guided transcriptome assembly of Italian ryegrass and meadow fescue (update data set)

<p>Transcriptome sequencing was performed on ten samples (corresponding to six genotypes) of&nbsp;<em>Festuca pratensis</em>&nbsp;and ten samples (corresponding to six genotypes) of&nbsp;<em>Lolium multiflorum</em>&nbsp;and fourteen samples of<em>&nbsp;Lolium perenne</em>&nbsp;(corresponding to fourteen genotypes). Using the OGA approach, 18,952 non-redundant&nbsp;<em>F. pratensis</em>&nbsp;transcripts were assembled by combining the contigs of all six genotypes based on orthology with the&nbsp;<em>Brachypodium distachyon&nbsp;</em>proteome. Similarly,&nbsp;<em>19,036</em>&nbsp;non-redundant<em>&nbsp;L. multiflorum</em>&nbsp;transcripts were assembled and annotated. In total, 17,455 orthologous transcripts were shared between the transcriptomes of the two species. Out of these, 16,613 orthologous transcripts overlap with the previously published<em>&nbsp;L. perenne</em>&nbsp;transcriptome containing 19,279 non-redundant transcripts(fasta files). We identified SNPs, the following criteria were used to classify it as one of following three classes (1) intraspecific SNPs (INTRA), (2) interspecific SNPs in two-way comparison (INTER-2W) and (3) interspecific SNPs in three-way comparison (INTER-3W) (GFF files).</p>

opencc-zeroJun 2016View details →
zenodo40/100

De novo transcriptome assemblies for the spiny mouse (Acomys cahirinus)

<p>Transcriptome assemblies generated per https://dx.doi.org/10.1101/076067 (preprint) / https://dx.doi.org/10.17504/protocols.io.ghebt3e (protocol). Manuscript available at Scientific Reports.</p>

opencc-by-4.0Jun 2017View details →
dryad40/100

Kellet's whelk genome and transcriptome assembly

<p>Understanding genomic characteristics of non-model organisms can help bridge gaps in ecology and evolutionary sciences, but lack of a reference genome and transcriptome for these species challenges their study. We advance this goal by conducting the first full genome and transcriptome sequence assembly and analysis of the non-model organism Kellet's whelk, <em>Kelletia kelletii</em>, a marine gastropod and fisheries species exhibiting a northern range expansion along the US west coast that is potentially driven by climate change. We used a combination of Oxford Nanopore Technologies, PacBio, and Illumina platforms for sequencing, and integrated a set of bioinformatic pipelines to create a comprehensive and contiguous de novo genome assembly. Our results represent the most complete and continuous documented genome among the <em>Buccinoidea</em> superfamily to date. Genome validation revealed its relatively high completeness with low missing metazoan BUSCOs, and an average coverage of ~70x for all contigs, indicating a robust assembly. Characteristics of the <em>K. kelletii</em> genome showed that short-read data contributed significantly to genome coverage and accuracy; however, long-read data was imperative to the completeness and continuity of the genome assembly. Genome annotation identified a large number of protein-coding genes compared to other closely related species, suggesting the presence of a complex genome structure. We conducted the transcriptome assembly and analysis of individuals during their period of peak embryonic development, and revealed highly expressed genes associated with specific GO terms and metabolic pathways, most notably lipid, carbohydrate, glycan, and phospholipid metabolism. We also identified numerous heat shock proteins (HSPs) in the transcriptome and genome with a potential association between the transcriptional expansion of HSP families and the marine environment experienced by the sessile life history stage of the developing embryo. This study offers a valuable reference genome and transcriptome for conducting comprehensive bioinformatic analyses of the non-model organism <em>K. kelletii</em>. Such resources will enhance our understanding of its ecology and evolution, as well as that of other coastal marine species facing environmental changes.</p>

opencc-zeroNov 2023View details →
dryad40/100

De novo transcriptome assembly and discovery of drought-responsive genes in eastern white spruce (Picea glauca)

<p>Forests face an escalating threat from the increasing frequency of extreme drought events driven by climate change. To address this challenge, it is crucial to understand how widely distributed species of economic or ecological importance may respond to drought stress. Here, we used RNA-sequencing to investigate transcriptome responses at increasing levels of water stress in white spruce (<em>Picea glauca</em> (Moench) Voss), distributed across North America. We began by generating an expanded transcriptome assembly emphasizing short-term drought stress at different developmental stages. We also analyzed differential gene expression at four time points over 22 days in a controlled drought stress experiment involving 2-year-old plants and three genetically unrelated clones. De novo transcriptome assembly and gene expression analysis revealed a total of 33,287 transcripts (18,934 annotated unique genes), with 4,425 unique drought-responsive genes. Many transcripts that had predicted functions associated with photosynthesis, cell wall organization, and water transport were down-regulated under drought conditions, while transcripts linked to abscisic acid response and defense response were up-regulated. Our study highlights a previously uncharacterized effect of drought stress on lipid metabolism genes in conifers and significant changes in the expression of several transcription factors, suggesting a regulatory response potentially linked to drought response or acclimation. Our research represents a fundamental step in unraveling the molecular mechanisms underlying short-term drought responses in white spruce seedlings. In addition, it provides a valuable source of new genetic data that could contribute to genetic selection strategies aimed at enhancing the drought resistance and resilience of white spruce to changing climates.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Simulation data for benchmarking de novo long read transcriptome assembly software

<p>Method of simulation of differentially expressed biological replicates</p> <p>We first obtained a subset of transcripts that are widely expressed in the GTEx v9 dataset (92 samples) using Gencode comprehensive annotation (v44). We kept transcripts with more than 5 reads in at least 15 samples after Salmon quantification (18145 genes, 40509 transcripts), and stored their mean count per million (CPM) values as the control group&rsquo;s baseline expression. We then generated a perturbed set of CPM values where transcript expression was changed by: (1) randomly selecting 1000 genes and changing all transcripts belonging to that gene concordantly (500 genes 2 fold up and 500 genes 2 fold down), (2) selected another 1000 genes randomly, and then select 2 random transcripts from the gene and swap their expression, (3) selected another 1000 genes randomly, and then select 1 random transcript to change its expression (500 transcripts 2 fold up and 500 transcripts 2 fold down). The updated CPM were stored as the perturbed group baseline expression. We then generated a count matrix and CPM matrix for 3 control replicates and 3 perturbed replicates with gamma distribution, followed by a Poisson distribution <a href="https://www.zotero.org/google-docs/?cUP4ui">(Baldoni et al., 2024)</a>. Both long-read and short-read FASTQ files were simulated using SQANTI-SIM with default settings and ONT R9.4 cDNA error profile (v 0.2.1) <a href="https://www.zotero.org/google-docs/?Qyopst">(Mestre-Tom&aacute;s et al., 2023)</a>. The long read data contained 6 million reads in total, and an average read length of 1085 bp, and short read data was 100 bp paired-end. We then subsampled the short-read data to match the total number of base pairs in the long read data (6.5 billion bases). The simulated data was non-stranded, and contains 2000 DE genes, 2000 genes with DTU, 5927 transcripts with DTU and 6933 DE transcripts.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo40/100

de novo transcriptome assemblies for Peperomia dahlstedtii and Peperomia pellucida

<p><em>De novo</em> transcriptome assemblies generated using Trinity using young leaf tissue from <em>Peperomia dahlstedtii</em> and <em>Peperomia pellucida</em>.&nbsp;</p>

opencc-by-3.0-usDec 2021View details →
zenodo40/100

Supplemental Results for Assembly, Annotation, and Analysis from HiFi reads of Gulf Toadfish Genome and Transcriptome fOpsBet2.1

<p>This repository contains gzipped tarballs of the results of the various assembly, annotation, and analysis steps performed during the assembly of the fOpsBet2.1 genome assembly for Opsanus beta at the University of Miami Rosenstiel School of Marine, Atmospheric, and Earth Science for the McDonald Toadfish Lab. These results are too numerous to include as supplemental data for a journal publication and so are available here for review. In this repository you will find results for:</p><p>Scripts:</p><p>-all bash and LSF scheduler job scripts used as part of the analysis, both exploratory and final.&nbsp;</p><p>QC:</p><p>-GenomeScope2 estimation of genome metrics from HiFi Reads</p><p>-QUAST genome statistics for each assembly step</p><p>-BUSCO completeness assessments for each assembly step&nbsp;</p><p>-inspector logs for polishing of initial assembly</p><p>-logs from Kraken2 contaminant screen</p><p>Assembly and Scaffolding:</p><p>-ntLINKS logs and intermediates for initial scaffolding</p><p>-ragtag logs and metrics for super-scaffolding to the ThaAma1.1 T. amazonica reference assembly</p><p>-mitoHIFI results for mitogenome assembly from HiFi reads, primary assembly, and purged alternate assembly</p><p>Annotation:</p><p>-PASA directory with full input and output for SQLite PASA assembly of transcriptome for gene predictors</p><p>-Results folder for Funannotate::annotate for gene models, annotations, and CDS/mRNA/protein fastas</p><p>-InterProSCan5 results for protein annotation used as input into Funannotate</p><p>-ghostKOALA KEGG assignment results for predicted proteins from funannotate results</p><p>Repetitive Elements:</p><p>-tidk telomere repeat analysis results</p><p>-TRAH satellite DNA analysis with subsequent analysis with HiCAT and StainedGlass</p><p>-RepeatModeler results for de novo TE prediction</p><p>-repclassifier results for TE curation</p><p>Comparative Analysis:</p><p>-OrthoFinder ortholog search for O. beta to several other vertebrates</p><p>-CAFE5 gene family expansion and contraction of Orthogroups from OrthoFinder results</p><p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Updated spiny mouse transcriptome assembly (now includes embryo-specific transcripts)

<p><strong>Summary</strong></p> <p>Updated spiny mouse transcriptome. Embryo-specific contigs generated from BioProject&nbsp;PRJNA436818&nbsp;were added to the Trinity_v2.3.2&nbsp;spiny mouse&nbsp;<em>de novo&nbsp;</em>transcriptome assembly (https://doi.org/10.5281/zenodo.808870).</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p>Embryos were collected from female spiny mice (n=12) in accordance with the Australian Code of Practice for the Care and Use of Animals for Scientific Purposes with approval from the Monash Medical Centre Animal Ethics Committee. Female dams were staged from delivery of their previous litter (spiny mice conceive their next litter approximately 12h postpartum) and culled at specific time-points for embryo retrieval at the required stage: 2-cell at 48h postpartum (n=4), 4-cell at 52h postpartum (&#39;early&#39; 4-cell; n=2) or at 68h postpartum (&#39;late 4-cell&#39;; n=2), and 8-cell at 72h postpartum (n=4). Embryos were snap frozen in cell lysis solution per&nbsp;the Nugen SoLo protocol (version M01406v3; available from NuGEN).&nbsp;After ligation of cDNA, qPCR was performed on all samples to determine the number of amplification cycles required to ensure that amplification was in the linear range. Based on these results, each sample was amplified using 24 cycles. Final libraries were quantitated by Qubit and size profile determined by the Agilent Bioanalyzer. All libraries were in the expected size range (~320-360 bp).&nbsp;Custom &#39;AnyDeplete&#39; rRNA depletion probes were designed and produced by NuGEN Technologies, Inc (San Carlos, CA, USA) using rRNA sequences from our reference transcriptome (Mamrot et al., 2017; https://doi.org/10.5281/zenodo.808870). Prior to use, efficacy and off-target effects of the rRNA depletion probes were examined <em>in silico</em> by NuGEN. Samples were loaded using c-Bot (200pM per library pool) and run on 2 lanes of an Illumina HiSeq 3000 8-lane flow-cell. PhiX spike-in was not used directly due to incompatibility with the custom rRNA depletion probes, however it was incorporated into other lanes of the same HiSeq 3000 run. RNA-Seq data (100bp, paired-end reads) are available from the NCBI as Bioproject PRJNA436818.</p> <p>The quality of RNA-Seq reads was assessed using FastQC v0.11.6 (<a href="https://github.com/s-andrews/FastQC">https://github.com/s-andrews/FastQC</a>; 50f0c26), with MultiQC v1.4 (<a href="https://github.com/ewels/MultiQC">https://github.com/ewels/MultiQC</a>; baefc2e) reports available from Github (<a href="https://github.com/jpmam1">https://github.com/jpmam1</a>) (Ewels et al., 2016). Adapter sequences were trimmed from the reads using trim-galore v0.4.2 (<a href="https://github.com/FelixKrueger/TrimGalore">https://github.com/FelixKrueger/TrimGalore</a>; d6b586e), implementing cutadapt v1.12 (<a href="https://github.com/marcelm/cutadapt">https://github.com/marcelm/cutadapt</a>; 98f0e2f). Reads with a quality scores lower than 20 and read pairs in which either forward or reverse reads were trimmed to fewer than 35 nucleotides were discarded. Further trimming of poor quality reads was conducted using Trimmomatic v0.36 (<a href="http://www.usadellab.org/cms/index.php?page=trimmomatic">http://www.usadellab.org/cms/index.php?page=trimmomatic</a>) with settings &quot;LEADING:3 TRAILING:3 SLIDINGWINDOW:4:20 AVGQUAL:25 MINLEN:35&quot; (Bolger et al., 2014). Nucleotides with quality scores lower than 3 were trimmed from the 3&rsquo; and 5&rsquo; read ends. Reads with an average quality score lower than 25 or with a length of fewer than 35 nucleotides after trimming were removed. Error correction of reads was performed using Rcorrector v1.0.2 (<a href="https://github.com/mourisl/Rcorrector">https://github.com/mourisl/Rcorrector</a>; 144602f) (Song et al., 2015). FastQC was used to assess the improvement in read quality after trimming adapter removal; MultiQC reports are available from Github (<a href="https://github.com/jpmam1">https://github.com/jpmam1</a>).</p> <p>Error corrected reads were assembled using Trinity v2.4.0 (<a href="https://github.com/trinityrnaseq/trinityrnaseq">https://github.com/trinityrnaseq/trinityrnaseq</a>; 1603d80) with settings &quot;--max_memory 400G, --CPU 32 and --full_cleanup&quot; (Haas et al., 2013). Assembly statistics were computed using the TrinityStats.pl from the Trinity package, and summary statistics are provided in Table S1. All reads were aligned to this transcriptome assembly using Bowtie2 v2.2.5 (<a href="https://github.com/BenLangmead/bowtie2">https://github.com/BenLangmead/bowtie2</a>; e718c6f) with settings: &quot;--end-to-end, --score-min L,-0.1,-0.1, --no-mixed, --no-discordant, -k 100, -X 1000, --time, -p 24&quot; (Langmead &amp; Salzberg, 2012).</p> <p>Read-supported contigs were identified within the embryo-specific Trinity <em>de novo </em>transcriptome assembly using samtools &quot;idxstats&quot; v1.5 (contigs with &gt;=1 reads aligning were retained) (<a href="https://github.com/samtools/samtools">https://github.com/samtools/samtools</a>; f510fb1) (Li et al., 2009). The read-supported contigs from the embryo-specific assembly (n=54,660) were added to the reference spiny mouse transcriptome assembly previously described (Mamrot, J., Legaie, R., Ellery, S.J., Wilson, T., Seemann, T., Powell, D.R., Gardner, D.K., Walker, D.W., Temple-Smith, P., Papenfuss, A.T. and Dickinson, H., 2017. De novo transcriptome assembly for the spiny mouse (Acomys cahirinus). Scientific Reports, 7(1), p.8996).</p> <p>The updated transcriptome is comprised of&nbsp;2,274,638 transcripts in total.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

De novo transcriptome assembly from the killifish, Fundulus rathbuni (gill epithelium)

<p>De novo transcriptome assembly from the killifish, Fundulus rathbuni. Fish were acclimated to either brackish or fresh water then exposed to an acute brackish water challenge. Transcriptome data from gill epithelium tissue were collected. A reference transcriptome assembly was&nbsp;generated from all individuals then used to analyze transcriptional responses to salinity.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Transcriptome assemblies of Thalassiosira hyalina and Nitzschia frigida

<p>Transcriptome assemblies of Thalassiosira hyalina and Nitzschia frigida originating from a time course experiment, in which these two species were exposed to high light stress and monitored over 120h under low and high pCO2. The corresponding Sequencing data is deposited at the EBI ArrayExpress database under accession number E-MTAB-6999. Contigs were created with Trinity Assembler.</p> <p>The according publication is currently in review (8/6/2019): Higher sensitivity towards light stress and ocean acidification in an Arctic sympagic compared to a pelagic diatom;</p> <p>Author team:&nbsp;Ane C. Kvernvik,&nbsp;Sebastian D. Rokitta, Eva Leu, Lars Harms, Tove M. Gabrielsen, Bj&ouml;rn Rost&nbsp;and Clara J. M. Hoppe</p> <p>Do not hesitate to contact the authors if you like more information!</p>

opencc-by-4.0Aug 2019View details →
zenodo40/100

Assemblies and annotation products of a transcriptome of regenerating and non-regenerating Lumbriculus variegatus worms

<p>These are the assemblies and annotation products of a transcriptome assembled from regenerating and non-regenerating tissues of the California blackworm <em>Lumbriculus variegatus</em>. The assembly and annotation strategies are described in the article <strong>Transcriptome analysis during early regeneration of&nbsp;<em>Lumbriculus variegatus</em></strong>, published in Gene Reports (https://doi.org/10.1016/j.genrep.2021.101050).</p> <p>blastp.outfmt6: homology results from BLASTp (v2.10.0) of the predicted proteins</p> <p>blastx.outfmt6: homology results from BLASTx (v2.10.0) of the assembled transcripts</p> <p>LV_transcriptome.fasta: assembled transcriptome after duplicated sequences were filtered out with CD-HIT-EST (v4.7)</p> <p>LV_transcriptome.fasta.transdecoder.cds: coding sequences identified with TransDecoder (v5.5.0)</p> <p>LV_transcriptome.fasta.transdecoder.gff3: positional annotation of the ORFs identified with TransDecoder (v5.5.0)</p> <p>LV_transcriptome.fasta.transdecoder.pep: predicted proteins identified with TransDecoder (v5.5.0)</p> <p>LV_transcriptome_non_filtered.fasta: transcriptome assembled with Trinity (Galaxy v0.0.1)</p> <p>signalp.out: signal peptide predictions from signalP (v4.1)</p> <p>tmhmm.out: transmembrane domains predictions from tmHMM (v2.0)</p> <p>TrinotatePFAM.out: protein domains predictions from HMMER (v3.3)</p>

opencc-by-4.0Feb 2021View details →
dryad40/100

De novo transcriptome assembly and discovery of drought-responsive genes in eastern white spruce (Picea glauca)

Open the record for dataset details and reuse information.

publicMar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record