Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
167
datasets available to search
ShareScore release 0.7.1
Dataset results
167 results for “transcriptome assembly”
Kellet's whelk genome and transcriptome assembly
Open the record for dataset details and reuse information.
Karyorelict ciliate transcriptome assemblies from project Karyocode
<p>Transcriptome assemblies of karyorelict ciliates, as described in the publication "Karyorelict ciliates use an ambiguous genetic code with context-dependent stop/sense codons" (https://doi.org/10.24072/pcjournal.141).</p> <p>Files are named with internal short names for each read library. The corresponding INSDC accessions of the read libraries are documented in the following spreadsheet:</p> <p>"Dataset accessions for comparative analysis of ciliate genetic codes", <a href="https://doi.org/10.17617/3.XWMBKT" target="_blank" rel="noopener">https://doi.org/10.17617/3.XWMBKT</a>, Edmond, V1; Table_S1.xlsx [fileName]</p> <p>Files with the suffix ".polyA_min7.fasta" contain the subset of assembled transcripts with a poly-A tail of at least 7 nt; this was done to filter out likely bacterial contaminants.</p>
De novo transcriptome assembly Aegilops cylindrica
<p>De novo transcriptome assembly of Aegilops cylindrica was created using trinity v. 2.15.1. The assembly contains 285000 transcripts encoding for 174040 genes.</p>
De novo transcriptome assembly of the rockrose Helianthemum marifolium
<p>Illumina paired-end RNA sequences from leaves of 16 individuals of Helianthemum marifolium (four individuals from each of the four recognised taxonomic subspecies) were cleaned and assembled de novo using Trinity and Oases. EvidentialGene provides the curated assembled transcriptome containing 122002 transcripts, which are contained in the file Helianthemum_marifolium_transcriptome.fa.</p> <p>The predicted peptide sequences obtained with TransDecoder are in the file Helianthemum_marifolium_fa_transdecoder.pep.</p> <p>Gene ontology annotations from Trinotate are in the file Helianthemum_marifolium_trinotate_annotation.xls.</p>
Genomic, transcriptomic and proteomic comparison of MRSA CC398 isolates collected from human and wild animal samples (Genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following methicillin-resistant <em>Staphylococcus aureus</em> (MRSA) strains: MRSA CC398 isolates recovered from humans, namely C5621 and C9017, and from a wild boar, namely OR418.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB35102).</p>
Hydractinia symbiolongicarpus genome and transcriptome assemblies
<p>Scaffold-level genome assemblies for Hydractinia symbiolongicarpus histoincompatible siblings BC-3 and BC-15 from a back-cross population derived from wildtype individuals . Assemblies generated with ABySS short read assembler using Illumina 200-bp insert paired-end libaries and 3-Kbp insert mate pair 'long jumping distance' libraries. The trimmed read depths were ~36X and ~49X for BC-3 and BC-15, respectively.</p> <p>Transcriptome assemblies for Hydractinia symbiolongicarpus wildtype individuals HWB-103 and HWB-29. Assemblies generated with Trinity using Illumina mRNA libraries.</p>
Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2
<p><span>Long-read sequencing technologies have improved significantly since their emergence. Their read lengths, potentially spanning entire transcripts, is advantageous for reconstructing transcriptomes. Existing long-read transcriptome assembly methods are primarily reference-based and to date, there is little focus on reference-free transcriptome assembly. We introduce RNA-Bloom2, a reference-free assembly method for long-read transcriptome sequencing data. </span>RNA-Bloom2 is available on GitHub at: <a href="https://github.com/bcgsc/RNA-Bloom">https://github.com/bcgsc/RNA-Bloom</a>.</p> <p><span>We benchmarked the assembly quality and the computational performance of RNA-Bloom2 on simulated data. We prepared two mouse simulated datasets with Trans-NanoSim</span><span> for the cDNA and dRNA sequencing protocols model on experimental ONT data</span><span>. The datasets were simulated </span><span>based on the mouse ENSEMBL annotation for GRCm39.</span><span> To investigate the effect of sequencing depth, we subsampled each dataset to 2, 10, and 18 million reads, resulting in a total of six sets of reads for our benchmarking experiments. Using the simulated data, w</span><span>e showed that the transcriptome assembly quality of RNA-Bloom2 is competitive to those of reference-based methods.</span></p>
Assembled "juvenile head region" transcriptome from the hagfish Eptatretus burgeri (NCBI GenBank Accession: SRX2541845)
<p>1) The raw reads were dwonloaded from NCBI GenBank (SRA run accession: (SRR5234495) with sratoolkit.</p> <p>2) The assembly was performed with Trinity with the folllowing parameters:</p> <p>Trinity --seqType fq --max_memory 100G --left eb-hf-head-region-SRR5234495/eb-hf-head-region-SRR5234495_1.fastq --right eb-hf-head-region-SRR5234495/eb-hf-head-region-SRR5234495_1.fastq --CPU 10 --output trinity-transcriptome-eb/</p> <p>3) ORFs were predeicted with TransDecoder with the following settings:</p> <p>TransDecoder.LongOrfs -t transcriptome-eb-hf-head-region.fasta</p> <p>TransDecoder.Predict -t transcriptome-eb-hf-head-region.fasta</p> <p> </p>
Error, noise and bias in de novo transcriptome assemblies
<p><i>De novo</i> transcriptome assembly is a powerful tool, widely used over the last decade for making evolutionary inferences. However, it relies on two implicit assumptions: that the assembled transcriptome is an unbiased representation of the underlying expressed transcriptome, and that expression estimates from the assembly are good, if noisy approximations of the relative abundance of expressed transcripts. Using publicly available data for model organisms, we demonstrate that, across assembly algorithms and data sets, these assumptions are consistently violated. Bias exists at the nucleotide level, with genotyping error rates ranging from 30-83%. As a result, diversity is underestimated in transcriptome assemblies, with consistent under-estimation of heterozygosity in all but the most inbred samples. Even at the gene level, expression estimates show wide deviations from map-to-reference estimates, and positive bias at lower expression levels. Standard filtering of transcriptome assemblies improves the robustness of gene expression estimates but leads to the loss of a meaningful number of protein-coding genes, including many that are highly expressed. We demonstrate a computational method, length-rescaled CPM, to partly alleviate noise and bias in expression estimates. Researchers should consider ways to minimize the impact of bias in transcriptome assemblies.</p>
Data for: Trinity assembled transcriptome of a Eurasian (Myriophyllum spicatum) and a hybrid (M. spicatum × M. sibiricum) genotype of watermilfoil
<p>Aquatic plant managers frequently treat Eurasian watermilfoil (<em>Myriophyllum spicatum</em> L.; EWM) and hybrid watermilfoil (<em>Myriophyllum spicatum</em> L. × <em>Myriophyllum sibiricum</em> Komarov) with 2,4-dichlorophenoxyacetic acid (2,4-D) herbicide. However, watermilfoil genotypes can differ in their response to 2,4-D. In this study, we compared facultative and constitutive gene expression differences for two watermilfoil genotypes (one Eurasian and one hybrid) that differ in their sensitivity to 2,4-D. To do this, we compared between control and 0.5mg L-1 2,4-D treated plants at four time points after treatment. We also assembled the first de novo watermilfoil transcriptome. We found that while qualitatively similar, the facultative transcriptional response of the EWM genotype to 2,4-D treatment was much stronger than the hybrid genotype, indicated by a greater number and log-fold-change of differentially expressed genes at all time points after treatment. Further, we found that the EWM and hybrid genotype differed in their 9-cis-epoxycarotenoid dioxygenase (NCED) and abscisic acid (ABA) gene response, and that there was a greater amount of photosynthesis gene downregulation (both in number and log-fold-change) in the EWM than the hybrid genotype. At the constitutive level, overall, the hybrid expressed genes at a higher level than the EWM genotype, but not the genes of the 2,4-D response pathway. These differences in gene expression match with the degree of phenotypic difference in growth observed between these genotypes when exposed to 2,4-D. The hybrid genotype used here mitigates the effects of 2,4-D treatment better than the EWM genotype at both the molecular and phenotypic level. More study is needed to understand the mechanism(s) of mitigation and whether this is a cause of hybridity, or the specific genotypic backgrounds used here.</p>
Reference transcriptome assembly of a protogynous sex change fish, harlequin sandsmelt (Parapercis pulchella)
<p>Reference transcriptome sequences (superTranscripts) of a marine teleost fish <em>Parapercis pulchella</em>.</p> <p>This dataset is a part of our work, "Reference transcriptome assembly of a protogynous sex change fish, harlequin sandsmelt (<em>Parapercis</em> pulchella)" published in <em>Marine Genomics</em>.</p> <p>https://doi.org/10.1016/j.margen.2024.101086</p> <p> If you use this dataset, plese cite the above paper.</p> <p>Raw RNA-seq data and <em>de novo </em>assembled sequences generated by Trinity have been deposited in NCBI/DDBJ/EMBL under accession PRJDB16534.</p> <p>This dataset is generated from Trinity raw-assembled sequences using superTranscripts method (Corset, Lace). </p> <p>Functional annotations were conducted using eggNog-mapper, KEEG Automatic Annotation Server (KAAS), and reciprocal BLAST best-hit analysis against medaka's protein sequences.</p> <p> </p> <p>The codes for generating these data are deposited in GitHub (<a href="https://github.com/yaoakifumi/Ppul-reference-transcriptome">https://github.com/yaoakifumi/Ppul-reference-transcriptome</a>).</p> <p> </p>
Data for: Trinity assembled transcriptome of a Eurasian (Myriophyllum spicatum) and a hybrid (M. spicatum × M. sibiricum) genotype of watermilfoil
Open the record for dataset details and reuse information.
Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2
Open the record for dataset details and reuse information.
Error, noise and bias in de novo transcriptome assemblies
Open the record for dataset details and reuse information.
Chromosome-level assembly of two pearl millet (Cenchrus americanus) genomes, functional annotation and transcriptomes
Open the record for dataset details and reuse information.
Chromosome-Level genome assembly and transcriptome analysis of the ural owl, Strix uralensis Pallas, 1771
Open the record for dataset details and reuse information.
Transcriptome assemblies associated with: A cnidarian phylogenomic tree fitted with hundreds of 18S leaves
Open the record for dataset details and reuse information.
Draft transcriptome assembly of Aipysurus laevis
<p><em>De novo</em> Trinity transcriptome assembly from six <em>Aipysurus laevis</em> tissues from multiple individuals. This data is complementary to DOI:10.5281/zenodo.3975254.</p>
Data from: The sugarcane mitochondrial genome: assembly, phylogenetics and transcriptomics
<p><strong>Background:</strong> Chloroplast genomes provide insufficient phylogenetic information to distinguish between closely related sugarcane cultivars, due to the recent origin of many cultivars and the conserved sequence of the chloroplast. In comparison, the mitochondrial genome of plants is much larger and more plastic and could contain increased phylogenetic signals. We assembled a consensus reference mitochondrion with Illumina TruSeq synthetic long reads and ONT MinION long reads. Based on this assembly we also analyzed the mitochondrial transcriptomes of sugarcane and sorghum and improved the annotation of the sugarcane mitochondrion as compared with other species.</p> <p><strong>Methods</strong>: Mitochondrial genomes were assembled from genomic read pools using a bait and assemble methodology. The mitogenome was exhaustively annotated using BLAST and transcript datasets were mapped with HISAT2 prior to analysis with the Integrated Genome Viewer.</p> <p><strong>Results</strong>: The sugarcane mitochondrion is comprised of two independent chromosomes, for which there is no evidence of recombination. Based on the reference assembly from the sugarcane cultivar SP80-3280 the mitogenomes of four additional cultivars (R570, LCP85-384, RB72343 and SP70-1143) were assembled (with the SP70-1143 assembly utilizing both genomic and transcriptomic data and the R570 data based on MinION assembly). We demonstrate that the sugarcane plastome is completely transcribed and we assembled the chloroplast genome of SP80-3280 using transcriptomic data only. Phylogenomic analysis using mitogenomes allow closely related sugarcane cultivars to be distinguished and supports the discrimination between <em>Saccharum officinarum</em> and <em>Saccharum cultum</em> as modern sugarcane's female parent. From whole chloroplast comparisons, we demonstrate that modern sugarcane arose from a limited number of S. cultum female founders. Transcriptomic and spliceosomal analyses reveal that the two chromosomes of the sugarcane mitochondrion are combined at the transcript level and that splice sites occur more frequently within gene coding regions than without. We reveal one confirmed and one potential cytoplasmic male sterility factor in the sugarcane mitochondrion, both of which are transcribed</p> <p><strong>Conclusion</strong>: Transcript processing in the sugarcane mitochondrion is highly complex with diverse splice events, the majority of which span the two chromosomes. PolyA baited transcripts are consistent with the use of polyadenylation for transcript degradation. For the first time we annotate two cytoplasmic male sterility factors within the sugarcane mitochondrion and demonstrate that sugarcane possesses all the molecular machinery required for cytoplasmic male sterility and rescue. A mechanism of cross-chromosomal splicing based on guide RNAs is proposed. We also demonstrate that mitogenomes can be used to perform phylogenomic studies on sugarcane cultivars.</p>
Data from: "De novo transcriptome assembly of the mountain fly Drosophila nigrosparsa using short RNA-seq reads" in Genomic Resources Notes Accepted 1 August 2014-30 September 2014
Drosophila (Drosophila) nigrosparsa is a habitat specialist restricted to the European montane/alpine zone (Bächli 2008). Mountain biodiversity is considered highly vulnerable to ongoing climate warming (IPCC 2013), and organisms at high altitudes have only limited possibility to shift to cooler habitats at elevations above (Pertoldi & Bach 2007). For such species, rapid evolution may offer a solution for long-term survival. We are establishing D. nigrosparsa as a model system to test the extent and tempo of adaptive evolution under thermal stress in the laboratory. In this study, we used Illumina high-throughput sequencing to assemble the species' transcriptome using the pooled mRNA from 22 developmental and physiological stages.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.