Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
67
datasets available to search
ShareScore release 0.9.0
Dataset results
67 results for “transcriptome reference”
BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq
<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts – BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5’ and 3’ UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>
(Annotation and metric) Thalassiosirales reference transcriptomes
<p>Here are deposited the public data recompiled for the construction of a Thalassiosirales reference database as part of a Ph.D Thesis "TEMPERATURE ACCLIMATION CAPACITY AND COLD-ADAPTATION MECHANIMS IN THALASSIOSIRALES ANTARCTIC MEMBERS". Data correspond to the annotation of 53 transcriptomes from the MMETSP of Thalassiosirales members.</p>
Supplementary tables for publication "A reference-free algorithm discovers regulation in the plant transcriptome"
<p>Supplementary tables for publication "A reference-free algorithm discovers regulation in the plant transcriptome" (doi: https://doi.org/10.1101/2024.05.23.595613)</p> <p>Table A: complete list of significant anchors and associated genes from analysis of sorghum dataset</p> <p>Table B: complete list of significant anchors and associated genes from analysis of maize dataset</p> <p>Table C: complete list of significant anchors and associated genes from analysis of Arabidopsis P/Fe dataset</p> <p>Table D: complete list of significant anchors and associated genes from analysis of Arabidopsis FLOE1 dataset</p> <p>arabidopsis_floe1_ALL_anchors_satc_truncated.txt: data from the Arabidopsis FLOE1 dataset used to generate figures in the paper. Columns are sample ID, anchor, target, and counts of that anchor/target combination in that particular sample. </p> <p>arabidopsis_pfe_ALL_anchors_satc_truncated.txt: data from the Arabidopsis P/Fe dataset used to generate figures in the paper. Columns are sample ID, anchor, target, and counts of that anchor/target combination in that particular sample. </p> <p>maize_pollen_ALL_anchors_satc_truncated.txt: data from the maize dataset used to generate figures in the paper. Columns are sample ID, anchor, target, and counts of that anchor/target combination in that particular sample. </p> <p>sorghum_drought_ALL_anchors_satc_truncated.txt: data from the sorghum dataset used to generate figures in the paper. Columns are sample ID, anchor, target, and counts of that anchor/target combination in that particular sample.</p> <p>cryptic_splicing_anchors.tsv: list of anchors described in Supplementary Information section of the article that are examples of cryptic splicing. Columns are dataset name, gene name/ID, anchor sequence, target 1 sequence, and target 2 sequence. </p>
Neurogenomic divergence during speciation by reinforcement of mating behaviors in chorus frogs (Pseudacris) – De novo reference transcriptome: Assemblerd contigs and gene annotations
<p>Assembled contigs (Trinity) and gene annotations (Trinotate) of a reference transcriptome for the Upland Chorus Frog, <em>Pseudacris feriarum</em>. Data to assemble the contigs were obtained by sequencing four tissue types: Brain, eyes, testis, and somatic (liver/heart/lung/skin/muscle). Raw reads are stored in the NCBI-SRA database (BioProject PRJNA723357).</p>
Chaetoceros decipiens (UNC1416) reference transcriptome
<p>Reference transcriptome and associated annotations for <em>Chaetoceros decipiens </em>(UNC1416). </p> <p>A culture was grown into late exponential phase for filtration. Total RNA was extracted using the RNAqueous-4PCR Total RNA Isolation Kit (Ambion, Foster City, CA, USA) according to the manufacturer’s protocol with an initial bead beating step to disrupt cells. RNA libraries were created with either the Illumina TruSeq Stranded mRNA Library Preparation Kit. The library was sequenced on an Illumina MiSeq (300 bp, paired-end reads) and an Illumina HiSeq 2500 with one lane in high output mode (100 bp, paired-end reads) and another lane in rapid run mode (150 bp, paired-end reads).</p> <p>Raw reads were trimmed for quality with Trimmomatic v0.36 then assembled <em>de novo </em>with Trinity v2.5.1 with the default parameters for paired-reads and a minimum contig length of 90 bp. Contigs were clustered based on 99% similarity using CD-HIT-EST v4.7 and then protein sequences were predicted with GeneMark S-T. Protein sequences were annotated by best-homology (lowest E-value) with the KEGG (Release 86.0), UniProt (Release 2018_03), and PhyloDB (v1.076) databases via BLASTP v2.7.1 (E-value ≤ 10<sup>-5</sup>) and with Pfam 31.0 via HMMER v3.1b2 (Dataset S2). KEGG Ortholog (KO) annotations were assigned from the top hit with a KO annotation from the top 10 hits (<a href="https://github.com/ctberthiaume/keggannot">https://github.com/ctberthiaume/keggannot</a>).</p> <p>Provided here are predicted proteins as nucleotides and peptides. Raw reads are deposited in SRA (SRP234548).</p>
Neurogenomic divergence during speciation by reinforcement of mating behaviors in chorus frogs (Pseudacris) – De novo reference transcriptome raw data, contigs and gene annotations
<p>RNA-Seq raw data used in the assembly and annotation of a reference transcriptome for the Upland Chorus Frog, <em>Pseudacris feriarum</em>. Raw data were obtained by sequencing of four tissue types: Brain, eyes, testis, and somatic. Assembled contigs (Trinity) and gene annotations (Trinotate) are also provided.</p>
Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2
<p><span>Long-read sequencing technologies have improved significantly since their emergence. Their read lengths, potentially spanning entire transcripts, is advantageous for reconstructing transcriptomes. Existing long-read transcriptome assembly methods are primarily reference-based and to date, there is little focus on reference-free transcriptome assembly. We introduce RNA-Bloom2, a reference-free assembly method for long-read transcriptome sequencing data. </span>RNA-Bloom2 is available on GitHub at: <a href="https://github.com/bcgsc/RNA-Bloom">https://github.com/bcgsc/RNA-Bloom</a>.</p> <p><span>We benchmarked the assembly quality and the computational performance of RNA-Bloom2 on simulated data. We prepared two mouse simulated datasets with Trans-NanoSim</span><span> for the cDNA and dRNA sequencing protocols model on experimental ONT data</span><span>. The datasets were simulated </span><span>based on the mouse ENSEMBL annotation for GRCm39.</span><span> To investigate the effect of sequencing depth, we subsampled each dataset to 2, 10, and 18 million reads, resulting in a total of six sets of reads for our benchmarking experiments. Using the simulated data, w</span><span>e showed that the transcriptome assembly quality of RNA-Bloom2 is competitive to those of reference-based methods.</span></p>
PacBio IsoSeq reference transcriptomes for Pinus taeda L.
<p>Fusiform rust disease, caused by the endemic fungus <i>Cronartium quercuum</i> f. sp. <i>fusiforme</i>, is the most damaging disease affecting economically important pine species in the southeast United States. In this report, we detail the genomic localization and sequence-level discovery of candidate race-nonspecific broad-spectrum fusiform rust resistance genes in <i>Pinus taeda </i>L. Two full-sib families, each with ~1000 progeny, were challenged with a complex inoculum consisting of over 150 pathogen isolates. High-density linkage mapping revealed three QTL distributed on two linkage groups. The two QTL on linkage group 2 were additive with respect to their effects on the probability of disease outcome. All three QTL were validated using a population of 2057 cloned pine genotypes in a six-year-old multi-environmental field trial. As a complement to the QTL mapping approach, bulked segregant RNAseq analysis revealed a small number of candidate nucleotide binding leucine rich repeat genes harboring SNP significantly associated with disease resistance. The results of this study demonstrate that single qualitative resistance genes can confer effective resistance against genetically diverse mixtures of an endemic pathogen.</p>
Functional annotation of the reference transcriptome of Mesodinium rubrum strain JAMR
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Dinophysis acuminata strain DAVA01
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Mesodinium rubrum strain MBL-DK2009
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Teleaulax amphioxeia
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
Functional annotation of the reference transcriptome of Dinophysis ovum strain DoSS3195
<p>Raw reads were pre-processed by removing the adaptors and low-quality reads using BBMap. The filtered reads were normalized for depth based on kmer counts using BBNorm function. De novo transcriptomes were generated using both Trinity and velvet-oases. CD-HIT-EST was used to merge the two de novo transcriptomes and reduce the transcript redundancy to 98% similarity and generate unique genes. Transcriptome assembly completeness was evaluated with BUSCO (Benchmarking Universal Single Copy Orthologs) database. Functional annotation was done using blastp function of ncbi-blast using the nr database with evalue 1E-20 and num_alignments 3.</p>
The full-length Quillaja brasiliensis (Quillajaceae) reference transcriptome
<p><strong>Introduction</strong></p> <p>In this study, we used the PacBio Iso-Seq technology to uncover the full-length transcriptome of <em>Quillaja brasiliensis </em>(Quillajaceae), a tree species native to Brazil and of great importance for the extraction of bioactive saponins. Establishing a full-length transcriptome is essential for understanding the metabolic pathways leading to saponin biosynthesis in <em>Q. brasiliensis </em>and should set the stage for upcoming investigations.</p> <p> </p> <p><strong>Sample obtention and processing</strong></p> <p>We collected seeds of <em>Quillaja brasiliensis</em> (A.St.-Hil. & Tul.) Mart. (Quillajaceae) from the city of Canguçú (Rio Grande do Sul, Brazil) in March/2018. After <em>in vitro</em> germination, we explanted the seedlings and submitted them to callogenesis or transferred them to the grow room at 25°C under a 16h/day photoperiod. We detached leaves from 2.8-year-old individuals for the experiments hereby presented. After leaf cleanup, we detached approximately 15 leaves and exposed them to ultraviolet radiation (UV-C germicide lamp, ʎ maximum 254 nm) or white light (control) (Table 1). After incubation, we froze the plant material in liquid nitrogen and stored it at -80 °C until the next processing step. We obtained calluses from <em>in vitro </em>germinated explants and submitted them to cellular suspension induction. We maintained the cell cultures in MS medium supplemented with naphthaleneacetic acid (5 mg/L) and kinetin (0.1 mg/L) in the absence of light, using 250 mL culture flasks under the agitation of 120 RPM. After establishing the growth profile, we collected three flasks for each growth phase of the liquid cell cultures: lag phase (three days), log phase (seven days), and stationary phase (21 days). For the sample collection, cells were filtered using the Büchner funnel, washed with distilled water, and flash-frozen in liquid nitrogen. We stored the collected material at -80 °C until the next processing step. We extracted RNA samples from different tissues and growth conditions using the cetyltrimethylammonium bromide method (CTAB) and performed their purification using the RNeasy MinElute Cleanup Kit (QIAGEN, Hilden, Germany). We quality-controlled the samples using fluorometric and spectrophotometric methods before submitting them to sequencing by Novogene (Beijing, People's Republic of China). We made an equimolar pool of the five sample types for sequencing using the Iso-Seq strategy (PacBio, Menlo Park, US) in circular consensus sequencing mode for the obtention of long reads. The raw reads are available at the European Nucleotide Archive (ENA) under accession PRJEB58985.</p> <table> <tbody> <tr> <td> <p><strong>Sample</strong></p> </td> <td> <p><strong>Sample source</strong></p> </td> <td> <p><strong>Treatment or Growth phase</strong></p> </td> </tr> <tr> <td> <p>Leaf UV light</p> </td> <td> <p>Leaf</p> </td> <td> <p>Ultraviolet light treatment.</p> </td> </tr> <tr> <td> <p>Leaf white light</p> </td> <td> <p> </p> </td> <td> <p>White light treatment</p> </td> </tr> <tr> <td> <p>Cell suspension lag</p> </td> <td> <p>Cell suspension</p> </td> <td> <p>Lag phase (three days)</p> </td> </tr> <tr> <td> <p>Cell suspension log</p> </td> <td> <p> </p> </td> <td> <p>Log phase (seven days)</p> </td> </tr> <tr> <td> <p>Cell suspension stationary</p> </td> <td> <p> </p> </td> <td> <p>Stationary phase (21 days)</p> </td> </tr> </tbody> </table> <p><strong>Table 1</strong>. Brief description of pooled sequenced samples. </p> <p> </p> <p><strong>Zenodo repository content</strong></p> <p>This repository stores the full-length transcriptome FASTA file obtained after running the IsoSeq v3 pipeline, followed by one round of polishing by LoRDEC and transcript-collapsing by Cogent/Cupcake-ToFU (collapsed_isoforms.fa). We also performed coding sequence prediction and functional annotation using CodAn and TRAPID 2.0 (PLAZA 4.5 Dicots), respectively (transcriptome_functional_characterization.tsv). A brief description of column names for the functional characterization table is also available here.</p> <p> </p>
Reference transcriptome assembly of a protogynous sex change fish, harlequin sandsmelt (Parapercis pulchella)
<p>Reference transcriptome sequences (superTranscripts) of a marine teleost fish <em>Parapercis pulchella</em>.</p> <p>This dataset is a part of our work, "Reference transcriptome assembly of a protogynous sex change fish, harlequin sandsmelt (<em>Parapercis</em> pulchella)" published in <em>Marine Genomics</em>.</p> <p>https://doi.org/10.1016/j.margen.2024.101086</p> <p> If you use this dataset, plese cite the above paper.</p> <p>Raw RNA-seq data and <em>de novo </em>assembled sequences generated by Trinity have been deposited in NCBI/DDBJ/EMBL under accession PRJDB16534.</p> <p>This dataset is generated from Trinity raw-assembled sequences using superTranscripts method (Corset, Lace). </p> <p>Functional annotations were conducted using eggNog-mapper, KEEG Automatic Annotation Server (KAAS), and reciprocal BLAST best-hit analysis against medaka's protein sequences.</p> <p> </p> <p>The codes for generating these data are deposited in GitHub (<a href="https://github.com/yaoakifumi/Ppul-reference-transcriptome">https://github.com/yaoakifumi/Ppul-reference-transcriptome</a>).</p> <p> </p>
Genome and Transcriptome references based on hg19 from UCSC, 2015
<p>rsem.transcripts.nant2015.fa.gz - bgzipped FASTA reference of transcriptomes</p><p>genome.nant2015.fa.gz - bgzipped FASTA human genome reference, with several viral sequences added.</p><p>refseq.txt.gz - Exact sequence accessions and mapping coordinates for a RefSeq transcriptome based off the UCSC genome browser for hg19.</p><p>Coordinates are BED-style, with one row per transcript, and 1+ transcript per gene.</p><p>Column annotation</p><p>1. RefSeq Accession</p><p>2. Chromosome</p><p>3. Strand</p><p>4. thinStart (gene boundary, including UTR)</p><p>5. thinEnd (gene boundary, including UTR)</p><p>6. thickStart (CDS boundary)</p><p>7. thinStart (CDS boundary)</p><p>8. number of exons</p><p>9. comma separated exon starts</p><p>10. comma separate exon ends</p><p>11. common gene name</p><p>12. refseq gene id</p><p>13. 0 if non-primary transcript, 1 if primary transcript</p>
Simulated data from: Reference-free assembly of long-read transcriptome sequencing data with RNA-Bloom2
Open the record for dataset details and reuse information.
PacBio IsoSeq reference transcriptomes for Pinus taeda L.
Open the record for dataset details and reuse information.
Data from: The plover neurotranscriptome assembly: transcriptomic analysis in an ecological model species without a reference genome
We assembled a de novo transcriptome of short-read Illumina RNA-Seq data generated from telencephalon and diencephalon tissue samples from the Kentish plover, Charadrius alexandrinus. This is a species of considerable interest in behavioural ecology for its highly variable mating system and parental behaviour, but it lacks genomic resources and is evolutionarily distant from the few available avian draft genome sequences. We assembled and identified over 21 000 transcript contigs with significant expression in our samples, showing high homology to exonic sequences in avian draft genomes. From these, we identified >31 000 high-quality SNPs and > 2500 simple sequence repeats (SSRs). We also analysed expression patterns in our data to identify potential candidate genes related to differences in male and female behaviour, identifying over 200 nonoverlapping putative autosomal transcripts that show significant expression differences between males and females. Gene ontology analysis revealed that female-biased transcripts were significantly enriched for cerebral functions related to learning, cognition and memory, and male-biased transcripts were mostly enriched for terms related to neural function such as neuron projection and synapses. This data set provides one of the first de novo transcriptome assemblies from non-normalized short-read next-generation data and outlines an effective strategy for measuring sequence and expression variability simultaneously without the aid of a reference genome.
Data from: Genomics of Compositae crops: reference transcriptome assemblies, and evidence of hybridization with wild relatives
Although the Compositae harbours only two major food crops, sunflower and lettuce, many other species in this family are utilized by humans and have experienced various levels of domestication. Here we have used next generation sequencing technology to develop 15 reference transcriptome assemblies for Compositae crops or their wild relatives. These data allow us to gain insight into the evolutionary and genomic consequences of plant domestication. Specifically, we performed Illumina sequencing of Cichorium endivia, Cichorium intybus, Echinacea angustifolia, Iva annua, Helianthus tuberosus, Dahlia hybrida, Leontodon taraxacoides and Glebionis segetum, as well 454 sequencing of Guizotia scabra, Stevia rebaudiana, Parthenium argentatum and Smallanthus sonchifolius. Illumina reads were assembled using Trinity, and 454 reads were assembled using MIRA and CAP3. We evaluated the coverage of the transcriptomes using BLASTX analysis of a set of ultra-conserved orthologs (UCOs) and recovered most of these genes (88-98%). We found a correlation between contig length and read length for the 454 assemblies, and greater contig lengths for the 454 compared to the Illumina assemblies. This suggests that longer reads can aid in the assembly of more complete transcripts. Finally, we compared the divergence of orthologs at synonymous sites (Ks) between Compositae crops and their wild relatives and found greater divergence when the progenitors were self-incompatible. We also found greater divergence between pairs of taxa that had some evidence of post-zygotic isolation. For several more distantly related congeners, such as chicory and endive, we identified a signature of introgression in the distribution of Ks values.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.