Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

70

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

70 results for “transcript assembly”

Learn how ShareScore rates datasets ↗
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

<p>The European starling, <em>Sturnus vulgaris</em>, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (<em>S. vulgaris</em> vAU), and a second North American genome (<em>S. vulgaris</em> vNA), as complementary reference genomes for population genetic and evolutionary characterisation. <em>S. vulgaris</em> vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (<em>Taeniopygia guttata</em>) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using <em>S. vulgaris</em> vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BUSCOMP, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, <em>S. vulgaris</em> vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.</p>

opencc-zeroJul 2022View details →
zenodo40/100

Fig. 5 in Successful transcription but not translation or assembly of Solenopsis invicta virus 3 in a baculovirus-driven expression system

Fig. 5. Confirmation of heterologous expression of SINV-3 transcript by amplification of the 3' (A) and 5' termini (B) from RNA templates purified from SINV- 3-transfected Sf21 cells. Regions amplified are illustrated in the genome diagram between (A) and (B). (A) Three plaque preparations (AcSINV-3 CiC, AcSINV-3 CiD, and AcSINV-3 DiD) were separated by centrifugation into soluble and pelleted fractions, treated with DNase I, reverse transcribed, and the 3' end of the genome amplified by PCR. Lane assignments were as follows: 1 = mass marker (bp); 2, 6 = AcSINV-3 CiC; 3, 7 = AcSINV-3 CiD; 4, 8 = AcSINV-3 DiD; 5, 9 = mock infection; 10 = positive control (wild-type virus); 11 = negative control; 12 = non-template control. (B) AcSINV-3 plaque preparations (CiC and DiD) evaluated by PCR of the 5' end of the genome (RNA preparations). Lane assignments were as follows: 1 = mass marker (bp); 2, 3 = DNase treated, reverse transcribed; 4, 5 = without DNase treatment, reverse transcribed; 6, 7 = DNase treated, without reverse transcription; 8, 9 = without DNase treatment, without reverse transcription; 10 = positive control; 11 = negative control; 12 = non-template control. (C) PCR amplification of the entire SINV-3 genome from DNA preparations of AcSINV-3 CiC (lane 2) and AcSINV-3 DiD (lane 3). Lane 1, molecular markers; lane 4, non-template control. (D) Western blot to evaluate translation of the SINV-3 transcript by detection of viral capsid protein 2 (VP2). Lane assignments were as follows: 1 = AcSINV-3 CiC (4 dpi); 2 = AcSINV-3 CiC (5 dpi); 3 = AcSINV-3 CiD (4 dpi); 4 = AcSINV-3 DiD (5 dpi); 5 = mock infection (negative control); 6 = positive control (purified wild-type SINV-3; kDa).

opencc-by-4.0Sep 2015View details →
zenodo40/100

Fig. 3 in Successful transcription but not translation or assembly of Solenopsis invicta virus 3 in a baculovirus-driven expression system

Fig. 3. (A) Plaque assay results for recombinant SINV-3 (AcSINV-3) transfection of Sf21 cells indicating the dilution used for each plate. (B upper panel) Representa- tive plaque with Sf21 cells stained with neutral red 10 d afer transfection (magnified 100 times). (B lower panel) Corresponding mock-infected Sf21 cells (negative control) afer 10 d of exposure. Plaque areas identify infection of insect cells by virus with corresponding cell death.

opencc-by-4.0Sep 2015View details →
zenodo40/100

Fig. 2 in Successful transcription but not translation or assembly of Solenopsis invicta virus 3 in a baculovirus-driven expression system

Fig. 2. (A) pFastBac1_SINV-3 hybrid construct map. Locations of the restriction sites (black hash marks), bacterial transposon Tn7 sites (grey triangles), polyhedrin promoter (angled arrow corresponding to the sequence below), SINV-3 open reading frames (dark closed arrows), and approximate location of the area detected by the polyclonal antibody preparation (pAb) are shown. (B) Verified sequence of the pFastBac1_SINV-3 hybrid construct illustrating the late gene polyhedrin promoter (angled arrow). SINV-3 sequence is in bold font,with pFastBac1 sequence in normal font, and restriction sites are superscripted and corresponding sequences italicized.Underlined sequence represents the inserted late gene polyhedrin core promoter.Analyzed sequences of the 5'and 3'termini and restriction sites were identical to the wild-type virus.

opencc-by-4.0Sep 2015View details →
zenodo40/100

Fig. 1 in Successful transcription but not translation or assembly of Solenopsis invicta virus 3 in a baculovirus-driven expression system

Fig. 1. Schematic of the SINV-3 genome and sub-cloning strategy to assemble the pFastBac1 donor plasmid/SINV-3 construct. (A) Organization of the SINV-3 genome illustrating the 2 ORFs numbered 1 and 2 that encode for non-structural and structural proteins, respectively, and the genome sections sub-cloned. Oligonucleotide primers used to generate cDNA and amplify each section are indicated. Primers with introduced restriction sites are also indicated. (B) Unique restriction sites for each sub-clone and the assembly process employed to concatenate the entire SINV-3 genome in the pFastBac1 donor vector.

opencc-by-4.0Sep 2015View details →
zenodo40/100

Updated spiny mouse transcriptome assembly (now includes embryo-specific transcripts)

<p><strong>Summary</strong></p> <p>Updated spiny mouse transcriptome. Embryo-specific contigs generated from BioProject&nbsp;PRJNA436818&nbsp;were added to the Trinity_v2.3.2&nbsp;spiny mouse&nbsp;<em>de novo&nbsp;</em>transcriptome assembly (https://doi.org/10.5281/zenodo.808870).</p> <p>&nbsp;</p> <p><strong>Methods</strong></p> <p>Embryos were collected from female spiny mice (n=12) in accordance with the Australian Code of Practice for the Care and Use of Animals for Scientific Purposes with approval from the Monash Medical Centre Animal Ethics Committee. Female dams were staged from delivery of their previous litter (spiny mice conceive their next litter approximately 12h postpartum) and culled at specific time-points for embryo retrieval at the required stage: 2-cell at 48h postpartum (n=4), 4-cell at 52h postpartum (&#39;early&#39; 4-cell; n=2) or at 68h postpartum (&#39;late 4-cell&#39;; n=2), and 8-cell at 72h postpartum (n=4). Embryos were snap frozen in cell lysis solution per&nbsp;the Nugen SoLo protocol (version M01406v3; available from NuGEN).&nbsp;After ligation of cDNA, qPCR was performed on all samples to determine the number of amplification cycles required to ensure that amplification was in the linear range. Based on these results, each sample was amplified using 24 cycles. Final libraries were quantitated by Qubit and size profile determined by the Agilent Bioanalyzer. All libraries were in the expected size range (~320-360 bp).&nbsp;Custom &#39;AnyDeplete&#39; rRNA depletion probes were designed and produced by NuGEN Technologies, Inc (San Carlos, CA, USA) using rRNA sequences from our reference transcriptome (Mamrot et al., 2017; https://doi.org/10.5281/zenodo.808870). Prior to use, efficacy and off-target effects of the rRNA depletion probes were examined <em>in silico</em> by NuGEN. Samples were loaded using c-Bot (200pM per library pool) and run on 2 lanes of an Illumina HiSeq 3000 8-lane flow-cell. PhiX spike-in was not used directly due to incompatibility with the custom rRNA depletion probes, however it was incorporated into other lanes of the same HiSeq 3000 run. RNA-Seq data (100bp, paired-end reads) are available from the NCBI as Bioproject PRJNA436818.</p> <p>The quality of RNA-Seq reads was assessed using FastQC v0.11.6 (<a href="https://github.com/s-andrews/FastQC">https://github.com/s-andrews/FastQC</a>; 50f0c26), with MultiQC v1.4 (<a href="https://github.com/ewels/MultiQC">https://github.com/ewels/MultiQC</a>; baefc2e) reports available from Github (<a href="https://github.com/jpmam1">https://github.com/jpmam1</a>) (Ewels et al., 2016). Adapter sequences were trimmed from the reads using trim-galore v0.4.2 (<a href="https://github.com/FelixKrueger/TrimGalore">https://github.com/FelixKrueger/TrimGalore</a>; d6b586e), implementing cutadapt v1.12 (<a href="https://github.com/marcelm/cutadapt">https://github.com/marcelm/cutadapt</a>; 98f0e2f). Reads with a quality scores lower than 20 and read pairs in which either forward or reverse reads were trimmed to fewer than 35 nucleotides were discarded. Further trimming of poor quality reads was conducted using Trimmomatic v0.36 (<a href="http://www.usadellab.org/cms/index.php?page=trimmomatic">http://www.usadellab.org/cms/index.php?page=trimmomatic</a>) with settings &quot;LEADING:3 TRAILING:3 SLIDINGWINDOW:4:20 AVGQUAL:25 MINLEN:35&quot; (Bolger et al., 2014). Nucleotides with quality scores lower than 3 were trimmed from the 3&rsquo; and 5&rsquo; read ends. Reads with an average quality score lower than 25 or with a length of fewer than 35 nucleotides after trimming were removed. Error correction of reads was performed using Rcorrector v1.0.2 (<a href="https://github.com/mourisl/Rcorrector">https://github.com/mourisl/Rcorrector</a>; 144602f) (Song et al., 2015). FastQC was used to assess the improvement in read quality after trimming adapter removal; MultiQC reports are available from Github (<a href="https://github.com/jpmam1">https://github.com/jpmam1</a>).</p> <p>Error corrected reads were assembled using Trinity v2.4.0 (<a href="https://github.com/trinityrnaseq/trinityrnaseq">https://github.com/trinityrnaseq/trinityrnaseq</a>; 1603d80) with settings &quot;--max_memory 400G, --CPU 32 and --full_cleanup&quot; (Haas et al., 2013). Assembly statistics were computed using the TrinityStats.pl from the Trinity package, and summary statistics are provided in Table S1. All reads were aligned to this transcriptome assembly using Bowtie2 v2.2.5 (<a href="https://github.com/BenLangmead/bowtie2">https://github.com/BenLangmead/bowtie2</a>; e718c6f) with settings: &quot;--end-to-end, --score-min L,-0.1,-0.1, --no-mixed, --no-discordant, -k 100, -X 1000, --time, -p 24&quot; (Langmead &amp; Salzberg, 2012).</p> <p>Read-supported contigs were identified within the embryo-specific Trinity <em>de novo </em>transcriptome assembly using samtools &quot;idxstats&quot; v1.5 (contigs with &gt;=1 reads aligning were retained) (<a href="https://github.com/samtools/samtools">https://github.com/samtools/samtools</a>; f510fb1) (Li et al., 2009). The read-supported contigs from the embryo-specific assembly (n=54,660) were added to the reference spiny mouse transcriptome assembly previously described (Mamrot, J., Legaie, R., Ellery, S.J., Wilson, T., Seemann, T., Powell, D.R., Gardner, D.K., Walker, D.W., Temple-Smith, P., Papenfuss, A.T. and Dickinson, H., 2017. De novo transcriptome assembly for the spiny mouse (Acomys cahirinus). Scientific Reports, 7(1), p.8996).</p> <p>The updated transcriptome is comprised of&nbsp;2,274,638 transcripts in total.</p>

opencc-by-4.0Mar 2018View details →
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

Open the record for dataset details and reuse information.

publicJul 2022View details →
zenodo36/100

Paulinella micropora KR01 assembled transcripts

<p>Description of files:</p> <p><strong>5RACE_Combined_D_L_assembly.fa</strong></p> <p>Transcripts assembled from 5-prime enriched RNA sequencing reads (cap-switching transcripts). Read data available from NCBI&rsquo;s SRA repository (BioProject ID&nbsp;PRJNA730897).</p> <p><strong>Standard_RNAseq_transcripts.fa</strong></p> <p>Transcripts assembled from standard RNA sequencing reads (standard transcripts) that align to the <em>P. micropora&nbsp;</em>KR01 genome. Read data available from NCBI&rsquo;s SRA repository (BioProject ID&nbsp;PRJNA568118).</p> <p><strong>5RACE_Combined_D_L_assembly_mapped.seqnames.txt</strong></p> <p>Names of the cap-switching transcripts that align to the <em>P. micropora </em>KR01 genome.</p> <p><strong>5RACE_Combined_D_L_assembly_unmapped.seqnames.txt</strong></p> <p>Names of the cap-switching transcripts that no not align to the <em>P. micropora </em>KR01 genome.</p> <p><strong>5RACE_SL_Transcripts.seqnames.txt</strong></p> <p>Names of the cap-switching transcripts that align to the <em>P. micropora </em>KR01 genome and encode a <em>Paulinella</em> spliced-leader sequence.</p> <p><strong>5RACE_SL_Transcripts_unique2RACE.seqnames.txt</strong></p> <p>Names of the cap-switching transcripts that align to the <em>P. micropora </em>KR01 genome, encode a <em>Paulinella</em> spliced-leader sequence, and are unique to the cap-switching dataset (i.e., are not encoded in the standard transcripts).</p> <p><strong>Standard_RNAseq_SL_transcripts.seqnames.txt</strong></p> <p>Names of the standard transcripts that align to the <em>P. micropora </em>KR01 genome and encode a <em>Paulinella</em> spliced-leader sequence.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Assembled transcripts: Another lesson from unmapped reads – in depth analysis of RNA-Seq reads from various horse tissues

<p>Assembled transcripts &ndash; putative transcripts de novo assembled with Trinity software</p> <p>&nbsp;</p> <p>Data: loin adipose - AD, hoof lamina - LM, liver - LI, longissimus muscle - LO, left lung - LU, heart left ventricle - LV, ovary- OV and parietal cortex &ndash; PC<br> 1 &ndash; horse ECA_UCD_AH1<br> 2 &ndash; horse ECA_UCD_AH2</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Desmodemsus cf. armatus assembled transcripts

<p>Transcripts from <em>Desmodesmus </em>cf. <em>armatus</em> (isolated from a waste stabilisation pond and initially identified as <em>Scenedesmus</em> sp.) assembled from Illumina RNA-Seq (150 bp paired-end reads) using Trinity.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Chlorella vulgaris (UTEX 259) assembled transcripts

<p>Transcripts from <em>Chlorella vulgaris</em> (UTEX 259) assembled from Illumina RNA-Seq (150 bp paired-end reads) using Trinity.</p>

opencc-by-4.0May 2023View details →
dryad32/100

Data from: De novo and reference transcriptome assembly of transcripts expressed during flowering provide insight into seed setting in tetraploid red clover

Open the record for dataset details and reuse information.

publicNov 2017View details →
dryad28/100

Data from: Performance of gene expression analyses using de novo assembled transcripts in polyploid species

Motivation: Quality of gene expression analyses using de novo assembled transcripts in species that experienced recent polyploidization remains unexplored. Results: Differential gene expression (DGE) analyses using putative genes inferred by Trinity, Corset and Grouper performed slightly differently across five plant species that experienced various poly-ploidy histories. In species that lack recent polyploidy events that occurred in the past several millions of years, DGE analyses using de novo assembled transcriptomes identified 54–82% of the differen-tially expressed genes recovered by mapping reads to the reference genes. However, in species that experienced more recent polyploidy events, the percentage decreased to 21–65%. Gene co-expression network analyses using de novo assemblies vs. mapping to the reference genes recov-ered the same module that significantly correlated with treatment in one species that lacks recent polyploidization.

opencc-zeroSep 2019View details →
zenodo28/100

Fig. 4 in Successful transcription but not translation or assembly of Solenopsis invicta virus 3 in a baculovirus-driven expression system

Fig. 4. Quantitative PCR (absolute) results evaluating transcript production of SINV-3 by AcSINV-3-infected Sf21 cells. RNA preparations from AcSINV-3 (2 rep- licates, Ci and Di) were treated with DNase I, reverse transcribed, and amplified by qPCR. Results were compared with a series of plasmid constructs containing a portion of the SINV-3 genome (102–109 genome equivalents). The region amplified was at the 3'-most end of ORF2 (see Fig. 2). No amplification was detected in polyhedrin-negative AcRP23.lacZ preparations.

opencc-by-4.0Sep 2015View details →
dryad28/100

Data from: Performance of gene expression analyses using de novo assembled transcripts in polyploid species

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad28/100

Data from: Generation of transcript assemblies and identification of single nucleotide polymorphisms from seven lowland and upland cultivars of switchgrass

Open the record for dataset details and reuse information.

publicMar 2015View details →
geo24/100

Merging short and stranded long reads improves transcript assembly

GEO Series GSE215355. Homo sapiens. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2023View details →
geo24/100

SUPT3H-less SAGA coactivator can assemble and function without significantly perturbing RNA polymerase II transcription in mammalian cells

GEO Series GSE175901. Homo sapiens; Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2022View details →
geo24/100

Chaperone-mediated ordered assembly of the SAGA and NuA4 transcription co-activator complexes

GEO Series GSE128448. Schizosaccharomyces pombe. 36 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2019View details →
geo24/100

Deep sequencing and de novo assembly of the mouse oocyte transcriptome define the contribution of transcription to the DNA methylation landscape.

GEO Series GSE70116. Mus musculus. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record