Skip to main content
zenodoopen

Quantification of the effects of chimerism: datasets

<p>To aid in exploring the effects of chimerism on read mapping, differential expression analysis and&nbsp;<em>de novo</em>assembly, a base set of 26,680 transcripts containing all sequences ranging in length of between 300 and 5000 nt present within the fruit fly cDNA library was created from Ensembl&nbsp;release-100&nbsp;(https://www.ensembl.org/info/data/ftp/index.html)&nbsp;[1]. These transcripts along with the complete cDNA library from which they were compiled are located within the&nbsp;<strong><em>BaseSetTranscripts</em>.zip</strong> file.</p> <p>This base set of transcripts was used as a reference for simulating reads as required within subsequent sections of our paper (titled:&nbsp;<em>Quantification of the effects of chimerism on read mapping, differential expression and annotation following short-read&nbsp;de novo&nbsp;assembly.</em>). Such simulations often involved hundreds of replicate iterations due to the nature of the study, as well as for the associated creation of modified base sets containing varying portions of chimerism. Parameter values used for read simulations and modified reference sets used within iterations are described in detail within the paper (F1000&nbsp;<em>paper link to be provided when available.).</em></p> <p>To explore the effects of chimerism on the detection of differentially expressed transcripts ten read datasets, each consisting of five million read-pairs, were simulated using CSReadGen&nbsp;[2]&nbsp;from the base set as described section 2.2 of the manuscript. These are located within the&nbsp;<strong><em>DEReads.zip</em></strong>&nbsp;file. Using these reads differential expression analysis was repeated iteratively, where during each iteration ChimSim&nbsp;&nbsp;[3]&nbsp;was used to create a modified base set to be used as a reference. Within each modified base set created a portion of the transcripts present were made chimeric. The portions of chimerism introduced ranged from 5% to 95% chimeric in steps of five. These modified base sets are located within the d&nbsp;<em><strong>DEChimSimRefs.zip</strong>&nbsp;</em>file. In each case a titles file has also been provided that indicates which transcripts within the base set were made chimeric (if any, e.g. at 0% chimerism this file is empty) and the manner in which chimeras was introduced in accordance to the three types discussed in the paper. For example in the file titled chimeric_refs_0.1_SEQS.fasta 10% of the sequences are chimeric and the file titled chimeric_refs_0.1_TITLES.txt indicates which these are and the type of chimerism introduced.</p> <p>The base set was then used to simulate&nbsp;ten data sets consisting of ten million read-pairs that were each assembled using CStone&nbsp;[4], Trinity&nbsp;[5]&nbsp;and rnaSPAdes&nbsp;[6]. Parameters for read simulations are once again described in detail within our paper (Section 2.3). The assemblies produced by each assembler are contained within the <strong><em>DeNovoAssemblies_SimulatedData.zip</em></strong> file.&nbsp;The two whole body read datasets from Pang et al.&nbsp;[7], following filtering by&nbsp;Trimmomatic&nbsp;[8]&nbsp;as described in our paper, are within the files <em><strong>Reads_RealData_WholeBody_1.zip</strong></em> and <em><strong>Reads_RealData_WholeBody_2.zip</strong></em>, as are the assemblies produced by each of the three assemblers when using these reads as input (<em><strong>DeNovoAssemblies_RealData.zip</strong></em>).</p> <p>Related software to this project are:<br> 1.&nbsp;<a href="http://sourceforge.net/projects/cstone/">CStone</a>&nbsp;<br> 2.&nbsp;<a href="http://sourceforge.net/projects/csreadgen/">CSReadGen</a><br> 3.&nbsp;<a href="https://sourceforge.net/projects/cview/">CView</a><br> 4.&nbsp;<a href="https://sourceforge.net/projects/chimsim/">ChimSim</a>&nbsp;&lt;<br> 5.&nbsp;<a href="https://sourceforge.net/projects/tvscript/">TVScript</a></p> <p>A related poster discussing the the identification of&nbsp;chimerism during assembly is available&nbsp;<a href="https://zenodo.org/record/6022494#.YgY-ji2cbGI">here</a>&nbsp;(DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.6022493">10.5281/zenodo.6022493</a>) and one discussing the effects of chimerism&nbsp;is available&nbsp;<a href="https://zenodo.org/record/6023171#.YgZmCC2cZQL">here</a>&nbsp;(DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.6023170">10.5281/zenodo.6023170</a>).</p> <p>General details of the project are available&nbsp;<a href="https://cibio.up.pt/en/projects/de-novo-based-sequence-assembly-of-next-generation-sequence-data-without-chimeras-improved-annotation-gene-expression-profiles-and-haplotype-br-reconstruction/">here</a>.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Yates AD, Achuthan P, Akanni W, Allen J, Allen J, Alvarez-Jarreta J, et al. Ensembl 2020. Nucleic Acids Res. 2020;48: D682&ndash;D688. doi:10.1093/NAR/GKZ966</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Archer J. CSReadGen website. 2020. Available: https://sourceforge.net/projects/csreadgen/</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Linheiro, Raquel; Archer J. ChimSim website. 2021. Available: https://sourceforge.net/projects/chimsim/</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Linheiro R, Archer J. CStone: A de novo transcriptome assembler for short-read data that identifies non-chimeric contigs based on underlying graph structure. Pertea M, editor. PLOS Comput Biol. 2021;17: e1009631. doi:10.1371/JOURNAL.PCBI.1009631</p> <p>5.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Grabherr MG, Haas BJ, Yassour M, Levin JZ, Thompson DA, Amit I, et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 2011 297. 2011;29: 644&ndash;652. doi:10.1038/nbt.1883</p> <p>6.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Bushmanova E, Antipov D, Lapidus A, Prjibelski AD. rnaSPAdes: a de novo transcriptome assembler and its application to RNA-Seq data. Gigascience. 2019;8: 1&ndash;13. doi:10.1093/GIGASCIENCE/GIZ100</p> <p>7.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Pang TL, Ding Z, Liang SB, Li L, Zhang B, Zhang Y, et al. Comprehensive Identification and Alternative Splicing of Microexons in Drosophila. Front Genet. 2021;12. doi:10.3389/fgene.2021.642602</p> <p>8.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Bolger AM, Lohse M, Usadel B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics. 2014;30: 2114&ndash;2120. doi:10.1093/BIOINFORMATICS/BTU170</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics