Skip to main content
zenodoopen

sequenced cfDNA samples, extracted DoC for TSSs (HK + PAU) and TFBSs (LYL1 and GRHL2)

<p>File contains CSV files which store depth of coverage values (cfDNA fragment&#39;s central 60 bp if length &gt; 60 bp; entire fragment otherwise) extracted from 4 samples (sample BAM files deposited at the European Genome Phenome Archive, accession: <span> </span><a href="https://ega-archive.org/studies/EGAS00001006963">EGAS00001006963</a><span><span>) before and after correction of sample specific GC sequence content bias (GC bias) by GCparagon software using a custom Pysam implementation. The Pysam implementation can read alignment tags for DoC computation which is required for the signal after bias correction.</span></span></p> <p><span><span>Extracted loci are top 10k reported transcription factor binding sites for LYL1 and GRHL2 as well as transcription start sites of housekeeping genes (HK genes) and unexpressed genes as assessed from the Protein atlas (PAU genes).<br> Pysam implementation by Kapidzic Faruk can be found here: </span></span><a href="https://github.com/Faruk-K/pysam">https://github.com/Faruk-K/pysam</a></p> <p>Linked to the <a href="https://github.com/BGSpiegl/GCparagon/tree/including_EGAS00001006963_results-DEV">GCparagon tool code repository</a> which is available on GitHub.</p> <p>For details see publication which is linked to the EGA dataset and the GitHub repository. (not published at moment of submission to Zenodo)</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4