Atlas of nascent RNA transcripts reveals enhancer to gene linkages
<p>Data associated with the paper "Atlas of nascent RNA transcripts reveals enhancer to gene linkages"</p> <p>GitHub repository for the analyses: <a href="https://github.com/Dowell-Lab/DBNascent_Analysis">https://github.com/Dowell-Lab/DBNascent_Analysis</a></p> <p>Below are the summaries of the files associated with this publication. </p> <p> </p> <p>1. muMerge calls for each paper used in the merging</p> <p><strong>paper_mumerge_calls</strong></p> <p>- The calls are separated by the bidirectional caller (dreg, tfit)</p> <p>- In the folders are bed files (e.g. <em>Allen2014global_hg38_dreg_MUMERGE.bed</em>) for each paper and species (hg38, mm10)</p> <p> </p> <p>2. Base content for regions called by dREG and Tfit in each paper in mouse and human</p> <p><strong>mumerge_base_composition</strong></p> <p>- The base content for each paper after the first round of muMerge</p> <p>- The files contain the id and the base nucleotide content in 300bp around the center region (id, cg, at)</p> <p> </p> <p>3. Bidirectional regions called by Tfit and dREG after merging. Regions are for mouse and human datasets. (See <a href="https://github.com/Dowell-Lab/bidirectionals_merged">https://github.com/Dowell-Lab/bidirectionals_merged</a>)</p> <p><strong>bidirectional_regions</strong></p> <p>- Bidirectional regions called after muMerge and filtering</p> <p>- Calls for both human and mouse datasets are reported (<em>hg38_tfit_dreg_bidirectionals.bed.gz</em> and <em>mm10_tfit_dreg_bidirectionals.bed.gz</em>)</p> <p>- The bed files are in bed6 format with the following columns:</p> <p>chromosome, start, stop, bidirectional, number of papers a bidirectional was called, strand (it is . since bidirectionals are not stranded)</p> <p> </p> <p>4. Metadata for samples used in the SPECS and correlation analysis</p> <p><strong>metadata</strong></p> <p><span> - Sample metadata for filtered samples (human_samples_QC_GC_protocol_filtered.tsv.gz) in the downstream analyses </span></p> <p> </p> <p>5. SPECS scores across genes and bidirectional regions</p> <p><strong>specs_scores</strong></p> <p><span> - The SPECS scores for all tissues analyzed (filt_qc123_all_specs_all.txt.gz) are reported,</span></p> <p><span> - Along with the maximum (filt_qc123_all_specs_maxval.txt.gz) </span></p> <p><span> - And minimum SPECS scores (filt_qc123_all_specs_minval.txt.gz).</span></p> <p><span> - The SPECS scores were also split by disease vs non-disease samples</span></p> <p><span> - The TPMs summaries are also included</span></p> <p> </p> <p><span>6. Normalized counts </span></p> <p><strong><span>normalized_counts</span></strong></p> <p><span> - Gene and bidirectional region normalized counts (gene_bidir_tpm.tsv.gz)</span></p> <p> </p> <p>7. Bidirectional Region and gene pairs (See https://github.com/Dowell-Lab/bidir_gene_pairs)</p> <p><strong>bidirectional_gene_pairs</strong></p> <p>- Gene and bidirectional region pairs (<em>dbnascent_pairs.txt.gz</em>) across tissues in high-quality samples. </p> <p>- The pairs are reported in a bed12 file</p> <p> - Where the first 6 columns are gene coordinates and the following 6 are bidirectional coordinates.</p> <p> - The remaining columns are the summary statistics for correlation and the relationship between the gene and bidirectional.</p> <p><span> - Additional columns note whether the pair overlaps eQTLs from GTEx (eQTL) or polII ChIA-PET loops</span></p> <ul> <li>transcript1_chrom : Gene chromosome</li> <li>transcript1_start : Gene start coordinate</li> <li>transcript1_stop : Gene stop coordinate</li> <li>transcript_1 : Gene id</li> <li>transcript1_score : Gene score (. since none was assigned)</li> <li>transcript1_strand : Gene strand</li> <li>transcript2_chrom : Bidirectional chromosome</li> <li>transcript2_start : Bidirectional start coordinate</li> <li>transcript2_stop : Bidirectional stol coordinate</li> <li>transcript_2 : Bidirectional id</li> <li>transcript2_score : Bidirectional score (i.e. the number of papers that support a bidirectional from muMerge)</li> <li>transcript2_strand : Bidirectional strand (. since these are not stranded)</li> <li>pcc : Pearsons correlation coefficient</li> <li>pval : P-value</li> <li>adj_p_BH : Adjusted p-value (Benjamini-Hochberg correction)</li> <li>nObs : Number of observations in correlation analysis</li> <li>t : T statistic</li> <li>distance_tss : Distance between the gene start (TSS) and the bidirectional start coordinate</li> <li>distance_tes : Distance between the gene stop (TES) and the bidirectional start coordinate</li> <li>position : Is the bidirectional upstream or downstream of the TSS</li> <li>tissue : Tissue id based on metadata for tissue-derived correlations (labeled All_samples if all samples are used)</li> <li>percent_transcribed_both : Percent of the number of observed samples used in the analysis</li> <li><span>pair_id : Gene:Transcript~Bidirectional pair name</span></li> <li><span>gene_id : Gene id</span></li> <li><span>chiapet : Binary indicator for whether pair overlaps overlap polII ChIA-PET </span></li> <li><span>gtex : Bindary Indicator whether a pair is overlapping GTEx pairs</span></li> </ul> <p> </p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0