Intergenic RNAPII Atlas : output data
<p>This dataset represents the RNAPII (RNAP2) Atlas of potentially transcribed intergenic regions of the human genome by integrating 906 high quality human Chromatin-ImmunoPrecipitation sequencing (ChIP-seq) biosamples targeting the RNA Polymerase II, obtained from public data warehouses. </p> <p><strong>Github Code available here : </strong><a href="https://github.com/benoitballester/Pol2Atlas">https://github.com/benoitballester/Pol2Atlas</a> </p> <p><strong>The dataset consists of 5 zipped folders described below: </strong></p> <p><strong>./pol2_consensuses/:</strong><br> consensuses.bed:<br> Location of intergenic RNAP2 consensuses in bed format for the hg38 assembly. <br> First three columns are genomic locations, 4th column is consensus ID, <br> 5th column is the number of datasets with RNAP2 observed at this RNAP2 consensus,<br> 6th column is strand (not used), 7-8th columns is consensus centroid.<br> consensusesHg19.bed:<br> Location of intergenic RNAP2 consensuses in hg19 assembly. ~1000 are missing due to liftover.<br> Consensus ID is matching with the hg38 one.<br> matrix.mtx:<br> RNAP2 occupancy consensus-dataset binary matrix in sparse matrix market format.<br> Corresponding row annotation are RNAP2 consensuses.<br> Corresponding column annotation are datasets stored in dataset.txt.<br> datasets.txt:<br> See matrix.mtx<br> clusterConsensuses_Labels.txt:<br> Assigned cluster for each RNAP2 consensus.<br> intersectIntergPol2.tsv:<br> RNAP2 consensuses with cluster ID and intersections with reference databases.<br> cluster_bed/:<br> consensuses.bed splitted per cluster.<br> saf_files/:<br> Files typically used for read counting with featureCounts. Suffixes:<br> _500 : RNAP2 consensuses standardized to 1kbp.<br> _all : All RNAP2 consensuses including genic.<br> Hg19 : Intergenic RNAP2 Lifted to Hg19.</p> <p> </p> <p><strong>./rnap2_all_peaks/:</strong><br> all_peaks.bed.gz:<br> Concatenated bed file with all POLR2A peaks from all experiments, genome wide, for the hg38 assembly. <br> Peaks are filtered with a MACS2 qvalue > 1e-5, datasets with less than 100 peaks in intergenic regions are removed.<br> First three columns are genomic locations, 4th column contains sample of origin of the peak, 5th column is the<br> MACS2 q-value, 6th column is dna strand (not used), 7-8th are peak "summit". 9th column contains an r,g,b value<br> corresponding to the biotype of origin (Blood / Immune, Brain, Embryo...) for easy visualization in a genome browser.<br> Legend is available in legend.png. Conversion table between rgb values and biotype in palette.csv.<br> Note that singletons are removed when creating consensus peaks.<br> all_peaks_interg.bed.gz:<br> Same as above, but for intergenic regions only (excluding 1kb before TSS and 1kb after TES).</p> <p> </p> <p><strong>./count_tables_rnaseq/:</strong><br> ENCODE/:<br> counts.mtx.gz:<br> Count table in sparse matrix market format. Row corresponds to samples, columns to Pol II probes (Pol2_500.saf).<br> samples.csv.gz:<br> Matching row annotation for count matrix.<br> encode_total_rnaseq_annot.tsv.gz:<br> Sample annotation (not ordered!).<br> GTEx/:<br> counts.mtx.gz:<br> Count table in sparse matrix market format. Row corresponds to samples, columns to Pol II probes (Pol2_500.saf).<br> samples.csv.gz:<br> Matching row annotation for count matrix.<br> sample_annot.tsv.gz:<br> Sample annotation (not ordered!).<br> TCGA/:<br> counts.mtx.gz:<br> Count table in sparse matrix market format. Row corresponds to samples, columns to Pol II probes (Pol2_500.saf).<br> samples.csv.gz:<br> Matching row annotation for count matrix.<br> annotation_table.tsv.gz:<br> Sample annotation (not ordered!).<br> </p> <p><br> <strong>./cancer_markers/:</strong><br> bed/:<br> DE_Tumor_vs_Normal/:<br> TCGA-*/:<br> allWithStats.bed:<br> FDR, mean difference in pearson residuals and log2(FC) for each RNAP2 probe. Warning: probes are prefiltered to have > 1 read in 3 samples, make sure to use row index to match <br> with RNAP2 consensuses.<br> allDE.bed:<br> All DE (cancer vs normal) probes in bed format for this cancer. <br> 5th column has been replaced by enrichment p-value.<br> DE_downreg.bed:<br> Downregulated (cancer vs normal) probes in bed format for this cancer. <br> 5th column has been replaced by enrichment p-value.<br> DE_upreg.bed:<br> Upregulated (cancer vs normal) probes in bed format for this cancer. <br> 5th column has been replaced by enrichment p-value.<br> classifier_TCGA-*:<br> Performance of a machine learning tumor-normal tissues classifier using Pol II probes as input.<br> globally_DE.bed:<br> Probes DE in 7+ cancers (FPR permutation threshold). Last column indicates the number of cancers this probe is DE in.<br> globally_Down regulated.bed:<br> Probes DE in 6+ cancers (FPR permutation threshold). Last column indicates the number of cancers this probe is Down regulated in.<br> globally_Up regulated.bed:<br> Probes DE in 5+ cancers (FPR permutation threshold). Last column indicates the number of cancers this probe is Up regulated in.<br> subtypes/:<br> BRCA/:<br> allWithStats_BRCA.*.bed:<br> FDR, mean difference in pearson residuals and log2(FC) for each RNAP2 probe for DE test of sample from this subtype against normal samples. <br> Warning: probes are prefiltered to have > 1 read in 3 samples, make sure to use row index to match <br> with RNAP2 consensuses.<br> bed_BRCA.*.bed:<br> All DE (subtype vs normal) probes in bed format for this cancer. <br> bed_uniqueDE_BRCA.*.bed:<br> All DE (subtype vs normal) probes in bed format for this cancer and not DE in any other subtype. </p> <p> TCGA_survival/:<br> TCGA-*/:<br> prognostic.bed:<br> All probes associated with survival for this cancer. <br> 5th column has been replaced by p-value.<br> stats.csv:<br> Cox linear model statistics for each Pol II probe. <br> Warning: probes are prefiltered to have > 1 read in 3 samples, make sure to use row index to match <br> with RNAP2 consensuses.<br> globally_prognostic.bed:<br> Probes associated with survival in 5+ cancers (FPR permutation threshold). 5th column has been replaced with <br> the number of cancers this probe is associated with survival in.<br> tabular/:<br> Same as above but stored in a tabular binary format for DE and survival.</p> <p> </p> <p><strong>./metacluster_markers/:</strong><br> bed/:<br> allPol2_datasetCount:<br> For each tissue, all Pol II consensuses, with 5th column indicating the number of<br> datasets (RNAP2, GTEx, ENCODE, TCGA tumour and normal) in which the RNAP2 consensus is<br> considered a marker.<br> robust_2_datasets_per_tissue:<br> For each tissue, Pol II consensuses considered marker in 2+ datasets out of 5 <br> (RNAP2, GTEx, ENCODE, TCGA tumour and normal).<br> tabular/:<br> Each Pol II consensus with marker information stored in a binary format.<br> </p>
ShareScore
28/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 0