Datasets for chromatin hub prediction in six cell lines based on multiple genomic features
<p>Tables with features and classes for machine learning prediction of chromatin hubs. Genomic features include CTCF, EP300, H3K27me3, H3K36me3, H3K4me1, H3K4me2, H3K4me3, H3K9ac, H3K9me3, RAD21, RNAPol2, and RNA.Seq, while the classes are Hubs and Non-Hubs.</p> <p>The cell lines featured here are A549, H1ESC, HeLa, IMR90, K562, and MCF7. They happen to be the 6 cell lines out of 8 existing in our integrative database, GREG (https://doi.org/10.1093/database/baz162). The normalized read-coverages from features (variables) are mapped through genomic intervals of 2 Kbs, genome-wide. Such genomic intervals (bins), are classified as Hubs or Non-Hubs. Hubs are those bins with multiple chromatin interactions, including at least one long-range interaction (larger than 1Mb) or an inter-chromosomal interaction (tagged as Inf).</p> <p>Columns per table:<br> chr start end CTCF EP300 H3K27me3 H3K36me3 H3K4me1 H3K4me2 H3K4me3 H3K9ac H3K9me3 RAD21 RNA.Seq RNAPol2 Class</p> <p>Note that features may be inconsistent across different cell types, due to the availability of data. The BAM files have been sourced from ENCODE and NCBI repositories.</p> <p>The analysis following this data can be found at https://github.com/mora-lab/GREG-Hubs.</p>
ShareScore
40/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0