Skip to main content
zenodorestricted

Data from: High-resolution methylome analysis uncovers stress-responsive genomic hotspots and drought-sensitive TE superfamilies in the clonal Lombardy poplar

<p>The following dataset contains the processed data presented in the article<strong>&nbsp;&quot;High-resolution methylome analysis uncovers stress-responsive genomic hotspots and drought-sensitive TE superfamilies in the clonal Lombardy poplar&quot;</strong></p> <ul> <li><strong>Supplementary_methods.docx: </strong>contain detailed information for the experimental stress treatments, sequencing library preparation, sequencing and DMR calling.</li> </ul> <ul> <li><strong>BedGraph files&nbsp;(CpG.bed, CHG.bed, CHH.bed):</strong>&nbsp;contain methylation levels (%) for each cytosine in the Lombardy poplar genome, in the respective sequence context. The first three columns represent the genomic coordinates of the cytosine, the 56 following columns indicate&nbsp;the methylation levels for each of the samples. Missing values are represented with NA (when particular&nbsp;cytosines were not captured by the sequencing method).&nbsp;</li> <li><strong>DMR_annotation_Populus_nigra_Italica_after_biotic_and_abiotic_treatments.txt: </strong>contains all the identified regions that showed significant stress-induced differential methylation (DMR). The file include all annotation for each single DMR: genomic location, genomic feature, gene, TE, sequence context&nbsp;and stress treatment, besides other specific relevant information.</li> <li><strong>sample_IDs_basic_metadata.txt</strong>: contains the sample ID and the associated metadata (stress treatment and ortet location and ID) for all samples used in the analysis.</li> <li><strong>Supplementary_file_1_metadata_samples.xlsx:&nbsp;</strong>contains the metadata&nbsp;associated to each sample including:&nbsp;sequencing statistics before and after quality and adapter trimming, read mapping and coverage statistics, and number of interrogated cytosines on each sequence context (CpG, CHG, CHH).</li> <li><strong>Supplementary_file_2_GO_enrichments.xlsx: </strong>contains the complete results for the GO enrichment analysis for different gene datasets associated to: drought-CHH-DMRs, SINEs, MITEs, SINEs + MITEs.</li> <li><strong>italica_denovo_TE_280920.gff</strong>: contains the predicted TEs using the following methodology.&nbsp;First, TEs were de-novo annotated using the Extensive de-novo TE Annotator (EDTA) (version 1.8.3) (https://github.com/oushujun/EDTA) with default parameters, except for option --sensitive:&nbsp;1, which uses RepeatModeler (version 2.0.1) to identify remaining TEs. All the steps in EDTA pipeline were selected (filter, final and anno) in order to perform whole-genome annotation/analysis after the TE library was constructed. Then, in&nbsp;the&nbsp;annotated library from EDTA, we merged overlapping fragments and fragments&nbsp;located at a close distance (&lt;10bp) in a strand wise manner. The&nbsp;merged fragment was annotated&nbsp;as the family of longer merged fragment. Structural variants derived from nanopore data were used to redefine the boundaries of overlapping TE fragments to be more precise with actual predictions.&nbsp;LINE elements were identified independently by RepeatModeler&nbsp;in order to construct a more comprehensive de-novo TE library.</li> <li><strong>SaliS.fasta:</strong>&nbsp;contains the consensus sequences of <strong>Sali</strong>caceae <strong>S</strong>INE families (SaliS), the file was built by extracting information&nbsp;from the supplementary table 2 of the publication: &quot;Divergence of 3&prime; ends as a driver of short interspersed nuclear element (SINE) evolution in the Salicaceae&quot; (https://doi.org/10.1111/tpj.14721)</li> <li><strong>Pnigra_Italica_SaliS.bed: </strong>the file contains the annotated SaliS found by blastn over the P. nigra Italica reference genome&nbsp;(-qcov_hsp_perc 90 -perc_identity 70 -word_size 7).&nbsp;Column headers: chr, start, end, length, strand, perc_identity, SaliS family.</li> <li><strong>Pnigra_Italica_all_TEs_for_anno.bed:&nbsp;</strong>contains the merged information from&nbsp;<strong>italica_denovo_TE_280920.gff </strong>and<strong>&nbsp;Pnigra_Italica_SaliS.bed.</strong>&nbsp;Column headers: chr, start, end, length, strand, perc_identity (only for SaliS), TE superfamily.</li> <li><strong>CXX_ortet_DMRs_merged.bed</strong>: contains DMRs merged from all pairwise DMR callings between two ortets. One file per context. Column headers: chr, start, end, number of comparisons where the DMR occur, avg number of cytosines (when called in multiple DMR callings), avg differential methylation vs. control (when called in multiple DMR callings), avg adjusted p value (when called in multiple DMR callings), avg DMR length (when called in multiple DMR callings).</li> </ul> <p><strong>SCRIPTS</strong></p> <ul> <li><strong>cov_filtering.sh:</strong> to filter individual positions according to a custom threshold.</li> <li><strong>unionbedg_with_NAs.sh:</strong> to merge information from different samples in a single file taking into account the percentage of missing values per position across the given samples.</li> <li><strong>anovas_and_contrasts_boxplots_barplots_cld.r</strong>: to perform statistical tests for the effect of treatments and ortets on the average global methylation. Each sequence context was analyzed separately.</li> <li><strong>CHH_noise_filter.sh:</strong> to remove cytosines with invariable methylation values across 90% of the samples.</li> <li><strong>GlobalMethAvg_calculation.r</strong>: to calculate global average methylation given a methylation file (CpG.bed, CHG/bed or CHH.bed) and sample file.</li> <li><strong>Hclustering_and_PCAs_analysis.r:</strong> to perform hierarchical clustering, principal component analysis and plot the respective figures.</li> <li><strong>ICC_matrices_analysis.r:</strong> to calculate intraclass correlation coefficients among all pairwise combinations and plot colored grids</li> </ul> <p>&nbsp;</p> <p>Annotations are based on the de novo reference genome of the Populus nigra cv. Italica clone uploaded in the ENA project: PRJEB44889 (<a href="http://www.ebi.ac.uk/ena/browser/view/GCA_950102115">www.ebi.ac.uk/ena/browser/view/GCA_950102115</a>). Bisulfite sequencing data can be found under the ENA project: PRJEB51831</p>

ShareScore

16/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
0
Reuse readiness
0
Engagement
4

Topics