Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,916
datasets available to search
ShareScore release 0.7.1
Dataset results
4,916 results for “DNA methylation”
Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study (Supplementary Data)
<p>The zip file contains supplementary data for the publication - Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study, accepted for publication in Environmental Health Perspectives (DOI: 10.1289/EHP6174).</p> <p>The description of the files are noted below:</p> <p><strong>1. Readme File for SAPALDIA Noise and Air Pollution EWAS Single Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_SingleExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_SingleExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p> </p> <p><strong>General footnote for all files:</strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter <2.5 µm. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 µg/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from single exposure epigenome-wide linear mixed models, with random intercept at the level of participant. Each model was adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator (for Lden models) and leukocyte composition. In a preliminary step, DNA methylation β-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites.</p> <p>Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (“modified winsorization”). The “winsorized” data were then used as the dependent variables in the epigenome-wide association study.</p> <p> </p> <p><strong>2. Readme File for SAPALDIA Noise and Air Pollution EWAS Multi Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_MultiExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_MultiExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p><strong>General table footnotes: </strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter <2.5 µm. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 µg/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from multi-exposure epigenome-wide linear mixed models, with random intercept at the level of participant, and were adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator and leukocyte composition. Multi-exposure models included all five exposures (Aircraft, railway, road traffic Lden and respective truncation indicators, NO<sub>2</sub> and PM<sub>2.5</sub>) at the same time. In a preliminary step, DNA methylation β-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites. Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (“modified winsorization”). The “winsorized” data were then used as the dependent variables in the epigenome-wide association study.</p>
DNA methylation dynamics during stress-response in woodland strawberry (Fragaria vesca)
<p><strong>Genome sequence and annotation of Fragaria vesca cv. Reine des Vallées</strong></p> <p>In order to generate a reference genome for Fragaria vesca cv. Reine des Vallées, we used MinIon long-read sequencing data to substitute the <em>F. vesca</em> genome v.4.0.a2 genome. The detailed method used to obtain these results were the following:</p> <p><em>Genome sequencing and assembly NIL Fb2</em></p> <p>Genomic DNA from strawberry plants was extracted by a Hexadecyltrimethylammonium bromide (Cetrimonium bromide, CTAB) modified protocol (Healey, Furtado, Cooper, & Henry, 2014) and purified with Agencourt AMPure XP beads (cat# A63880). Long-read sequencing was performed for the genome assembly; Genomic DNA by Ligation (Oxford Nanopore, cat# SQK-LSK109) library was prepared as described by the manufacturer and sequenced on a MinION for 72 h (Oxford Nanopore).</p> <p><em>Reference genome polishing</em></p> <p>Reads obtained from nanopore were filtered with Filtlong v0.2.1 (<a href="https://github.com/rrwick/Filtlong">https://github.com/rrwick/Filtlong</a>) using --min_mean_q 80 and --min_length 200. Cleaned reads were then aligned to the most recent version of the <em>F. vesca</em> genome v4.0.a2, downloaded from the Genome Database for Rosaceae (GDR) (<a href="https://www.rosaceae.org/species/fragaria_vesca/genome_v4.0.a2">https://www.rosaceae.org/species/fragaria_vesca/genome_v4.0.a2</a>), using minimap2 v2.21 (H. Li, 2018) with parameters -aLx map-ont --MD -Y. The generated BAM file was then sorted and indexed with samtools v1.11 (H. Li et al., 2009). We used mosdepth v0.3.1 (Pedersen & Quinlan, 2018) to verify that coverage on chromosomic scaffolds was over 50 X. Sniffles v1.0.12a (Sedlazeck et al., 2018) with parameters -s 10 -r 1000 -q 20 --genotype -l 30 -d 1000 was used to detect structural variations larger than 30 bp. The VCF files obtained from Sniffles was sorted and filtered with BCFtools v1.14 (Danecek et al., 2021) to keep only structural variants (SV) with smaller than 200,00 bp (we observed that larger SV were most of the time false positive caused by misalignments in regions with gaps or Ns), supported by 10 or more reads and with allelic frequencies above 0.8 (we were interested in homozygous changes). The complete filtering command used is “bcftools view -q 0.8 -Oz -i '(SVTYPE = "DUP" || SVTYPE = "INS" || SVTYPE = "DEL" || SVTYPE = "TRA" || SVTYPE = "INV" || SVTYPE = "INVDUP") && %FILTER = "PASS" && FMT/DV>9 && SVLEN>29 && SVLEN<200000' “</p> <p>From the VCF listing all the structural variants that we detected in our <em>F. vesca </em>accession, we generated a substituted genome version based on the reference <em>F. vesca</em> genome v.4.0.a2. The reference genome was first indexed with samtools faidx v1.11(Danecek et al., 2021) and a sequence dictionary was generated with Picard CreateSequenceDictionary v2.25.6 (<a href="https://broadinstitute.github.io/picard">https://broadinstitute.github.io/picard</a>). The VCF containing the SV produced from our Nanopore sequencing was also indexed with gatk (Van der Auwera GA & O'Connor BD, 2020) IndexFeatureFile v4.2.0.0 (<a href="https://gatk.broadinstitute.org/hc/en-us/articles/360037262651-IndexFeatureFile">https://gatk.broadinstitute.org/hc/en-us/articles/360037262651-IndexFeatureFile</a>). FastaAlternateReferenceMaker v4.2.0.0 (<a href="https://gatk.broadinstitute.org/hc/en-us/articles/360037594571-FastaAlternateReferenceMaker">https://gatk.broadinstitute.org/hc/en-us/articles/360037594571-FastaAlternateReferenceMaker</a>) was then run with the reference genome and the VCF file to generate a substituted genome representative of our <em>Fragaria</em> accession.</p> <p>As substituting our genome with the detected structural variants changes genomic coordinates, we also corrected the public GFF genome annotation of <em>F. vesca</em> (Y, Pi, Gao, Liu, & Kang, 2019) using liftoff v1.6.1 (Shumate & Salzberg, 2021). Liftoff also detects and annotates duplications within the substituted genome.</p> <p>Transposable elements annotation was carried out using the EDTA transposable element annotation pipeline v. 1.9.6 (S. Ou et al., 2019) on the substituted genome using default parameters<em>.</em></p> <p><strong>Differentially methylated regions</strong></p> <p>The file Stress_vs_control_DMRs.zip file contains the DMRs that were called using the reads submitted to ENA (ERP135585) and obtained as follows:</p> <p>First, bedGraph files from wgbs pipeline were pre-filtered for a minimum coverage of 5 reads using awk command. These output files were then used as input for the EpiDiverse/dmr bioinformatics analysis pipeline for non-model plant species to define DMRs (Nunn <em>et al</em>., 2021) with default parameters (minimum coverage threshold 5; maximum q-value 0.05; minimum differential methylation level 10%; 10 as minimum number of Cs; Minimum distance (bp) between Cs that are not to be considered as part of the same DMR is 146 bp). The pipeline uses metilene v.0.2.6.1 (<a href="https://www.bioinf.uni-leipzig.de/Software/metilene/">https://www.bioinf.uni-leipzig.de/Software/metilene/</a>) for pairwise comparison between groups and R-packages ggplot2 v.3.3.5 and gplots v.3.1.1, for visualization results (Fig. S1). Based on our <em>F. vesca</em> genome transcript annotation and methylation data (overlapped regions with DNA methylation cytosines and DMRs), we detected the methylated genes, promoters, 3’ UTRs, 5’UTR and transposable elements in strawberry. Global DNA methylation and DMR plots were performed with R-package ggplot2. Gene analyses by methylation patterns and analysis of per-family TE DNA methylation profiles were performed with deepTools v.3.5.0 (Ramírez <em>et al</em>., 2014). DMRs comparison between treatments were done by the Venn diagram v.1.7.0 R-package.</p> <p>We produced several genome browsers tracks with DMRs that we integrated in our local instance of JBrowse available at the following url: <a href="https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub">https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub</a></p>
Single-molecule DNA methylation patterns of full-length human-specific LINE-1 (L1HS) retrotransposons in a panel of cell lines.
<p>We used bs-ATLAS-seq to comprehensively map the genomic location and assess the DNA methylation status of full-length human-specific LINE-1 elements (L1HS). The approach capture region 1-210 of L1HS elements, which corresponds to the most 5' end of its promoter sequence. This was performed in a panel of 12 human primary or transformed cell lines (BJ, IMR90, MRC5, H1, K562, HCT116, HeLa S3, HepG2, MCF7, HEK-293, HEK-293T, 2102Ep), many being shared with the encode project.</p> <p>These datasets provide a visualization for DNA methylation patterns at the single molecule level for each L1HS loci.</p>
Crossreactive probes on Illumina DNA methylation arrays: a large study on ALS shows that a cautionary approach is warranted in interpreting epigenome-wide association studies
<p>Data corresponding to the paper "Crossreactive probes on Illumina DNA methylation arrays: a large study on ALS shows that a cautionary approach is warranted in interpreting epigenome-wide association studies."<br> <br> Corresponding scripts can be found at: <a href="https://github.com/pjhop/dnamarray_crossreactivity">https://github.com/pjhop/dnamarray_crossreactivity</a><br> All downstream analyses in <a href="https://github.com/pjhop/dnamarray_crossreactivity/blob/master/analysis/c9_analysis.Rmd">c9_analysis.Rmd</a> and in<a href="https://github.com/pjhop/dnamarray_crossreactivity/blob/master/analysis/supplementary_note.Rmd"> supplementary_note.Rmd</a> can be reproduced using the deposited data as follows:</p> <ul> <li>Clone the dnamarray_crossreactivity repository: < git clone https://github.com/pjhop/dnamarray_crossreactivity.git ></li> <li>Download the data ('data.zip') and place it in the 'dnamarray_crossreactivity' folder.</li> <li>Unzip the data.zip folder</li> </ul> <p>Scripts used to generate the data in each subdirectory can be found at:</p> <ul> <li>data/processed/c9_matches/: <a href="https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/c9_matches">https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/c9_matches</a></li> <li>data/output/ewas/: <a href="https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/ewas">https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/ewas</a></li> <li>data/output/figs/: empty folder, running 'c9_analysis.Rmd' will save figures here.</li> <li>data/misc/: <a href="https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/other">https://github.com/pjhop/dnamarray_crossreactivity/tree/master/analysis/other</a></li> <li>data/extdata: <ul> <li>Zhou <em>et al.</em> annotations (EPIC.hg19.manifest.tsv.gz, HM450.hg19.manifest.pop.tsv.gz, HM450.hg19.manifest.tsv.gz) were downloaded from: <a href="https://zwdzwd.github.io/InfiniumAnnotation">https://zwdzwd.github.io/InfiniumAnnotation</a> (downloaded at 17/09/2020)</li> <li>Naeem <em>et al.</em><em> </em>data (12864_2013_7006_MOESM2_ESM.csv) was downloaded from: <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3943510/">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3943510/</a></li> <li>Chen <em>et al.</em> data (48639-non-specific-probes-Illumina450k.xlsx) was downloaded from <a href="https://github.com/Jfortin1/funnorm_repro/blob/master/bad_probes/48639-non-specific-probes-Illumina450k.xlsx">https://github.com/Jfortin1/funnorm_repro/blob/master/bad_probes/48639-non-specific-probes-Illumina450k.xlsx</a></li> <li>The anno_450k.txt.gz and anno_EPIC.txt.gz are subsets of the annotation files included in the following package respectively: <a href="https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylation450kanno.ilmn12.hg19.html">https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylation450kanno.ilmn12.hg19.html</a> and <a href="https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylationEPICanno.ilm10b2.hg19.html">https://bioconductor.org/packages/release/data/annotation/html/IlluminaHumanMethylationEPICanno.ilm10b2.hg19.html</a></li> </ul> </li> <li> data/genome_bs: Scripts used to generate these data can be found at <a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg19.R</a> and <a href="https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R">https://github.com/pjhop/DNAmCrosshyb/blob/master/data-raw/bisulfite_convert_hg38.R</a> .</li> <li> data/raw: Individual-level data is available upon access at: <a href="https://ega-archive.org/studies/EGAS00001004587">https://ega-archive.org/studies/EGAS00001004587</a></li> </ul>
Supporting data for "The methylome of Biomphalaria glabrata and other mollusks: enduring modification of epigenetic landscape and phenotypic traits by a new DNA methylation inhibitor"
<p>Methylome of the fresh water snail <em>Biomphalaria glabrata</em>. DNA was extracted from the feet of 10 individuals of <em>B. glabrata</em> originally isolated from Brazil. These snails have been cultivated in the laboratory since 1960. Tissue were grinded at 4°C and incubated in 1 ml volume of lysis buffer (20 mM TRIS pH 8; 1 mM EDTA; 100 mM NaCl; 0.5% SDS), with 0.3 mg of proteinase K at 55°C for 1 night. Afterwards, lysate was purified with phenol-chloroform and DNA was isopropanol precipitated. The extracted DNA (around 138ng/µL) was poled in equivalent amounts and Whole Genome Bisulfite Sequencing was done by GATC-biotech (www.gatc-biotech.com). The principle of this treatment is to convert non-methylated cytosines of gDNA into deoxy-uracil, whereas methylated cytosines remain intact. WGBS was done according to the Lister protocol (sequence 2 forward strands only). The reference genome (Biomphalaria-glabrata-BB02_SCAFFOLDS_BglaB1.fa) and annotation (Biomphalaria-glabrata-BB02_BASEFEATURES_BglaB1.3.gff3) used in this project are available on VectorBase (https://www.vectorbase.org/). To align our short reads, we chose to use two specific bisulfite mapping tools, BSMAP 1.0.0 (https://code.google.com/p/bsmap/) and Bismark 0.10.2 (www.bioinformatics.babraham.ac.uk /projects/bismark/), to compare their efficiency and convenience to finally work with the more suitable one on our datasets. IGV (Interactive Genomics Viewer, https://www.broadinstitute.org/igv/) was used to visualized final alignments.<br> BSMAP performed better than Bismark and was used for downstream analyses. Without default parameters alignement efficiency for BSMAP is 47.1%, allowing for 2 mismatches increases it to 55.6%. Methylation occurs predominantly in CpGs. (C methylated in CpG context: 12.4%, C methylated in CHG context: 0.5%, C methylated in CHH context: 0.5%) The major part of CpG sites, 95.7% were unmethylated, of the remaining 4.3% of CpG sites around 3.8% had low methylation, and 0.5% were completely methylated. Methylation is of the mosaic type. Methylation is relatively low with 1.2% of total cytosines. Our analyses suggested that conserved genes and genes with stable expression are localized in high methylated regions of the genome. Finally, we see that repetitive sequences were predominantly situated in low methylated regions of <em>B. glabrata</em>. </p> <p>Wiggle files were generated for CpG pairs only.</p> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>
Suplementary data, results and scripts: "Reconstruction of Cell-specific Models Capturing the Influence of Metabolism on DNA methylation in Cancer"
<p>This repository contains supplementary data, models and scripts associated with "Reconstruction of Cell-specific Models Capturing the Influence of Metabolism on DNA methylation in Cancer".</p><p>Folders content:</p><p>'data_results_matlabscripts': data, result files and scripts (original python scripts and adapted MATLAB scripts)</p><p>'supplementary_figures': supplementary figures</p><p>'supplementary_tables': supplementary tables</p>
Differential DNA methylation in the benign and cancerous prostate tissue of African American and European American men
<p>The data presented here are the summary statistics for the manuscript, "Differential DNA methylation in the benign and cancerous prostate tissue of African American and European American men." The study aims to improve our understanding of prostate cancer disparities between African American and European American men by comparing the DNA methylation features that distinguish tumor and paired, histologically benign tissue from a sample of African American and European American prostate cancer patients. The summary statistics presented here represent the results of a differential methylation analyses comparing tumor and benign tissue in each ancestry group as well as the results of an analysis of differential methylation by ancestry group within each tissue. </p> <p>The files included are: </p> <p>AA_TumorvBenign_Dummy.zip which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in African Americans based on a model that accounts for the paired nature of samples using a series of dummy variables for individual. </p> <p>AA_TumorvBenign_MixedModel.zip which which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in African Americans based on a model that accounts for the paired nature of samples using a linear mixed model that included patient as a random effect. </p> <p>EA_TumorvBenign_Dummy.zip which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in European Americans based on a model that accounts for the paired nature of samples using a series of dummy variables for individual. </p> <p>EA_TumorvBenign_MixedModel.zip which which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in European Americans based on a model that accounts for the paired nature of samples using a linear mixed model that included patient as a random effect. </p> <p>Benign_AncestryCompare_Summary.zip which contains the results of an association analysis between ancestry designation (African American vs European American with African American as the baseline) and individual CpG sites in benign tissue. </p> <p>Tumor_AncestryCompare_Summary.zip which contains the results of an association analysis between ancestry designation (African American vs European American with African American as the baseline) and individual CpG sites in tumore tissue. </p>
COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS
<p>We provide significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS. We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within ±1Mb, which we assumed they act through cis mechanisms. We include the model’s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value < 0.05 and R<sup>2</sup> > 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>
Variation in DNA methylation and response to short-term herbivory in Thlaspi arvense
<p>Plant metabolic pathways and gene networks involved in the response to herbivory are well-established, but the impact of epigenetic factors as modulators of those responses is less understood. Here, we studied the role of DNA cytosine methylation on phenotypic responses after short-term herbivory in <em>Thlaspi arvense</em> plants with two contrasting flowering phenotypes. We investigated the effect of experimental demethylation and herbivory treatments following a 2x3 factorial design. First, half the seeds were submerged in a water solution of the demethylating agent 5-azacytidine and the other half only in water, as controls. Then, we assigned control and demethylated plants to three herbivory categories (i) insect herbivory, (ii) artificial herbivory, and (iii) undamaged plants. The effects of the demethylation and herbivory treatments were assessed by quantifying genome-wide global DNA cytosine methylation, concentration of leaf glucosinolates, final stem biomass, fruit and seed production, and seed size. For most of the plant traits analysed, individuals from the two plant-types responded differently. In late-flowering plants, global DNA methylation did not differ between control and demethylated plants but it was significantly reduced by herbivory. Conversely, in early-flowering plants, demethylation at seed stage was still evident in leaf genomes of reproductive individuals whereas herbivory did not affect their global DNA methylation.</p>
Data and code for the publication "DNA methylation underpins the epigenomic landscape regulating genome transcription in Arabidopsis"
<p>The zipped file of this repository contains code and data to reproduce the results of the publication:</p> <p>Zhao et al, DNA methylation underpins the epigenomic landscape regulating genome transcription in Arabidopsis. Genome Biology (2022). </p> <p>All sequence data have been deposited in NCBI GEO accession codes GSE183987 and GSE169497.</p> <p> </p> <p>Please see the README document for detailed:</p> <p>- Descriptions of the code and data provided</p> <p>- Lists of the required dependencies</p>
Meta-analysis results of epigenome-wide association studies in neonates reveals widespread differential DNA methylation associated with birthweight
<p>Birthweight is associated with health outcomes across the life course, DNA methylation may be an underlying mechanism. In this meta-analysis of epigenome-wide association studies of 8,825 neonates from 24 birth cohorts in the Pregnancy And Childhood Epigenetics Consortium, DNA methylation in neonatal blood is associated with birthweight at 914 sites, with a difference in birthweight ranging from -183 to 178 grams per 10% increase in methylation (P<sub>Bonferroni</sub><1.06x10<sup>-7</sup>).</p>
Data and Code for "Cell Type-specific Genome Scans of DNA Methylation Diversity Indicate an Important Role for Transposable Elements"
<p>This is a release of the gitlab repository "meta-methylome" (https://gitlab.com/okartal/meta-methylome.git) that, in addition to the code, also contains the resulting genomic data.</p> <p>Extract the directory on the command line using</p> <pre><code class="language-bash">$ tar -xhzvf meta-methylome.tar.gz</code></pre> <p>to preserve the symbolic links.</p>
Opioid medication use and blood DNA methylation: epigenome-wide association meta-analysis
<p>We conducted the first large-scale epigenome-wide meta-analysis of blood DNA methylation and recent use of opioid medications. There were five participating studies (10,842 individuals; 9,886 European ancestry and 956 African ancestry participants) including four that used the newer Illumina EPIC/850K array and one that used the older Illumina 450K array. We identified novel loci differentially methylated in relation to opioid medication use.</p>
Structural dynamics of DNA depending on methylation pattern: Simulation dataset
<p>Dataset of molecular dynamics simulations of double-stranded DNA<br> (repeats of CpG dinucleotides with regular patterns of methylation).</p> <p>- Input: Parameters and initial structures<br> - Output: Trajectories</p> <p>NAMD 2.13 (multi-core with CUDA) was used for the simulations.</p>
Data from: Non-invasive age estimation based on fecal DNA using methylation-sensitive high-resolution melting for Indo-Pacific bottlenose dolphins
<p class="MsoNormal"><span>Age is necessary information for the study of life history of wild animals. A general method to estimate the age of odontocetes is counting dental growth layer groups (GLGs). However, this method is highly invasive as it requires the capture and handling of individuals to collect their teeth.</span><span> Recently, the development of DNA-based age </span><span>estimation methods has been actively studied as an alternative to such invasive methods, of which many have used biopsy samples. However, if DNA-based age estimation can be developed from fecal samples, age estimation can be performed without touching or disrupting individuals, thus establishing an entirely non-invasive method. </span><span>We developed an age estimation model using the methylation rate of two gene regions, <em>GRIA2</em> and <em>CDKN2A,</em> measured through methylation-sensitive high-resolution melting (MS-HRM) from fecal samples of wild Indo-Pacific bottlenose dolphins (<em>Tursiops aduncus</em>). The age of individuals was known through conducting longitudinal individual identification surveys underwater. Methylation rates were quantified from 36 samples. Both gene regions showed a significant correlation between age and methylation rate. The age estimation model was constructed based on the methylation rates of both genes which achieved sufficient accuracy (after LOOCV: MAE = 5.08, <em>R<sup>2</sup></em> = 0.34) for the ecological studies of the Indo-Pacific bottlenose dolphins, with a lifespan of 40-50 years. This is the first study to report the use of non-invasive fecal samples to estimate the age of marine mammals.</span></p>
Age estimation of captive Asian elephants (Elephas maximus) based on DNA methylation: An exploratory analysis using methylation-sensitive high-resolution melting (MS-HRM)
<p>Age is an important parameter for bettering the understanding of biodemographic trends-development, survival, reproduction and environmental effects-critical for conservation. However, current age estimation methods are challenging to apply to many species, and no standardised technique has been adopted yet. This study examined the potential use of methylation-sensitive high-resolution melting (MS-HRM), a labour, time, and cost-effective method to estimate chronological age from DNA methylation in Asian elephants (<em>Elephas maximus</em>). The objective of this study was to investigate the accuracy and validation of MS-HRM use for age determination in long-lived species, such as Asian elephants. The average lifespan of Asian elephants is between 50-70 years but some have been known to survive for more than 80 years. DNA was extracted from 53 blood samples of captive Asian elephants across 11 zoos in Japan, with known ages ranging from a few months to 65 years. Methylation rates of two candidate age-related epigenetic genes, <em>RALYL</em> and <em>TET2,</em> were significantly correlated with chronological age. Finally, we established a linear, unisex age estimation model with a mean absolute error (MAE) of 7.36 years. This exploratory study suggests an avenue to further explore MS-HRM as an alternative method to estimate the chronological age of Asian elephants.</p>
Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models
<p><span>Knowledge of individual age can help both in-situ and ex-situ conservation programs to design more efficient and suitable management plans for targeted wildlife species. DNA methylation is one of the epigenetic aging markers that has emerged as a promising tool that can estimate age with high accuracy using only a tiny amount of biological material, which can be collected in a minimally invasive way. Here, we sequenced five targeted genetic regions and used </span><span>8–23</span><span> selected CpG sites to build age estimation models with machine learning methods </span><span>with about only $3–7 per sample</span><span>, using blood samples of seven Felidae species—ranging from small to big, and domestic to endangered species: domestic cats (<em>Felis catus</em>, 139 samples), Tsushima leopard cats (<em>Prionailurus bengalensis euptilurus</em>, 84 samples), and five<em> Panthera </em>species (96 samples). </span><span>The models built achieved satisfactory accuracy—the mean absolute error of the best models was 1.966, 1.348, and 1.552 years in domestic cats, Tsushima leopard cats, and <em>Panthera</em> spp., respectively.</span><span> Our models in domestic cats and Tsushima leopard cats were applicable to individuals regardless of health conditions, indicating the high applicability of our models to samples collected from diverse situations, e.g., rescued individuals in the context of conservation. We also showed the possibility of developing universal age estimation models for the five<em> Panthera</em> spp. using two of the five genetic regions, suggesting an even lower cost to use our models for future applications.</span></p>
Early developmental carry-over effects on exploratory behaviour and DNA methylation in wild great tits (Parus major)
<p>Adverse, postnatal conditions experienced during development are known to induce lingering effects on morphology, behaviour, reproduction and survival. Despite the importance of early developmental stress for shaping the adult phenotype, it is largely unknown which molecular mechanisms allow for the induction and maintenance of such phenotypic effects once the early environmental conditions are released. Here we aimed to investigate whether lasting early developmental phenotypic changes are associated with post-developmental DNA methylation changes. We used a cross-foster and brood size experiment in great tit (Parus major) nestlings, which induced post-fledging effects on biometric measures and exploratory behaviour, a validated personality trait. We investigated whether these post-fledging effects are associated with DNA methylation levels of CpG sites in erythrocyte DNA. Individuals raised in enlarged broods caught up on their developmental delay after reaching independence and became more explorative as days since fledging passed, while the exploratory scores of individuals that were raised in reduced broods remained stable. Although we previously found that brood enlargement hardly affected pre-fledging methylation levels, we found 420 CpG sites that were differentially methylated between fledged individuals that were raised in small versus large sized broods. A considerable number of the affected CpG sites were located in or near genes involved in metabolism, growth, behaviour and cognition. Since the biological functions of these genes line up with the observed post-fledging phenotypic effects of brood size, our results suggest that DNA methylation provides organisms the opportunity to modulate their condition once the environmental conditions allow it. In conclusion, this study shows that nutritional stress during early development associates with indirect, carry-over effects on DNA methylation. We propose that treatment-associated DNA methylation differences arise as a consequence of pre-fledging phenotypic changes, rather than that they cause early environmentally-induced effects.</p>
Genetic architecture of immune cell DNA methylation in the rhesus macaque
<p><strong>Complete model outputs from rhesus macaque (<em>Macaca mulatta</em>) whole blood meQTL and eQTL analyses in article, "Genetic architecture of immune cell DNA methylation in the rhesus macaque". </strong></p> <p><strong><em>cis</em> meQTL model output (SNP-CpG associations):</strong> </p> <ol> <li>IMAGE_573_meqtl_model_res_wPVE.txt: <ul> <li>Model results from IMAGE meQTL mapping including all genome, chromatin state annotations, and PVE estimates</li> </ul> </li> <li>pqlseq_allimagesnps_res_converged_wpve.txt: <ul> <li>Model results from PQLseq meQTL mapping including PVE estimates </li> </ul> </li> </ol> <p><strong><em>cis</em> eQTL model output (SNP-gene associations): </strong></p> <ol> <li>eqtl_res_sva5_gemma_172samples_qvalue.txt: <ul> <li>Model results from GEMMA eQTL mapping </li> </ul> </li> </ol> <p> </p>
Multi-cell type deconvolution using a probabilistic model for single-molecule DNA methylation haplotypes
<p>Files required to run deconvolution with CelFIE-ISH and Epistate, in U250 regions from Loyfer et al. 2023, in both "pat" and "epiread" formats. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.