Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
123
datasets available to search
ShareScore release 0.9.0
Dataset results
123 results for “hydroxymethylation”
Hydroxymethylation profile of cell free DNA is a biomarker for early colorectal cancer
<p>The files in this data release represent processed data from the FORESEE study conducted by Cambridge Epigenetix Ltd, and reported in the preprint manuscript: "Hydroxymethylation profile of cell free DNA is a biomarker for early colorectal cancer" (<a href="https://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a>). </p> <p> </p> <p>As described in the manuscript, classifiers were trained and validated on genomic features extracted from sequencing datasets across cases and controls. Several classes of genomic features were constructed for training and validation data sets which are described below:</p> <p> </p> <p><strong>CRC_enhancer_znorm_training_matrix_v1.csv<br> CRC_enhancer_znorm_validation_matrix_v1.csv</strong></p> <p>Columns contain sample names, rows contain genomic features. </p> <p><em>Description of feature generation process. </em></p> <p>To calculate 5hmC levels at gene enhancers, we first calculated read counts using Bam readcounts v0.01. RPKM were calculated over candidate gene-enhancers downloaded from GeneCards v4.4. 5hmC enrichment was computed as the log2 ratio between the hydroxymethylome library RPKM and the input library RPKM after the inclusion of pseudocounts. Feature scaled (z-score normalization) 5hmC levels of enhancers quantile-normalized over samples.</p> <p><br> <strong>CRC_cegxdelfi_znorm_training_matrix_v1.tsv<br> CRC_cegxdelfi_znorm_validation_matrix_v1.tsv</strong></p> <p>Columns contain sample names, rows contain genomic features. <br> </p> <p><em>Description of feature generation process. </em><br> We divided the genome into 100KB bins and quantified cfDNA fragment sizes per bin. We removed blacklisted regions, genomic gaps (UCSC table) and non-standard chromosomes a priori. We excluded outlier bins in fragment size, only retaining fragments between 100nt to 220nt length. Finally, we split the genome into 100KB bins (in total 26170 non-overlapping genomic regions) and calculated the following characteristics of fragment size distribution per genomic bin: number of short fragments (100-150nt), number of long fragments (151-220nt), ratio between short and long fragments and the total number of fragments. This approach generates 26170 features per metric and per sample. The last step is the averaging of the 100 KB bins into larger non-overlapping genomic regions of 5 MB (in total 512 bins).</p> <p><br> <strong>CRC_cegxnps_znorm_training_matrix_v1.tsv</strong></p> <p><strong>CRC_cegxnps_validation_matrix_v1.tsv</strong></p> <p> Columns contain sample names, rows contain genomic features. <br> <em>Description of feature generation process. </em><br> Further detail in the manuscript: <a href="http://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a></p> <p><br> <strong>FORESEE_sample_description.tsv</strong></p> <p>This file holds sample data for colorectal cancer and control samples described in <a href="http://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a></p> <ul> <li>The sample_name column links to the column names in the *_matrix.tsv files</li> <li>The columns denoted raw_file1 and raw_file2 link the sample metadata with the enhancer readcount files contained in the gh_readcount_training.tar and gh_readcount_validation.tar.</li> </ul> <p>The columns in the table are briefly described below:</p> <p><em>sample_name</em>:<em> </em>Sample identifier<br> <em>Title</em>: Composed of the the disease name, gender and sample_name<br> <em>Source_name</em>: Tissue source<br> <em>Organism</em>: Contains the term: “Homo sapiens”<br> <em>Characteristics_indication</em>: Disease indication <br> <em>Characteristics_stage</em>: Cancer stage where appropriate. Indicated by roman numerals (I,II,III,IV)<br> <em>Characteristics_gender</em>: Described as “Female” or “Male”<br> <em>Characteristics_ethnicity</em>: Ethnicity description<br> <em>Characteristics_age_at_collection</em>: Age value in years<br> <em>Molecule</em>: Contains the value “cell free DNA”<br> <em>Description</em>: Contains value: “Training sample” or “Validation sample”</p> <p><em>Processed_data_file</em>: Contains the term: “CRC_enhancer_training_matrix” or “CRC_enhancer_validation_matrix”. <br> <em>raw_file1</em>: Refers to the readcount file from the 5hmC capture library<br> <em>raw_file2</em>: Refers to the readcount file from the Input control (shallow sequenced) library</p> <p> </p> <p><strong>gh_readcount_training.tar<br> gh_readcount_validation.tar</strong></p> <p>These tar files include the raw read counts computed across enhancer regions for case and control data and are referenced in the FORESEE_sample_description.tsv file.</p> <p> </p> <p><strong>Manuscript Abstract</strong></p> <p>Our classifier discriminated CRC samples from controls with an area under the receiver operating characteristic curve (AUC) of 90% (sensitivity was 55% at 95% specificity). Performance was similar for early stage 1 (AUC 89%) and late stage 4 CRC (AUC 94%). Performance was independent of the proportion of tumor-DNA in the cell free DNA. </p> <p>We expanded the classifier to include information about cell free DNA fragment size and abundance across the genome. Overall performance was similar (AUC 91%), with gains in sensitivity (63% at 95% specificity). </p> <p>The 5-hydroxymethylcytosine signal allows detection of CRC, even in cell free DNA samples with undetectable tumor DNA. Including 5-hydroxymethylcytosine in multi-analyte screening, will improve sensitivity for early-stage cancer. </p>
Early Diagnosis of Oral Cancer by Detecting p16 Hydroxymethylation
ClinicalTrials.gov study NCT02967120. IPD Sharing: YES. Countries: 1. Publications: 3.
Associations of NPPA promoter true methylation and hydroxymethylation with ischemic stroke and its functional outcome
Open the record for dataset details and reuse information.
Parkinson’s disease-associated shifts between DNA methylation and DNA hydroxymethylation in human brain
GEO Series GSE267937. Homo sapiens. 200 samples. Type: Methylation profiling by genome tiling array.
Epigenetic mapping of the somatotropic axis reveals differential DNA hydroxymethylation marks with potential implication in growth
GEO Series GSE158910. Oreochromis niloticus. 5 samples. Type: Expression profiling by high throughput sequencing.
X-Ray induced DNA-Hydroxymethylation changes
GEO Series GSE111437. Homo sapiens. 40 samples. Type: Methylation profiling by high throughput sequencing; Expression profiling by high throughput sequencing.
Divergent originations of parental DNA hydroxymethylation in human preimplantation embryos
GEO Series GSE224618. Homo sapiens. 11 samples. Type: Methylation profiling by high throughput sequencing; Expression profiling by high throughput sequencing.
Obesity and dyslipidemia are associated with partially reversible modifications to DNA hydroxymethylation in swine adipose-derived mesenchymal stem/stromal cells [Sus scrofa]
GEO Series GSE216950. Sus scrofa. 9 samples. Type: Methylation profiling by high throughput sequencing.
Sequencing DNA methylation and hydroxymethylation at co-occurring chromatin features
GEO Series GSE296587. Mus musculus. 31 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.
Genome-wide mapping of DNA hydroxymethylation in osteoarthritic chondrocytes [methylation]
GEO Series GSE64393. Homo sapiens. 8 samples. Type: Methylation profiling by high throughput sequencing.
Hydroxymethylation at gene regulatory regions directs stem cell commitment during erythropoiesis
GEO Series GSE40243. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Tet-dependent 5-hydroxymethyl-Cytosine modification of mRNA regulates the axon guidance genes robo2 and slit in Drosophila [hMeRIP]
GEO Series GSE225979. Drosophila melanogaster. 10 samples. Type: Expression profiling by high throughput sequencing.
DNA methylation and hydroxymethylation assessment through Illumina EPIC array analysis of paired bisulfite and oxidative-bisulfite conversion
GEO Series GSE144129. Homo sapiens. 454 samples. Type: Methylation profiling by genome tiling array.
Quantifying propagation of DNA methylation and hydroxymethylation with iDEMS
GEO Series GSE193681. Mus musculus. 6 samples. Type: Other.
Unified analysis of DNA methylation, hydroxymethylation and chromatin accessibility through targeted DNA labeling [TOP-Seq]
GEO Series GSE231928. Mus musculus. 9 samples. Type: Other.
Alterations of 5-hydroxymethylation in circulating cell-free DNA reflect molecular distinctions of subtypes of non-Hodgkin lymphoma
GEO Series GSE155228. Homo sapiens. 73 samples. Type: Methylation profiling by high throughput sequencing.
A paradigm for post-embryonic Oct4 re-expression: E7-induced hydroxymethylation regulates Oct4 expression in cervical cancer
GEO Series GSE248367. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
TET1-mediated hydroxymethylation facilitates hypoxic gene induction in neuroblastoma
GEO Series GSE55391. Homo sapiens. 14 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Other.
Gene body DNA hydroxymethylation restricts magnitude of transcriptional changes during aging [direct RNA-seq]
GEO Series GSE221122. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.
Roles of DNA hydroxymethylation versus DNA formylation and carboxylation in neural stem cells [RNA-seq]
GEO Series GSE287654. Mus musculus. 7 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.