Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

123

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

123 results for “hydroxymethylation”

Learn how ShareScore rates datasets ↗
zenodo32/100

Hydroxymethylation profile of cell free DNA is a biomarker for early colorectal cancer

<p>The files in this data release represent&nbsp;processed data from the FORESEE study conducted by Cambridge Epigenetix Ltd, and reported in the preprint manuscript:&nbsp;&nbsp;&quot;Hydroxymethylation profile of cell free DNA is a biomarker for early colorectal cancer&quot; (<a href="https://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a>).&nbsp;</p> <p>&nbsp;</p> <p>As described in the manuscript, classifiers were trained and validated on genomic features extracted from sequencing datasets across&nbsp;cases and controls.&nbsp; Several classes of genomic features were constructed for training and validation data sets which are described below:</p> <p>&nbsp;</p> <p><strong>CRC_enhancer_znorm_training_matrix_v1.csv<br> CRC_enhancer_znorm_validation_matrix_v1.csv</strong></p> <p>Columns contain sample names, rows contain genomic features.&nbsp;</p> <p><em>Description of feature generation process.&nbsp;</em></p> <p>To calculate 5hmC levels at gene enhancers, we first calculated read counts using Bam readcounts v0.01. RPKM were calculated over candidate gene-enhancers downloaded from GeneCards v4.4.&nbsp;5hmC enrichment was computed as the log2 ratio between the hydroxymethylome library RPKM and the input library RPKM after the inclusion of pseudocounts. Feature scaled (z-score normalization) 5hmC levels&nbsp;of enhancers quantile-normalized over samples.</p> <p><br> <strong>CRC_cegxdelfi_znorm_training_matrix_v1.tsv<br> CRC_cegxdelfi_znorm_validation_matrix_v1.tsv</strong></p> <p>Columns contain sample names, rows contain genomic features.&nbsp;<br> &nbsp;</p> <p><em>Description of feature generation process.&nbsp;</em><br> We divided the genome into 100KB bins and quantified cfDNA fragment sizes per bin. We removed blacklisted regions, genomic gaps (UCSC table) and non-standard chromosomes a priori.&nbsp;We excluded outlier bins in fragment size, only retaining fragments between 100nt to 220nt length. Finally, we split the genome into 100KB bins (in total 26170 non-overlapping genomic regions)&nbsp;&nbsp;and calculated the following characteristics of fragment size distribution per genomic bin: number of short fragments (100-150nt), number of long fragments (151-220nt), ratio between short and&nbsp;&nbsp;long fragments and the total number of fragments. This approach generates 26170 features per metric and per sample. The last step is the averaging of the 100 KB bins into larger non-overlapping&nbsp;&nbsp;genomic regions of 5 MB (in total 512 bins).</p> <p><br> <strong>CRC_cegxnps_znorm_training_matrix_v1.tsv</strong></p> <p><strong>CRC_cegxnps_validation_matrix_v1.tsv</strong></p> <p>&nbsp; Columns contain sample names, rows contain genomic features. &nbsp;<br> <em>Description of feature generation process.&nbsp;</em><br> &nbsp; &nbsp; Further detail in the manuscript:&nbsp;<a href="http://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a></p> <p><br> <strong>FORESEE_sample_description.tsv</strong></p> <p>This file holds sample data for colorectal cancer and control samples described in <a href="http://www.researchsquare.com/article/rs-667874/v1">Walker et al. 2021</a></p> <ul> <li>The sample_name column&nbsp;links to the column names in the *_matrix.tsv files</li> <li>The columns denoted raw_file1 and raw_file2 link the sample metadata with the enhancer&nbsp;readcount files contained in the gh_readcount_training.tar and gh_readcount_validation.tar.</li> </ul> <p>The columns in the table are briefly described below:</p> <p><em>sample_name</em>:<em> </em>Sample identifier<br> <em>Title</em>: Composed of the the disease name, gender and sample_name<br> <em>Source_name</em>: Tissue source<br> <em>Organism</em>: Contains the term: &ldquo;Homo sapiens&rdquo;<br> <em>Characteristics_indication</em>: Disease indication&nbsp;<br> <em>Characteristics_stage</em>: Cancer stage where appropriate. Indicated by roman numerals (I,II,III,IV)<br> <em>Characteristics_gender</em>: Described as &ldquo;Female&rdquo; or &ldquo;Male&rdquo;<br> <em>Characteristics_ethnicity</em>: Ethnicity description<br> <em>Characteristics_age_at_collection</em>: Age value in years<br> <em>Molecule</em>: Contains the value &ldquo;cell free DNA&rdquo;<br> <em>Description</em>: Contains value: &ldquo;Training sample&rdquo; or &ldquo;Validation sample&rdquo;</p> <p><em>Processed_data_file</em>: Contains the term: &ldquo;CRC_enhancer_training_matrix&rdquo; or &ldquo;CRC_enhancer_validation_matrix&rdquo;. &nbsp;<br> <em>raw_file1</em>: Refers to the readcount file from the 5hmC capture library<br> <em>raw_file2</em>: Refers to the readcount file from the Input control (shallow sequenced) library</p> <p>&nbsp;</p> <p><strong>gh_readcount_training.tar<br> gh_readcount_validation.tar</strong></p> <p>These tar files include the raw read counts computed across enhancer regions for case and control data and are referenced in the FORESEE_sample_description.tsv file.</p> <p>&nbsp;</p> <p><strong>Manuscript Abstract</strong></p> <p>Our classifier discriminated CRC samples from controls with an area under the receiver operating characteristic curve (AUC) of 90% (sensitivity was 55% at 95% specificity). Performance was similar&nbsp;for&nbsp;early stage 1 (AUC 89%) and late stage 4 CRC (AUC 94%). Performance was independent of the proportion of tumor-DNA in the cell free DNA.&nbsp;&nbsp;</p> <p>We expanded the classifier to include information about cell free DNA fragment size and abundance across the genome. Overall performance was similar (AUC 91%), with gains in sensitivity (63% at 95% specificity).&nbsp;</p> <p>The 5-hydroxymethylcytosine signal&nbsp;allows detection of CRC, even&nbsp;in&nbsp;cell free DNA&nbsp;samples with undetectable tumor DNA.&nbsp;Including&nbsp;5-hydroxymethylcytosine in multi-analyte&nbsp;screening, will improve&nbsp;sensitivity&nbsp;for early-stage cancer.&nbsp;</p>

opencc-by-4.0Aug 2021View details →
ClinicalTrials.gov32/100

Early Diagnosis of Oral Cancer by Detecting p16 Hydroxymethylation

ClinicalTrials.gov study NCT02967120. IPD Sharing: YES. Countries: 1. Publications: 3.

controlledIPD-YESFeb 2026View details →
dryad32/100

Associations of NPPA promoter true methylation and hydroxymethylation with ischemic stroke and its functional outcome

Open the record for dataset details and reuse information.

publicNov 2024View details →
geo24/100

Parkinson’s disease-associated shifts between DNA methylation and DNA hydroxymethylation in human brain

GEO Series GSE267937. Homo sapiens. 200 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenDec 2025View details →
geo24/100

Epigenetic mapping of the somatotropic axis reveals differential DNA hydroxymethylation marks with potential implication in growth

GEO Series GSE158910. Oreochromis niloticus. 5 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2021View details →
geo24/100

X-Ray induced DNA-Hydroxymethylation changes

GEO Series GSE111437. Homo sapiens. 40 samples. Type: Methylation profiling by high throughput sequencing; Expression profiling by high throughput sequencing.

openGEO-OpenJan 2019View details →
geo24/100

Divergent originations of parental DNA hydroxymethylation in human preimplantation embryos

GEO Series GSE224618. Homo sapiens. 11 samples. Type: Methylation profiling by high throughput sequencing; Expression profiling by high throughput sequencing.

openGEO-OpenJun 2024View details →
geo24/100

Obesity and dyslipidemia are associated with partially reversible modifications to DNA hydroxymethylation in swine adipose-derived mesenchymal stem/stromal cells [Sus scrofa]

GEO Series GSE216950. Sus scrofa. 9 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenMay 2023View details →
geo24/100

Sequencing DNA methylation and hydroxymethylation at co-occurring chromatin features

GEO Series GSE296587. Mus musculus. 31 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.

openGEO-OpenJan 2026View details →
geo24/100

Genome-wide mapping of DNA hydroxymethylation in osteoarthritic chondrocytes [methylation]

GEO Series GSE64393. Homo sapiens. 8 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenJun 2015View details →
geo24/100

Hydroxymethylation at gene regulatory regions directs stem cell commitment during erythropoiesis

GEO Series GSE40243. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMar 2014View details →
geo24/100

Tet-dependent 5-hydroxymethyl-Cytosine modification of mRNA regulates the axon guidance genes robo2 and slit in Drosophila [hMeRIP]

GEO Series GSE225979. Drosophila melanogaster. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2024View details →
geo24/100

DNA methylation and hydroxymethylation assessment through Illumina EPIC array analysis of paired bisulfite and oxidative-bisulfite conversion

GEO Series GSE144129. Homo sapiens. 454 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenJan 2021View details →
geo24/100

Quantifying propagation of DNA methylation and hydroxymethylation with iDEMS

GEO Series GSE193681. Mus musculus. 6 samples. Type: Other.

openGEO-OpenOct 2022View details →
geo24/100

Unified analysis of DNA methylation, hydroxymethylation and chromatin accessibility through targeted DNA labeling [TOP-Seq]

GEO Series GSE231928. Mus musculus. 9 samples. Type: Other.

openGEO-OpenDec 2023View details →
geo24/100

Alterations of 5-hydroxymethylation in circulating cell-free DNA reflect molecular distinctions of subtypes of non-Hodgkin lymphoma

GEO Series GSE155228. Homo sapiens. 73 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenFeb 2021View details →
geo24/100

A paradigm for post-embryonic Oct4 re-expression: E7-induced hydroxymethylation regulates Oct4 expression in cervical cancer

GEO Series GSE248367. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2023View details →
geo24/100

TET1-mediated hydroxymethylation facilitates hypoxic gene induction in neuroblastoma

GEO Series GSE55391. Homo sapiens. 14 samples. Type: Expression profiling by high throughput sequencing; Methylation profiling by high throughput sequencing; Other.

openGEO-OpenMay 2014View details →
geo24/100

Gene body DNA hydroxymethylation restricts magnitude of transcriptional changes during aging [direct RNA-seq]

GEO Series GSE221122. Mus musculus. 8 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2023View details →
geo24/100

Roles of DNA hydroxymethylation versus DNA formylation and carboxylation in neural stem cells [RNA-seq]

GEO Series GSE287654. Mus musculus. 7 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record