Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

69,051

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

69,051 results for “Cancer”

Learn how ShareScore rates datasets ↗
zenodo44/100

Patient-derived and artificial ascites have minor effects on MeT-5A mesothelial cells and do not facilitate ovarian cancer cell adhesion

<p>Raw data of &quot;Patient-derived and artificial ascites have minor effects on MeT-5A mesothelial cells and do not facilitate ovarian cancer cell adhesion&quot;.</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Data for: Machine learning identifies robust matrisome markers and regulatory mechanisms in cancer

<p>The expression and regulation of matrisome genes - the ensemble of extracellular matrix, ECM, ECM-associated proteins and regulators as well as cytokines, chemokines and growth factors - is of paramount importance for the many biological processes and signals within the tumor microenvironment. The availability of large and diverse multi-omics data enables mapping and understanding the regulatory circuitry governing the tumor matrisome to an unprecedented level, though such a volume of information requires robust approaches to data analysis and integration. In this study, we show that combining Pan-Cancer expression data from The Cancer Genome Atlas (TCGA) with genomics, epigenomics and microenvironmental features from TCGA and other sources enables the identification of &ldquo;landmark&rdquo; matrisome genes and machine learning-based reconstruction of their regulatory networks in 74 clinical and molecular subtypes of human cancers and approx. 6700 patients. These results, enriched for prognostic genes and cross-validated markers at the protein level, unravel the role of genetic and epigenetic programs in governing the tumor matrisome and allow the prioritization of tumor-specific matrisome genes (and their regulators) for the development of novel therapeutic approaches.</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Supplementary files for Molecular Differences Between Squamous Cell Carcinoma and Adenocarcinoma Cervical Cancer Subtypes: Potential Prognostic Biomarkers

<p>Supplementary files for Molecular Differences Between Squamous Cell Carcinoma and Adenocarcinoma Cervical Cancer Subtypes: Potential Prognostic Biomarkers</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Fluorescent Confocal Laser Scanning Microscopy of White Blood Cells, Cancer Cell Line MCF7, and Mixtures of these Cells: A Model System for Circulating Tumor Cell Biomarker Evaluation V.1

<p>This is a confocal laser scanning microscopy data set of white blood cells (leukocytes), the cancer cell line MCF7, and mixtures of these cells acquired on a Zeiss LSM 780 microscope in the University of Colorado Anschutz Medical Campus Advanced Light Microscopy Core. Cells are fluorescently labeled for DNA with DAPI (Sigma D9542), lipids with Bodipy 495/503 (Thermo Fisher D3922), the filament protein cytokeratin (CK) with pan-cytokertain-alexa555 antibodies (Cell Signaling Technologies 3478S) and the surface membrane antigen CD45 with CD45-alexa647 antibodies (Biolegend 304020). Bodipy was excited with a continuous wave (CW) 488 nm laser, alexa555 was excited with CW 561 nm laser, and alexa647 was excited with a CW 633 nm laser. The acquiring instrument does not have a CW 405 nm source so DAPI was excited by two photon process using a Coherent Cameleon ultrafast pulsed laser tuned to 765 nm. The objective used was a Zeiss Plan-Apochromat 20x, 0.8 NA, air.</p> <p>The data consists of 4 channel 8x8 mosaic z-stacks. The Zeiss software performed stitching of the mosaics. These stitched data images are included and marked with _Stitched at the end. Those interested in performing the stitching themselves can do this with the raw data files (without the _Stitched). The jpeg images are processed from the stitched LSM images. The LSM files contain additional meta data on the experiment including power levels and acquisition settings.</p> <p>The _Stiched .lsm files will load in ImageJ (tested with V.1.49) as 4 channel 3 stack images.</p> <p>This data is a model system for evaluating the DNA/Lipids/CK/CD45 biomarker panel to identify circulating tumor cells (CTCs). The D- population of the model is the WBCs and the D+ population is the MCF7 cancer cell line. The amount of separation the biomarker panel plus analysis algorithm can produce between these populations (D+/D-) is an estimate the sensitivity and specificity of the biomarker panel plus algorithm to CTCs.</p> <p>Experiments generating the data were performed over the course of 15 days. Peripheral blood samples were collected from the Gynecological Tissue and Fluid Bank (COMIRB 07-0935 / COMIRB 05-1081)&nbsp;from consenting patients undergoing surgery at the University of Colorado Hospital. Blood samples were used the same day they were collected. Blood samples were collected from 3 patients with benign conditions, labeled WBBN#, and 3 patients with ovarian cancer, labeled WBCA#. We do not expect there to be any difference in the isolated white blood cells samples prepared from the cancer and benign patients. Samples were stored at room temperature until white blood cells were isolated. Mixed samples were prepared by passaging a MCF7 flask and mixing it with isolated white blood cells before fixation. A schedule showing the time duration between collection, processing and imaging is included as &ldquo;experimental schedule.gif&rdquo;.</p> <p>The MCF7 cancer cell line was a kind gift from Dr. Heide Ford. Genomic DNA was isolated from the MCF7 cell line after the experiment and sent for cell line authentication. The gDNA was a match to MCF7. The authentication report and data are included in this submission.</p> <p>CD45 antibodies were exhausted on day 7. New antibody was purchased and received on day 8. The day 7 images only has labels for DAPI and Bodipy. The samples prepared with the old antibodies on days 4 and 7 were relabeled and imaged with the new antibodies on days 14 and 15. This labeling was also done to confirm the pan-CK antibodies remained good since they are dim in the MCF7 cells imaged on days 12 and 13. The pan-CK on days 14 and 15 looks the same as it did on days 5 and 7 confirming the antibodies are good.</p> <p>Four of the filters containing cells were not sufficiently flat to be acquired with a 3 slice z-stack so a 5 slice z-stack was used. These files have been zipped to compress them under the 2 GB limit permitted by zenodo.org</p> <p>Further information on how these samples were prepared, processed, and analyzed can be found in our associated 2016 SPIE Photonics West BIOS conference proceeding titled, &ldquo;Quantitative image cytometry measurements of lipids, DNA, CD45 and cytokeratin for circulating tumor cell identification in a model system&rdquo;, http://dx.doi.org/10.1117/12.2222317.</p> <p>This work was supported by funding provided to the University of Colorado Cancer Center by the American Cancer Society and awarded as Institutional Research Grant Number 57-001-53, by funding provided by the Defense Advanced Research Projects Agency under grant number N66001-10-4035, and by funding provided by NIH/NCATS Colorado CTSI Grant Number TL1 TR001081. The University of Colorado Anschutz Medical Campus Advanced Light Microscopy Core is also supported in part by NIH/NCATS Colorado CTSI Grant Number UL1 TR001082. The funders had no role in the study design, data collection, analysis, or&nbsp;decision to publish.</p>

opencc-by-4.0Apr 2016View details →
zenodo44/100

DWCox: A Density-Weighted Cox Model for Outlier-Robust Prediction of Prostate Cancer Survival

<p>This package, <strong>DWCox</strong>, implements a <strong>d</strong>ensity-<strong>w</strong>eighted <strong>Cox</strong> regression model that is more robust against outliers in the training data. DWCox gives more accurate predictions than the standard Cox regression on prostate cancer survival, especially in cases where the training data are expected to contain a lot of outliers. More details can be found in our paper (coming soon) and the README file inside this package.</p>

openmit-licenseNov 2016View details →
zenodo44/100

Suplementary data, results and scripts: "Reconstruction of Cell-specific Models Capturing the Influence of Metabolism on DNA methylation in Cancer"

<p>This repository contains supplementary data, models and scripts associated with "Reconstruction of Cell-specific Models Capturing the Influence of Metabolism on DNA methylation in Cancer".</p><p>Folders content:</p><p>'data_results_matlabscripts': data, result files and scripts (original python scripts and adapted MATLAB scripts)</p><p>'supplementary_figures': supplementary figures</p><p>'supplementary_tables': supplementary tables</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Supplementary material - Optical Diffraction Tomography and Raman Confocal Microscopy for the Investigation of Vacuoles Associated with Cancer Senescent Engulfing Cells

<p>Supplementary material containing the data used in the manuscript &quot;Optical Diffraction Tomography and Raman Confocal Microscopy for the Investigation of Vacuoles Associated with Cancer Senescent Engulfing Cells&quot;</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

BeBOD estimates of incidence, prevalence, and years lived with disability for 57 cancer types, 2004-2021

<p><strong>Belgian National Burden of Disease Study</strong></p><p><strong>Estimates of the morbidity burden of disease for 57 cancer sites</strong></p><p><i>Incidence</i></p><p>Data on new cancer cases in Belgium are collected by the&nbsp;<a href="https://kankerregister.org/Annual%20Tables">Belgian Cancer Registry</a> (BCR). For the current study, we selected 80 ICD-10 (C00.0-96.9 and chronic myeloid neoplasms) codes resulting in 57 cancer sites. Data were extracted by year (from 2004 to 2021), age group (5-years), sex and region (N=3). We excluded "Respiratory system and intrathoracic organs, NOS (not otherwise specified)" from further analyses because of too few cases.</p><p><i>Prevalence</i></p><p>Prevalence estimates were estimated using the above-described incidence estimates and the survival estimates also provided by BCR, derived from linkage with the Belgian Crossroads Bank for Social Security. We used a 10-year prevalence perspective meaning that from the year 2013 onwards, we were able to define the prevalence in a given year as the sum of person-months spent in the different health states. Specifically, we used a microsimulation approach to simulate future health states for each year-, age-, sex-, region- and cancer-specific cohort of incident cases.</p><p>See for more details: <a href="https://doi.org/10.1186/s12885-021-09109-4">https://doi.org/10.1186/s12885-021-09109-4</a></p><p><i>Years&nbsp;Lived with Disability</i></p><p>Years Lived with Disability (YLDs) were calculated using both an incidence and prevalence perspective&nbsp;as a measure of morbidity. YLDs are calculated as the product of the number of prevalent cases with the disability weight (DW), averaged over the different health states of the disease. The DWs reflect the relative reduction in quality of life, on a scale from 0 (perfect health) to 1 (death). We calculate YLDs using the Global Burden of Disease DWs.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Identification of biomarkers for the early detection of non-small cell lung cancer: a systematic review and meta-analysis

<p>We sought to identify the best biomarkers for the early diagnosis of LC, using a systematic review of seven databases. We identified 79 articles that focused on the identification and assessment of diagnostic biomarkers and then performed a meta-analysis. This work has been submitted for publication.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Cancer-specific RBP-USE pairs

<ol><li>all_predictions_soft_tr_paired.tsv - table of all predictions of cancer-specific RBP-USE pairs made by applying soft thresholds to dPSI values. No eCLIP support is required. Columns contain info about gene names, USE position, splicing and gene expression changes of genes, its USEs and regulators between cancer vs normal tissue or KD vs control samples.</li><li>best_predictions_soft_tr_paired.tsv - subset of the first table, containing only rows with |log2FC|&gt;0.3 and eCLIP support.</li></ol>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Socioeconomic status and adiposity in childhood cancer survivors: A cross-sectional retrospective study

<p>This dataset contains information on selected indicators of socioeconomic status and anthropometric indicators of adiposity in a population of childhood cancer survivors from the Late Effect Outpatient Clinic at St. Anne's Hospital in Brno, Czech Republic.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

CanVaxKB: A Web-based Cancer Vaccine Knowledgebase

<p>CanVaxKB is a web-based cancer vaccine knowledgebase. CanVaxKB collects, annotates and analyzes various types of cancer vaccines around the world. Currently it contains all cancer vaccines stored in the VIOLIN vaccine database. CanVaxKB also provides a user-friendly web interface for users to interactively search, compare, and analyze different cancer vaccines. The CanVaxKB website is here: https://violinet.org/canvaxkb.&nbsp;&nbsp;</p> <p>The Vaccine Ontology (VO) also includes the CanVaxKB stored cancer vaccine information, which is accessible at: https://github.com/vaccineontology/VO.&nbsp;</p> <p>The four supplemental files provided here are for the NCI Cancer paper about CanVaxKB. The citation for the CanVaxKB NCI Cancer is here:</p> <p>Eliyas Asfaw*, Asiyah Yu Lin*, Anthony Huffman*, Siqi Li*, Madison George*, Chloe Darancou, Madison Kalter, Nader Wehbi, Davis Bartels, Elyse Fleck, Nancy Tran, Daniel Faghihnia, Kimberly Berke, Ronak Sutariya, Farah Reyal, Youssef Tammam, Bin Zhao, Edison Ong, Zuoshuang Xiang, Virginia He, Justin Song, Andrey I. Seleznev, Jinjing Guo, Yuanyi Pan, Jie Zhang, Yongqun He. CanVaxKB: A Web-based Cancer Vaccine Knowledgebase. NCI Cancer. In press.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

2 million histological images of breast cancer tumors with her2 labels

<p><strong>Data Description</strong><br> This is a 2 million set of non-overlapping image patches from hematoxylin &amp; eosin (H&amp;E) stained histological images of human breast cancer tumor tissue.</p> <p>The anonymized dataset comes from a cohort of BC patients from the A. C. Camargo Cancer Center (ACCCC, N = 504). All patients were treated for breast cancer at the ACCCC between 2019 and 2021. As part of their diagnosis, in HER2 IHC score 2+ cases, patients&#39; HER2 status was determined following the ASCO guidelines updated in 2018, with visual evaluation of IHC assay and either a FISH or DDISH test. All cases with metastasis or neoadjuvant treatment were excluded.</p> <p>A total of 426 H&amp;E stained high resolution images (40x magnification) were scanned from biopsy and resection tissue samples with a Leica Aperio AT2 scanner. Ethical approval of the ACCCC study was given by the ethics committee of the Funda&ccedil;&atilde;o Ant&ocirc;nio Prudente. We divided the cases into the following 3 groups according to the results of the IHC and ISH tests: HER2-negative, HER2-low and HER2-high.</p> <p>The slides were divided into 256 px x 256 px tiles at 0.5 um/pixel magnification. Then, we used a custom trained ConvNext-tiny neural network to only include tiles from the tumor region and its environment, generating a total of 2051877 image patches.</p> <p>A sample is considered her2-negative with an IHC score of 0; her2-low with an IHC score of 1+ or an IHC score of 2+ with a negative ISH-based test result, and her2-high with an IHC score of 2+ with a positive ISH-based test or an IHC score of 3+.</p> <p>The accompanying code used for training&nbsp;the models is available at https://github.com/tojallab/wsi-mil</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Artificial Intelligence Enables Precision Diagnosis of Cervical Cytology Grades and Cervical Cancer

<p>This repository includes source data used to genrtate all tables and figures&nbsp; for published stduy "Artificial Intelligence Enables Precision Diagnosis of Cervical Cytology Grades and Cervical Cancer". Besides, a small set of digital images for different class of cervical smear samples are included.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Differential DNA methylation in the benign and cancerous prostate tissue of African American and European American men

<p>The data presented here are the summary statistics for the manuscript,&nbsp;"Differential DNA methylation in the benign and cancerous prostate tissue of African American and European American men." The study aims to improve our understanding of prostate cancer disparities between African American and European American men by comparing the DNA methylation features that distinguish tumor and paired, histologically benign tissue from a sample of African American and European American prostate cancer patients. The summary statistics presented here represent the results of a differential methylation analyses comparing tumor and benign tissue in each ancestry group as well as the results of an analysis of differential methylation by ancestry group within each tissue. &nbsp;</p> <p>The files included are:&nbsp;</p> <p>AA_TumorvBenign_Dummy.zip which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in African Americans based on a model that accounts for the paired nature of samples using a series of dummy variables for individual.&nbsp;</p> <p>AA_TumorvBenign_MixedModel.zip which which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in African Americans based on a model that accounts for the paired nature of samples using a linear mixed model that included patient as a random effect.&nbsp;</p> <p>EA_TumorvBenign_Dummy.zip which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in European Americans based on a model that accounts for the paired nature of samples using a series of dummy variables for individual.&nbsp;</p> <p>EA_TumorvBenign_MixedModel.zip which which contains the results of the association analysis between tumor vs benign (benign as the default) tissue status and individual CpG sites in European Americans based on a model that accounts for the paired nature of samples using a linear mixed model that included patient as a random effect.&nbsp;</p> <p>Benign_AncestryCompare_Summary.zip which contains the results of an association analysis between ancestry designation (African American vs European American with African American as the baseline) and individual CpG sites in benign tissue.&nbsp;</p> <p>Tumor_AncestryCompare_Summary.zip which contains the results of an association analysis between ancestry designation (African American vs European American with African American as the baseline) and individual CpG sites in tumore tissue.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data to "Phantom-based quality assurance for multicenter quantitative MRI in locally advanced cervical cancer"

<p>This record includes the DICOM images and analysed data that were used in the multicenter QA program for quantitative MRI in cervical cancer as published (<a href="https://www.sciencedirect.com/science/article/pii/S0167814020307854?via%3Dihub">https://doi.org/10.1016/j.radonc.2020.09.013</a> ).</p> <p>The DICOM data includes the acquired DICOM data for each institute selected to those that were used in the publication. Acquisitions that were not used were removed. Data was anonymized with conquest dicom server tools.</p> <p>The analyzed data files are included giving per measurement the estimated quantitative parameter values as well as the position of the ROIs and extracted signal intensity values per phantom sample. An explanation of the structure of the files is added in the readme file. The analysis was done with in-house written code in matlab.</p> <p>Included are a description of the sequence parameters for each institute (IQEMBRACE_PhantomQA_OverviewInstitutionalSequenceParameters_20241114) and details on the choices in the analysis of the data (IQEMBRACE_PhantomQA_OverviewPhantomData_20241114). As background also the description of the measurements was added, giving more information on how the measurements were performed.</p> <p>This work was in preparation for the IQ-EMBRACE trial (clinicaltrials.gov NCT03210428)</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

OpenAIRE Dataset for SciLake cancer research pilot

<p>This dataset is related to the subset of the OpenAIRE graph relevant to the cancer research pilot. The dataset is built according to the <a href="https://graph.openaire.eu/docs/data-model/">data model</a> of the OpenAIRE Graph dataset.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Catalog of stool metagenome-assembled genomes from patients with different cancer types

<p><strong>A non-redundant catalog of 3,816 genomes with at least 75% completeness and no more than 15% contamination assembled from metagenomes. Samples of 976 metagenomes were obtained from patients receiving immunotherapy for the treatment of different types of cancers.</strong></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Multi-scale simulations for optimizing cancer treatment

<p>The dataset comprises the output of several simulations of a model of tumor growth with different parameter values (10.1101/2021.12.17.473136). The model is a multi-scale agent-based model of a tumor spheroid that is treated with periodic pulses of the cytokine tumor necrosis factor (TNF). The multi-scale model&nbsp;simulates processes including i) the diffusion, uptake, and secretion of molecular entities such as oxygen, or TNF; ii) the mechanical interaction between cells; and iii) cellular processes including cell life cycle, cell death models, signal transduction.</p> <p>The multi-scale model was implemented and simulated using&nbsp;the PhysiBoSS&nbsp;framework (Letort et al. 2019). The dataset corresponds to 425 different simulations launched and automatically tagged as Interesting/Non-Interesting based on the effect of the parameters on the simulation (see README file). Each simulation&nbsp;was tagged by them with the following parameters (in that order):</p> <ul> <li>oxygen_necrotic, oxygen_critical</li> <li>oxygen_no_proliferation</li> <li>oxygen_reference</li> <li>initial_uptake_rate</li> <li>protein_threshold</li> <li>secretion_rate</li> <li>oxygen_concentration</li> <li>tnf_concentration</li> </ul> <p>Details on how these files are built can be found in the <strong>Biological Use Case</strong> output format file (<a href="https://zenodo.org/record/3921049">https://zenodo.org/record/3921049</a>).&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Oral cancer speech corpus for paper "Detecting and analysing spontaneous oral cancer speech in the wild"

<p>This is the oral cancer speech corpus used in the paper <em>&quot;Detecting and analysing spontaneous oral cancer speech in the wild&quot;.</em></p> <p><strong>Description</strong></p> <p>This dataset contains approximately 3 hours of oral cancer speech data collected from YouTube, including a file with additional metadata. We use this dataset to perform an oral cancer speech detection task in our paper.</p> <p><strong>Funding</strong></p> <p>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under Marie Sklodowska-Curie grant agreement No 766287. The Department of Head and Neck Oncology and surgery of the Netherlands Cancer Institute receives a research grant from Atos Medical (Horby, Sweden),<br> which contributes to the existing infrastructure for quality of life research.</p> <p><strong>Citation:</strong></p> <p>If you use this dataset please cite:</p> <pre><code>@misc{halpern2020detecting, title={Detecting and analysing spontaneous oral cancer speech in the wild}, author={Bence Mark Halpern and Rob van Son and Michiel van den Brekel and Odette Scharenborg}, year={2020}, eprint={2007.14205}, archivePrefix={arXiv}, primaryClass={eess.AS} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Mar 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record