Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

30

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

30 results for “clinical coding”

Learn how ShareScore rates datasets ↗
zenodo44/100

CodiEsp corpus: gold standard Spanish clinical cases coded in ICD10 (CIE10) - eHealth CLEF2020

<p><strong>Introduction</strong></p> <p>These are the train, development and test sets of the CodiEsp corpus. Train, development and test have gold standard annotations. In addition, the unannotated background set is also distributed. All documents&nbsp;are released in the context of the CodiEsp track for CLEF ehealth 2020 (<a href="http://temu.bsc.es/codiesp/">http://temu.bsc.es/codiesp/</a>).</p> <p>The CodiEsp corpus contains manually coded clinical cases. All documents are in Spanish language and CIE10 is the coding terminology (it is the Spanish version of ICD10-CM and ICD10-PCS). The CodiEsp corpus has been randomly sampled into three subsets: the train, the development, and the test set. The train set contains 500 clinical cases, and the development and test set 250 clinical cases each. CodiEsp participants must submit predictions for the test and background set, but they will only be evaluated on the test set.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estap&eacute; and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</p> <pre><code>@inproceedings{miranda2020overview, title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2020} }</code></pre> <p>&nbsp;</p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 88.6% for diagnosis&nbsp;coding, 88.9% for procedure coding&nbsp;and 80.5% for the textual reference annotation. For more information, see the <a href="http://ceur-ws.org/Vol-2696/paper_263.pdf">paper</a>.</p> <p><br> <strong>Zip structure</strong><br> Four folders: train, dev, test and background. Each one of them contains the files for the train, development, test and background corpora, respectively.</p> <ul> <li><strong>train, dev and test</strong> folders have: <ul> <li>3 tab-separated files with the annotation information relevant for each of the 3 sub-tracks of CodiEsp.&nbsp;</li> <li>A subfolder named <em>text_files</em>&nbsp;with the plain text files of the clinical cases.</li> <li>A subfolder named <em>text_files_en</em>&nbsp;with the plain text files machine-translated to English. Due to the translation process, the text files are sentence-splitted.</li> </ul> </li> <li>The <strong>background</strong>&nbsp;folder has only <em>text_files</em>&nbsp;and <em>text_files_en</em>&nbsp;subfolders with the plain text files.</li> </ul> <p><br> <strong>Format</strong><br> The CodiEsp corpus is distributed in plain text in UTF8 encoding, where each clinical case is stored as a single file whose name is the clinical case identifier. Annotations are released in a tab-separated file. Since the CodiEsp track has 3 sub-tracks, every set of documents (train and test) has 3 tab-separated files associated with it.&nbsp;</p> <p>For the sub-tracks CodiEsp-D and CodiEsp-P, the file has the following fields:</p> <pre>articleID ICD10-code </pre> <p>Tab-separated files for the sub-track CodiEsp-X contain extra fields that provide the text-reference and its position:</p> <pre>articleID label ICD10-code text-reference reference-position</pre> <p><br> <strong>Corpus summary statistics</strong><br> The final collection of 1000 clinical cases that make up the corpus had a total of 16504 sentences, with an average of 16.5 sentences per clinical case. It contains a total of 396,988 words, with an average of 396.2 words per clinical case.</p> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/codiesp/">Web</a></strong></li> <li><strong><a href="http://ceur-ws.org/Vol-2696/paper_263.pdf">Citation</a>:&nbsp;</strong>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estap&eacute; and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</li> <li><strong><a href="https://doi.org/10.5281/zenodo.3859869">Silver Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3730566">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhA0crlSVCYMPqMUWd4mXc4x"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/codiesp/index.php/participants-systems/"><strong>Participant codes</strong></a></li> </ul> <p>&nbsp;</p> <p>For more information, visit the track webpage: http://temu.bsc.es/codiesp/&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>&nbsp;</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Data and Code Supplement to: "Processes of change in a randomized clinical trial of Radically Open Dialectical Behavior Therapy (RO DBT) for adults with treatment refractory depression"

<p>Dataset to support secondary&nbsp;analyses reported in&nbsp;&quot;Processes of change in a randomized clinical trial of Radically Open Dialectical&nbsp;Behavior Therapy (RO DBT) for adults with treatment refractory depression&quot; in the&nbsp;Journal of Consulting and Clinical Psychology</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Additional datasets and code accompanying the article "Evaluation of in silico algorithms for use with ACMG/AMP clinical variant interpretation guidelines"

<p> </p> <p> </p> <p>Datasets and code accompanying the manuscript titled “Evaluation of in silico algorithms for use with ACMG/AMP clinical variant interpretation guidelines” by Rajarshi Ghosh, Ninak Oak and Sharon E. Plon.</p>

opencc-by-4.0Oct 2017View details →
ClinicalTrials.gov36/100

Geriatric Core Dataset (G-CODE) for Clinical Research in Elderly Cancer Patients

ClinicalTrials.gov study NCT03976531. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo32/100

CodiEsp Silver Standard: Participant predictions in eHealth CLEF2020 - Spanish clinical cases coded in ICD10 (CIE10)

<p><strong>Introduction</strong></p> <p>Predictions in the background set of <a href="https://temu.bsc.es/codiesp/">eHealth CLEF 2020 Task 1</a> participants.</p> <p>&nbsp;</p> <p><strong>Zip structure</strong></p> <p>One directory per CodiEsp subtask. Within each CodiEsp subtask directory, there is one directory per team that contains the prediction runs.</p> <p>&nbsp;</p> <p><strong>Format</strong><br> The text documents are distributed in plain text files, UTF-8 encoding.<br> The CodiEsp Silver Standard annotations&nbsp;have the following format:</p> <p>For the sub-tracks CodiEsp-Diagnostic and CodiEsp-Procedure, the file files have&nbsp;the following fields:</p> <pre>articleID ICD10-code </pre> <p>Tab-separated files for the sub-track CodiEsp-X (explainability) contain extra fields that provide the text-reference and its position:</p> <pre>articleID label ICD10-code text-reference reference-position</pre> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/codiesp/">Web</a></strong></li> <li><strong>Citation:&nbsp;</strong>Miranda-Escalada, A., Farr&eacute;, E., &amp; Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In&nbsp;<em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><strong><a href="https://zenodo.org/record/3837305#.X7T9KVlKg5k">Gold Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p>&nbsp;</p> <p>All credit&nbsp;to CodiEsp participants</p>

opencc-by-4.0May 2020View details →
zenodo32/100

Cantemist Silver Standard: Participant predictions in SEPLN IberLEF2020 - Spanish oncology clinical cases coded in ICD-O

<p><strong>Introduction</strong></p> <p>Predictions in the background set of Cantemist&nbsp;participants.</p> <p>&nbsp;</p> <p><strong>Zip structure</strong></p> <p>One directory per Cantemist subtask. Within each Cantemist subtask directory, there is one directory per team that contains the prediction runs.</p> <p>&nbsp;</p> <p><strong>Format</strong></p> <p>The text documents are distributed in plain text files, UTF-8 encoding.<br> The CodiEsp Silver Standard annotations&nbsp;have the following format:</p> <p>For the sub-tracks Cantemist-NER and Cantemist-Norm, the files are in Brat format.</p> <p>For the sub-track Cantemist-Coding files have&nbsp;the following fields:</p> <pre>articleID ICDO-code </pre> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/cantemist/">Web</a></strong></li> <li><strong>Citation:&nbsp;</strong>Miranda-Escalada, A., Farr&eacute;, E., &amp; Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In&nbsp;<em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><a href="https://doi.org/10.5281/zenodo.3773228"><strong>Gold Standard corpus</strong></a></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p>&nbsp;</p> <p>All credit&nbsp;to Cantemist participants.&nbsp;</p> <p>&nbsp;</p> <p>For more information, visit the track webpage: <a href="http://temu.bsc.es/cantemist/">http://temu.bsc.es/cantemist/</a> or email us at encargo-pln-life@bsc.es</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Original features and code of clinical prediction model

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
ClinicalTrials.gov32/100

Clinical Significance of Heterozygosity for Mutations of the SLC12A3 Gene Coding for the Thiazide Sensitive Na-Cl Cotransporter

ClinicalTrials.gov study NCT02035046. IPD Sharing: Not stated. Countries: 1. Publications: 6.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

A Clinical Study of PD-L1 Antibody ZKAB001(Drug Code) in Recurrent or Metastatic Cervical Cancer

ClinicalTrials.gov study NCT03676959. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

A Clinical Study of PD-L1 Antibody ZKAB001(Drug Code) in Limited Stage of High-grade Osteosarcoma

ClinicalTrials.gov study NCT03676985. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Clinical Evaluation of Polyherbal Coded Formulation Obesecure for Leptin Regulation and Obesity Management

ClinicalTrials.gov study NCT04443790. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov28/100

AcrySof IQ Toric A-Code Post-Market Clinical Study

ClinicalTrials.gov study NCT03350503. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
geo24/100

Long non-coding RNA HOTAIR is expressed in small cell lung cancer and is relevant to cellular proliferation, invasiveness and clinical relapse: tissue and cell-line analyses

GEO Series GSE43877. Homo sapiens. 2 samples. Type: Expression profiling by array.

openGEO-OpenMar 2014View details →
geo24/100

Profile of protein-coding and long non-coding RNA expression and their clinical importance in T-cell acute lymphoblastic leukemia

GEO Series GSE216117. Homo sapiens. 35 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2024View details →
geo24/100

Comprehensive Analysis of Recurrence-Associated Small Non-Coding RNAs in Esophageal Cancer [clinical study, Illumina]

GEO Series GSE66258. Homo sapiens. 108 samples. Type: Expression profiling by array.

openGEO-OpenJun 2016View details →
geo24/100

A functional genomics atlas enhanced by convolutional neural networks facilitates clinical interpretation of disease relevant variants in non-coding regulatory elements [ATAC-seq]

GEO Series GSE263338. Homo sapiens. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenNov 2025View details →
geo24/100

A functional genomics atlas enhanced by convolutional neural networks facilitates clinical interpretation of disease relevant variants in non-coding regulatory elements [ChIP-seq]

GEO Series GSE263337. Homo sapiens. 16 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenNov 2025View details →
geo24/100

A functional genomics atlas enhanced by convolutional neural networks facilitates clinical interpretation of disease relevant variants in non-coding regulatory elements [wt RNA-seq]

GEO Series GSE267549. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2025View details →
geo24/100

Transcriptome analysis of spinal tuberculosis associated long non coding RNAs in human clinical bone tissues

GEO Series GSE268366. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2024View details →
geo24/100

A functional genomics atlas enhanced by convolutional neural networks facilitates clinical interpretation of disease relevant variants in non-coding regulatory elements [STARR-RNA-seq]

GEO Series GSE263335. Homo sapiens. 16 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record