Skip to main content
zenodoopen

Chest Radiograph at Diverse Institutes (CRADI) dataset

<p><strong>Introduction to the Chest Radiograph at Diverse Institutes (CRADI) dataset</strong></p> <p>&nbsp;</p> <p><strong>Background</strong></p> <p>&nbsp;</p> <p>Chest radiography is extensively used to screen and diagnose pulmonary and cardiac diseases. The advantages of its clinical practicality, efficiency, and cost-effectiveness make chest radiography the most accessible imaging test for pulmonary disorders, especially in primary hospitals. Currently, the interpretation of a chest radiograph mainly relies on radiologists.</p> <p>&nbsp;</p> <p>With the development of algorithms, convolutional neural networks (CNNs) have shown the ability to detect a single disorder in chest radiography, e.g., pneumothorax, lung cancer, pneumonia, and tuberculosis. Traditionally, expert annotation is applied to establish a CNN model for classifying medical images. This manual labeling procedure is time-consuming and highly demanding. Importantly, beyond the detection of a single disease, multi-label classification is necessary to interpret a chest radiograph in clinical practice.</p> <p>&nbsp;</p> <p>In order to promote the development of the artificial intelligence-assisted diagnosis of chest radiography, we launched the Chest Radiograph at Diverse Institutes (CRADI) dataset. This dataset is comprised of a large number of chest radiographs. Each radiograph has a 25-label disorder annotation, that was established by the terms adopted from the Fleischner&rsquo;s glossary, and extracted from the original diagnostic report by natural language processing (NLP) and radiologist expertise.</p> <p>&nbsp;</p> <p>At present, the data of the CRADI dataset comes from two academic hospitals and multiple community clinics in Shanghai. The cases in the CRADI dataset are comprised of in-patients, out-patients, and screening participants.</p> <p>&nbsp;</p> <p>The CRADI dataset provides a better understanding of the multiple and different clinical data sources for chest radiography, which is potentially helpful for the training and test of CNN models.</p> <p>&nbsp;</p> <p>We welcome more data into the CRADI dataset. If you find it helpful to your research work or you want to contribute to this dataset, please feel free to contact us.</p> <p>&nbsp;</p> <p><strong>Data source</strong></p> <p><strong>Number of images</strong></p> <p><strong>Data_source</strong></p> <p>Academic hospital 1</p> <p>74,082</p> <p>0</p> <p>In- and out-patient from Academic hospital 2</p> <p>5,996</p> <p>1</p> <p>Screening participants from Academic hospital 2</p> <p>2,130</p> <p>2</p> <p>Community clinics</p> <p>1,804</p> <p>3</p> <p>Note. Each case includes one posterior-anterior (PA) view chest radiograph and the corresponding text label of disorder findings.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Preprocessing methods</strong></p> <p>Images: transformed and resized from DCM format to PNG, changed from 12-bit grayscale to 8-bit. All patient- or institute-related information is de-identified.</p> <p>Label: labels are extracted from the original diagnostic reports. The regular expression is applied by NLP and rules-based extraction methods. In total, 25 labels are extracted for each image.</p> <p><strong>Data</strong></p> <p><strong>Overview: </strong>All images are compressed into one file. All classification labels are listed in one CSV file. Data order is as following way:</p> <p><strong>Data organization</strong></p> <p><strong>Classification result</strong></p> <p>Each image links to the label by an item of &lsquo;patientID&rsquo;.</p> <p>data_resource<a href="#_msocom_1">[1]</a>&nbsp; stands for the resource of data. Data sources are listed in the previous table.</p> <p>Result table format: |pateintID|data_resource<a href="#_msocom_2">[2]</a>&nbsp;|label1|label2|.......|label25|</p> <p>The order of the 25 labels is as the following:</p> <p>1) pneumothorax, 2) emphysema, 3) pulmonary parenchymal calcification, 4) PICC implant, 5) aortic unfolding, 6) aortic arteriosclerosis, 7) aortic abnormalities, 8) small consolidation, 9) cardiomegaly, 10) patchy consolidation, 11) consolidation, 12) cavity, 13) mass, 14) prominent bronchovascular marking, 15) pulmonary edema, 16) pulmonary nodule, 17) hilar adenopathy, 18) pleural effusion, 19) pleural thickening, 20) pleural adhesion, 21) pleural calcification, 22) pleural abnormalities, 23) scoliosis, 24) pacemaker implant, 25) interstitial involvement.</p> <p>Data_resource or data_source?</p> <p>同上</p> <p><strong>Introduction to the Chest Radiograph at Diverse Institutes (CRADI) dataset</strong></p> <p>&nbsp;</p> <p><strong>Background</strong></p> <p>&nbsp;</p> <p>Chest radiography is extensively used to screen and diagnose pulmonary and cardiac diseases. The advantages of its clinical practicality, efficiency, and cost-effectiveness make chest radiography the most accessible imaging test for pulmonary disorders, especially in primary hospitals. Currently, the interpretation of a chest radiograph mainly relies on radiologists.</p> <p>&nbsp;</p> <p>With the development of algorithms, convolutional neural networks (CNNs) have shown the ability to detect a single disorder in chest radiography, e.g., pneumothorax, lung cancer, pneumonia, and tuberculosis. Traditionally, expert annotation is applied to establish a CNN model for classifying medical images. This manual labeling procedure is time-consuming and highly demanding. Importantly, beyond the detection of a single disease, multi-label classification is necessary to interpret a chest radiograph in clinical practice.</p> <p>&nbsp;</p> <p>In order to promote the development of the artificial intelligence-assisted diagnosis of chest radiography, we launched the Chest Radiograph at Diverse Institutes (CRADI) dataset. This dataset is comprised of a large number of chest radiographs. Each radiograph has a 25-label disorder annotation, that was established by the terms adopted from the Fleischner&rsquo;s glossary, and extracted from the original diagnostic report by natural language processing (NLP) and radiologist expertise.</p> <p>&nbsp;</p> <p>At present, the data of the CRADI dataset comes from two academic hospitals and multiple community clinics in Shanghai. The cases in the CRADI dataset are comprised of in-patients, out-patients, and screening participants.</p> <p>&nbsp;</p> <p>The CRADI dataset provides a better understanding of the multiple and different clinical data sources for chest radiography, which is potentially helpful for the training and test of CNN models.</p> <p>&nbsp;</p> <p>We welcome more data into the CRADI dataset. If you find it helpful to your research work or you want to contribute to this dataset, please feel free to contact us.</p> <p><strong>Data source&nbsp;&nbsp;</strong></p> <p>training data data source: 0</p> <p>In- and out-patient from external hospital data source: 1</p> <p>Screening participants from external hospital data source : 2</p> <p>Community clinics datasource: 3</p> <p>Note. Each case includes one posterior-anterior (PA) view chest radiograph and the corresponding text label of disorder findings.</p> <p><strong>Preprocessing methods</strong></p> <p>Images: transformed and resized from DCM format to PNG, changed from 12-bit grayscale to 8-bit. All patient- or institute-related information is de-identified.</p> <p>Label: labels are extracted from the original diagnostic reports. The regular expression is applied by NLP and rules-based extraction methods. In total, 25 labels are extracted for each image.</p> <p><strong>Data</strong></p> <p><strong>Overview: </strong>All images are compressed into one file. All classification labels are listed in one CSV file. Data order is as following way:</p> <p><strong>Classification result</strong></p> <p>Each image links to the label by an item of &lsquo;patientID&rsquo;.</p> <p>data_resource<a href="#_msocom_1">[1]</a>&nbsp; stands for the resource of data. Data sources are listed in the previous table.</p> <p>Result table format: |pateintID|data_resource<a href="#_msocom_2">[2]</a>&nbsp;|label1|label2|.......|label25|</p> <p>The order of the 25 labels is as the following:</p> <p>1) pneumothorax, 2) emphysema, 3) pulmonary parenchymal calcification, 4) PICC implant, 5) aortic unfolding, 6) aortic arteriosclerosis, 7) aortic abnormalities, 8) small consolidation, 9) cardiomegaly, 10) patchy consolidation, 11) consolidation, 12) cavity, 13) mass, 14) prominent bronchovascular marking, 15) pulmonary edema, 16) pulmonary nodule, 17) hilar adenopathy, 18) pleural effusion, 19) pleural thickening, 20) pleural adhesion, 21) pleural calcification, 22) pleural abnormalities, 23) scoliosis, 24) pacemaker implant, 25) interstitial involvement.</p> <p>Data will be released after anonymilization process.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0