Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
100
datasets available to search
ShareScore release 0.9.0
Dataset results
100 results for “Medical Dataset”
Multi-modality medical image dataset for medical image processing in Python lesson
<p>This dataset contains a collection of medical imaging files for use in the <a href="https://github.com/esciencecenter-digital-skills/medical-image-processing">"Medical Image Processing with Python" lesson</a>, developed by the <a href="https://www.esciencecenter.nl/">Netherlands eScience Center</a>. </p> <p>The dataset includes:</p> <ol> <li>SimpleITK compatible files: MRI T1 and CT scans (<em>training_001_mr_T1.mha, training_001_ct.mha</em>), digital X-ray (<em>digital_xray.dcm</em> in DICOM format), neuroimaging data (<em>A1_grayT1.nrrd, A1_grayT2.nrrd</em>). Data have been downloaded from <a href="https://insightsoftwareconsortium.github.io/SimpleITK-Notebooks/Python_html/00_Setup.html">here</a>. </li> <li>MRI data: a T2-weighted image (<em>OBJECT_phantom_T2W_TSE_Cor_14_1.nii</em> in NIfTI-1 format). Data have been downloaded from <a href="../records/6467772">here</a>. </li> <li>Example images for the machine learning lesson: chest X-rays (<em>rotatechest.png, other_op.png</em>), cardiomegaly example (<em>cardiomegaly_cc0.png</em>).</li> <li>Array data: Array data for the Intro to Medical Imaging lesson. Numpy arrays were created by processing and manipulation of publicly available data i.e. from <a href="https://doi.org/10.1109/TNS.1974.6499235">the Schepp Logan phantom</a> and from the <a href="https://fastmri.med.nyu.edu/">NYU FastMRI dataset</a></li> <li>Data for the anonymization exercises: ultrasound (<em>identifiable_us.jpg</em>) dowloaded from <a href="https://www.flickr.com/photos/jcarter/2461223727">here</a>, and DICOM data (<em>our_sample_dicom.dcm</em>) shared for this course specifically by a colleague</li> <li>Histopathology data: histopathology slide images from <a href="https://openslide.org/">openslide</a> library samples in the freely distributable test data </li> </ol> <p>These files represent various medical imaging modalities and formats commonly used in clinical research and practice. They are intended for educational purposes, allowing students to practice image processing techniques, machine learning applications, and statistical analysis of medical images using Python libraries such as scikit-image, pydicom, and SimpleITK.</p>
RIGA+ Dataset for Unsupervised Domain Adaptation in Medical Image Segmentation
<p>Different from the previous combined multi-domain dataset for unsupervised domain adaptation (UDA) in medical image segmentation, this multi-domain fundus image dataset contains annotations made by the same group of ophthalmologists. Hence the annotator bias among different datasets can be mitigated. Therefore, this dataset can provide a relatively fair benchmark for evaluating UDA methods in fundus image segmentation.</p> <p>This dataset is based on the RIGA[1] dataset and MESSIDOR[2] dataset. We appreciate their efforts devoted by the authors of [1] and [2].</p> <p>The six duplicated cases in the RIGA dataset are filtered out according to the <a href="https://www.adcis.net/en/third-party/messidor/">Errata</a>. We also remove the duplicated cases that exist in both the RIGA dataset and the MESSIDOR dataset by hash value matching.</p> <table align="center"> <caption>Details of the RIGA+ dataset</caption> <thead> <tr> <th scope="row">Domain</th> <th scope="col">Dataset</th> <th scope="col"> <p>Labeled Samples</p> <p>(Train+Test)</p> </th> <th scope="col"> <p>Unlabeled</p> <p>Samples</p> </th> </tr> </thead> <tbody> <tr> <th scope="row">Source</th> <td>BinRushed</td> <td>195 (195+0)</td> <td>0</td> </tr> <tr> <th scope="row">Source</th> <td>Magrabia</td> <td>95 (95+0)</td> <td>0</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE1</td> <td>173 (138+35)</td> <td>227</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE2</td> <td>148 (118+30)</td> <td>238</td> </tr> <tr> <th scope="row">Target</th> <td>MESSIDOR-BASE3</td> <td>133 (106+27)</td> <td>252</td> </tr> </tbody> </table> <p>[1] Almazroa A, Alodhayb S, Osman E, et al. Retinal fundus images for glaucoma analysis: the RIGA dataset[C]//Medical Imaging 2018: Imaging Informatics for Healthcare, Research, and Applications. International Society for Optics and Photonics, 2018, 10579: 105790B.</p> <p>[2] Decencière E, Zhang X, Cazuguel G, et al. Feedback on a publicly distributed image database: the Messidor database[J]. Image Analysis & Stereology, 2014, 33(3): 231-234.</p> <p>If you find this dataset useful for your research, please consider citing the paper as follows:</p> <pre><code>@inproceedings{hu2022domain, title={Domain Specific Convolution and High Frequency Reconstruction based Unsupervised Domain Adaptation for Medical Image Segmentation}, author={Shishuai Hu and Zehui Liao and Yong Xia}, booktitle={International Conference on Medical Image Computing and Computer-Assisted Intervention}, year={2022}, organization={Springer} }</code></pre> <p> </p>
Chest X-Ray Image Dataset: A Resource for Medical Diagnosis and Machine Learning
<p>The Chest X-Ray Image Dataset is an extensive collection designed to support medical research and the development of diagnostic tools for COVID-19 detection. It consists of two distinct classes: COVID-19 affected X-ray images and normal X-ray images of the chest area, each covering the full lungs. This dataset provides a diverse range of X-ray images, capturing the unique characteristics of both healthy and COVID-19 affected lungs, making it an invaluable resource for training and testing machine learning models in medical image classification and analysis.</p>
Dataset for the paper "Social support and avoidance explain positive and negative effects of emotion recognition ability on well-being in medical students"
<p>Full reference of the paprer TBD</p>
Dataset for the manuscript titled "Assessment of Anti-Tuberculosis Medication Adherence and Associated Factors among Patients attending the Sokoto Specialist Hospital Tuberculosis Treatment Center, Nigeria – 2019".
<p>Dataset of the research entitled "Assessment of Anti-Tuberculosis Medication Adherence and Associated Factors among Patients attending the Sokoto Specialist Hospital Tuberculosis Treatment Center, Nigeria – 2019".</p>
Development of a Feedback Evaluation Tool for the Assessment of Telerehabilitation as a Teaching-Learning Tool for Medical Students Dataset
<p>Acknowledging that telemedicine is a growing medium of instruction, it is important for their be specific tools to measure student outcomes. This study aims to create a feedback evaluation tool to aid students express their concerns on telerehabilitation as a method of instruction, as well as aid the educators in receiving specific points to improve upon.</p> <p>The file attached is the collected data for the research on the Development of a Feedback Evaluation Tool for the Assessment of Telerehabilitation as a Teaching-Learning Tool for Medical Students. It contains the processed data from the Data Collection Forms as well as the results of the processed interview.</p>
Dataset for: Extent and types of gender-based discrimination against female medical students and physicians at five university hospitals in Germany – Results of an online survey
<p><strong><span>Objective</span></strong><span>: There is a gap in research on gender-based discrimination (GBD) in medical education and practice in Germany. This study therefore examines the extent and forms of GBD among female medical students and physicians in Germany. Causes, consequences and possible interventions of GBD are discussed.</span></p> <p><strong><span>Methods</span></strong><span>: Female medical students (n=235) and female physicians (n=157) from five university hospitals in northern Germany were asked about their personal experiences with GBD in an online survey on self-efficacy expectations and individual perceptions of the "glass ceiling effect" using an open-ended question regarding their own experiences with GBD. The answers were analyzed by content analysis using inductive category formation and relative category frequencies. </span></p> <p><strong><span>Results</span></strong><span>: From both interviewed groups, approximately 75% of each reported having experienced GBD. Their experiences fell into five main categories: sexual harassment with subcategories of verbal and physical, discrimination based on existing/possible motherhood with subcategories of structural and verbal, direct preference for men, direct neglect of women, and derogatory treatment based on gender.</span></p> <p class="MsoNoSpacing"><strong><span>Conclusion</span></strong><span>: The study contributes to filling the aforementioned research gap. At the hospitals studied, GBD is a common phenomenon among both female medical students and physicians, manifesting itself in multiple forms. Transferability of the results beyond the hospitals studied to all of Germany seems plausible. Much is known about the causes, consequences and effective countermeasures against GBD. Those responsible for training and employers in hospitals should fulfill their responsibility by implementing measures from the set of empirically evaluated interventions.</span></p>
Mapping of the dataset of the German National Regsitry for Rare Diseases (NARSE) to Observational Medical Outcomes Partnership Common Data Model (OMOP CDM)
<p>Mapping between the data set of the German National Registry for Rare Diseases ("Nationales Register für Seltene Erkrankungen"; <a href="https://www.narse.de/">NARSE</a>) to Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) using international standards.</p>
Dataset for MedCodER: A Generative AI Assistant for Medical Coding
Open the record for dataset details and reuse information.
Dataset for: Extent and types of gender-based discrimination against female medical students and physicians at five university hospitals in Germany – Results of an online survey
Open the record for dataset details and reuse information.
MedIMeta: A comprehensive and easy-to-use multi-domain multi-task medical imaging meta-dataset
<p>We introduce the Medical Imaging Meta-Dataset (MedIMeta), a novel multi-domain, multi-task meta-dataset designed to facilitate the development and standardised evaluation of ML models and cross-domain few-shot learning algorithms for medical image classification. MedIMeta contains 19 medical imaging datasets spanning 10 different domains and encompassing 54 distinct medical tasks, offering opportunities for both single-task and multi-task training. All tasks are standardised to the same format and readily usable in PyTorch or other ML frameworks. All datasets have been previously published with an open license that allows redistribution or we obtained an explicit permission to do so.</p> <p>Each dataset within the MedIMeta dataset is standardized to a size of 224 × 224 pixels which matches image size commonly used in pre-trained models. Furthermore, the dataset comes with pre-made splits to ensure ease of use and standardized benchmarking. We release a user-friendly Python package to directly load images for use in PyTorch.<br><br></p> <h3>Links</h3> <ul> <li>Project website: <a href="https://www.woerner.eu/projects/medimeta/" target="_blank" rel="noopener">https://www.woerner.eu/projects/medimeta/</a></li> <li>Data loading code (medimeta Python package): <a href="https://github.com/StefanoWoerner/medimeta-pytorch" target="_blank" rel="noopener">https://github.com/StefanoWoerner/medimeta-pytorch</a></li> <li>Data creation code: <a href="https://github.com/StefanoWoerner/medimeta-dataset-scripts" target="_blank" rel="noopener">https://github.com/StefanoWoerner/medimeta-dataset-scripts</a></li> </ul> <p> </p> <h3>Dataset Overview</h3> <table> <tbody> <tr> <td><strong>Dataset Name</strong></td> <td><strong>Dataset ID</strong></td> <td><strong>License</strong></td> <td><strong>Domain</strong></td> <td><strong>Task Names</strong></td> <td><strong>Task Targets</strong></td> <td><strong># Labels</strong></td> </tr> <tr> <td>AML Cytomorphology</td> <td>aml</td> <td>CC BY-SA 4.0</td> <td>Microscopy</td> <td>morphological class</td> <td>multi-class classification</td> <td>15</td> </tr> <tr> <td>Breast Ultrasound</td> <td>bus</td> <td>CC BY-SA 4.0</td> <td>Breast ultrasound</td> <td>case category<br>malignancy</td> <td>multi-class classification<br>binary classification</td> <td>3<br>2</td> </tr> <tr> <td>Colorectal Cancer Histopathology</td> <td>crc</td> <td>CC BY-SA 4.0</td> <td>Histopathology</td> <td>tissue class</td> <td>multi-class classification</td> <td>9</td> </tr> <tr> <td>Chest X-ray Multi-disease</td> <td>cxr</td> <td>CC BY-SA 4.0</td> <td>Chest X-ray</td> <td>disease labels<br>patient sex</td> <td>multi-label classification<br>binary classification</td> <td>14<br>2</td> </tr> <tr> <td>Dermatoscopy</td> <td>derm</td> <td>CC BY-SA 4.0</td> <td>Dermatoscopy</td> <td>disease category</td> <td>multi-class classification</td> <td>7</td> </tr> <tr> <td>Diabetic Retinopathy (Regular Fundus)</td> <td>dr_regular</td> <td>CC BY-SA 4.0</td> <td>Retinal fundus</td> <td>DR level<br>Overall quality<br>Artifact<br>Clarity<br>Field definition</td> <td>ordinal regression<br>binary classification<br>ordinal regression<br>ordinal regression<br>ordinal regression</td> <td>5<br>2<br>6<br>5<br>5</td> </tr> <tr> <td>Diabetic Retinopathy (Ultra-widefield Fundus)</td> <td>dr_uwf</td> <td>CC BY-SA 4.0</td> <td>Retinal fundus</td> <td>DR level</td> <td>ordinal regression</td> <td>5</td> </tr> <tr> <td>Fundus Multi-disease</td> <td>fundus</td> <td>CC BY-SA 4.0</td> <td>Retinal fundus</td> <td>disease presence<br>disease labels</td> <td>binary classification<br>multi-label classification</td> <td>2<br>45</td> </tr> <tr> <td>Glaucoma-specific fundus images</td> <td>glaucoma</td> <td>CC BY-SA 4.0</td> <td>Retinal fundus</td> <td>Glaucoma suspect</td> <td>binary classification</td> <td>2</td> </tr> <tr> <td>Mammography (Calcifications)</td> <td>mammo_calc</td> <td>CC BY-SA 4.0</td> <td>Mammography</td> <td>pathology<br>calc type<br>calc distribution</td> <td>binary classification<br>multi-label classification<br>multi-label classification</td> <td>2<br>14<br>5</td> </tr> <tr> <td>Mammography (Masses)</td> <td>mammo_mass</td> <td>CC BY-SA 4.0</td> <td>Mammography</td> <td>pathology<br>mass shape<br>mass margins</td> <td>binary classification<br>multi-label classification<br>multi-label classification</td> <td>2<br>8<br>5</td> </tr> <tr> <td>OCT</td> <td>oct</td> <td>CC BY-SA 4.0</td> <td>OCT</td> <td>disease class<br>urgent referral</td> <td>multi-class classification<br>binary classification</td> <td>4<br>2</td> </tr> <tr> <td>Axial Organ Slices</td> <td>organs_axial</td> <td>CC BY-NC-SA 4.0</td> <td>Abdominal CT</td> <td>organ label</td> <td>multi-class classification</td> <td>11</td> </tr> <tr> <td>Coronal Organ Slices</td> <td>organs_coronal</td> <td>CC BY-NC-SA 4.0</td> <td>Abdominal CT</td> <td>organ label</td> <td>multi-class classification</td> <td>11</td> </tr> <tr> <td>Sagittal Organ Slices</td> <td>organs_sagittal</td> <td>CC BY-NC-SA 4.0</td> <td>Abdominal CT</td> <td>organ label</td> <td>multi-class classification</td> <td>11</td> </tr> <tr> <td>Peripheral Blood Cells</td> <td>pbc</td> <td>CC BY-SA 4.0</td> <td>Microscopy</td> <td>cell class</td> <td>multi-class classification</td> <td>8</td> </tr> <tr> <td>Pediatric Pneumonia</td> <td>pneumonia</td> <td>CC BY-SA 4.0</td> <td>Chest X-ray</td> <td>pneumonia presence<br>disease class</td> <td>binary classification<br>multi-class classification</td> <td>2<br>3</td> </tr> <tr> <td>Skin Lesion Evaluation (Dermoscopy)</td> <td>skinl_derm</td> <td>CC BY-SA 4.0</td> <td>Dermatoscopy</td> <td>Diagnosis<br>Diagnosis grouped<br>Pigment Network<br>Blue Whitish Veil<br>Vascular Structures<br>Vascular Structures grouped<br>Pigmentation<br>Pigmentation grouped<br>Streaks<br>Dots and Globules<br>Regression Structures<br>Regression Structures grouped</td> <td>multi-class classification<br>multi-class classification<br>multi-class classification<br>binary classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>binary classification</td> <td>15<br>5<br>3<br>2<br>8<br>3<br>5<br>3<br>3<br>3<br>4<br>2</td> </tr> <tr> <td>Skin Lesion Evaluation (Clinical Photography)</td> <td>skinl_photo</td> <td>CC BY-SA 4.0</td> <td>Clinical skin imaging</td> <td>Diagnosis<br>Diagnosis grouped<br>Pigment Network<br>Blue Whitish Veil<br>Vascular Structures<br>Vascular Structures grouped<br>Pigmentation<br>Pigmentation grouped<br>Streaks<br>Dots and Globules<br>Regression Structures<br>Regression Structures grouped</td> <td>multi-class classification<br>multi-class classification<br>multi-class classification<br>binary classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>multi-class classification<br>binary classification</td> <td>15<br>5<br>3<br>2<br>8<br>3<br>5<br>3<br>3<br>3<br>4<br>2</td> </tr> </tbody> </table>
Dataset Of Diagnosing Medical Score Calculator Apps Paper
<p>Dataset Of Diagnosing Medical Score Calculator Apps Paper</p>
Dataset related to article:MULTIMODAL DIGITAL HEALTH TREATMENTS FOR CHRONIC MIGRAINE ASSOCIATED WITH MEDICATION OVERUSE HEADACHE: A LITERATURE APPRAISAL AND RESULTS OF A SINGLE‑ARM OPEN TRIAL (THE BE‑HOME PROGRAM)
<p><strong><span>Variables related to baseline, months 3, 6 and 12 referred to headache frequency, medication intake, GSE, HIT-6, BDI-II, MSQ and PCS, of patients with CM and MOH who participated to the BE-HOME trial.</span></strong></p>
Dataset: Justify your disability! A simulated medical evaluation as a robust novel stress induction paradigm in chronic pain patients.
<p><span>Maladaptive stress responses may exacerbate chronic widespread pain (CWP) and deserve further investigations. Yet, existing stress induction paradigms lack relevance for individuals with this condition. Hence, we developed the Social Benefits Stress Test (SBST), adapted from the Trier Social Stress Test. <span>Instead of a job interview, the main task consists in justifying the inability to work.</span> </span></p> <p><span><span>Forty women with CWP in the context of hypermobile Ehlers-Danlos syndrome or hypermobility spectrum disorders were included. They underwent a 30-min baseline, the new stress task and a recovery period. The psychophysiological stress response was captured using self-reported stress ratings, salivary cortisol and α-amylase levels, as well as continuous physiological monitoring of heart rate variability (HRV) and electrodermal activity (EDA).</span></span><span> </span></p> <p><span>Compared to baseline, the analysis revealed </span><span>a significant and transient increase in stress ratings </span><span>during the stress task,</span><span> </span><span>associated with a peak in salivary biomarkers concentrations.</span><span> </span><span>The </span><span>HRV signal analysis showed a significant </span><span>decrease in high frequency power (HF), and </span><span>increases in heart rate,</span><span> </span><span>low frequency power (LF) and in LF/HF ratio. The EDA analysis revealed a significant increase in skin conductance level (SCL) tonic component and skin conductance response (SCR). Subjective </span><span>stress ratings positively correlated with changes in salivary biomarkers, LF/HF ratio and EDA outcomes</span><span>.</span></p> <p><span>The SBST induced a reproducible moderate stress response across subjective and physiological measures in a population of CWP patients, validating this task as a relevant experimental model of social stress in chronic pain. The SBST is a useful tool to study the relationship between stress and chronic pain.</span></p>
Dataset for the paper "Association between professional identity, burnout, and mental health in medical students: A cross-sectional study"
<p><strong>Full reference of the paprer TBD</strong></p>
Dataset related to article "WITHDRAWAL FAILURE IN PATIENTS WITH CHRONIC MIGRAINE AND MEDICATION OVERUSE HEADACHE"
<p>The file contains baseline and follow-up data (SPSS FILES .sav) referred to the MOH-COST study, used for a secondary analysis </p>
Dataset for the paper: "Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims"
<p><strong>Overview</strong></p> <p>This dataset of medical misinformation was collected and is published by <a href="https://kinit.sk/">Kempelen Institute of Intelligent Technologies (KInIT)</a>. It consists of approx. 317k news articles and blog posts on medical topics published between January 1, 1998 and February 1, 2022 from a total of 207 reliable and unreliable sources. The dataset contains full-texts of the articles, their original source URL and other extracted metadata. If a source has a credibility score available (e.g., from Media Bias/Fact Check), it is also included in the form of annotation. Besides the articles, the dataset contains around 3.5k fact-checks and extracted verified medical claims with their unified veracity ratings published by fact-checking organisations such as Snopes or FullFact. Lastly and most importantly, the dataset contains 573 manually and more than 51k automatically labelled mappings between previously verified claims and the articles; mappings consist of two values: <em>claim presence </em>(i.e., whether a claim is contained in the given article) and <em>article stance </em>(i.e., whether the given article supports or rejects the claim or provides both sides of the argument).</p> <p>The dataset is primarily intended to be used as a training and evaluation set for machine learning methods for claim presence detection and article stance classification, but it enables a range of other misinformation related tasks, such as misinformation characterisation or analyses of misinformation spreading.</p> <p>Its novelty and our main contributions lie in (1) focus on medical news article and blog posts as opposed to social media posts or political discussions; (2) providing multiple modalities (beside full-texts of the articles, there are also images and videos), thus enabling research of multimodal approaches; (3) mapping of the articles to the fact-checked claims (with manual as well as predicted labels); (4) providing source credibility labels for 95% of all articles and other potential sources of weak labels that can be mined from the articles' content and metadata.</p> <p>The dataset is associated with the research paper "Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims" accepted and presented at ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '22). </p> <p>The accompanying <a href="https://github.com/kinit-sk/medical-misinformation-dataset">Github repository</a> provides a small static sample of the dataset and the dataset's descriptive analysis in a form of Jupyter notebooks.</p> <p>In order to obtain an access to the full dataset (in the CSV format), please, request the access by following the instructions provided below.</p> <p> </p> <p><strong>Note: </strong>Please, check also our <a href="https://doi.org/10.5281/zenodo.7737982">MultiClaim Dataset</a> that provides a more recent, a larger, and a highly multilingual dataset of fact-checked claims, social media posts and relations between them.</p> <p> </p> <p><strong>References</strong></p> <p>If you use this dataset in any publication, project, tool or in any other form, please, cite the following papers:</p> <pre><code>@inproceedings{SrbaMonantPlatform, author = {Srba, Ivan and Moro, Robert and Simko, Jakub and Sevcech, Jakub and Chuda, Daniela and Navrat, Pavol and Bielikova, Maria}, booktitle = {Proceedings of Workshop on Reducing Online Misinformation Exposure (ROME 2019)}, pages = {1--7}, title = {Monant: Universal and Extensible Platform for Monitoring, Detection and Mitigation of Antisocial Behavior}, year = {2019} }</code></pre> <pre><code>@inproceedings{SrbaMonantMedicalDataset, author = {Srba, Ivan and Pecher, Branislav and Tomlein Matus and Moro, Robert and Stefancova, Elena and Simko, Jakub and Bielikova, Maria}, booktitle = {Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '22)}, numpages = {11}, title = {Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims}, year = {2022}, doi = {10.1145/3477495.3531726}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3477495.3531726}, } </code></pre> <p><br><strong>Dataset creation process</strong></p> <p>In order to create this dataset (and to continuously obtain new data), we used our research platform <a href="https://rome2019.github.io/papers/Srba_etal_ROME2019.pdf">Monant</a>. The Monant platform provides so called data providers to extract news articles/blogs from news/blog sites as well as fact-checking articles from fact-checking sites. General parsers (from RSS feeds, Wordpress sites, Google Fact Check Tool, etc.) as well as custom crawler and parsers were implemented (e.g., for fact checking site Snopes.com). All data is stored in the unified format in a central data storage.<br><strong> <br> <br>Ethical considerations</strong></p> <p>The dataset was collected and is published for research purposes only. We collected only publicly available content of news/blog articles. The dataset contains identities of authors of the articles if they were stated in the original source; we left this information, since the presence of an author's name can be a strong credibility indicator. However, we anonymised the identities of the authors of discussion posts included in the dataset. </p> <p>The main identified ethical issue related to the presented dataset lies in the risk of mislabelling of an article as supporting a false fact-checked claim and, to a lesser extent, in mislabelling an article as not containing a false claim or not supporting it when it actually does. To minimise these risks, we developed a labelling methodology and require an agreement of at least two independent annotators to assign a claim presence or article stance label to an article. It is also worth noting that we do not label an article as a whole as false or true. Nevertheless, we provide partial article-claim pair veracities based on the combination of claim presence and article stance labels.</p> <p>As to the veracity labels of the fact-checked claims and the credibility (reliability) labels of the articles' sources, we take these from the fact-checking sites and external listings such as Media Bias/Fact Check as they are and refer to their methodologies for more details on how they were established.</p> <p>Lastly, the dataset also contains automatically predicted labels of claim presence and article stance using our baselines described in the next section. These methods have their limitations and work with certain accuracy as reported in this paper. This should be taken into account when interpreting them. <br><strong> <br> <br>Reporting mistakes in the dataset</strong><br>The mean to report considerable mistakes in raw collected data or in manual annotations is by creating a new issue in the accompanying <a href="https://github.com/kinit-sk/medical-misinformation-dataset">Github repository</a>. Alternately, general enquiries or requests can be sent at info [at] kinit.sk.</p> <p><br><strong>Dataset structure</strong></p> <p><em><strong>Raw data</strong></em></p> <p>At first, the dataset contains so called raw data (i.e., data extracted by the Web monitoring module of Monant platform and stored in exactly the same form as they appear at the original websites). Raw data consist of articles from news sites and blogs (e.g. naturalnews.com), discussions attached to such articles, fact-checking articles from fact-checking portals (e.g. snopes.com). In addition, the dataset contains feedback (number of likes, shares, comments) provided by user on social network Facebook which is regularly extracted for all news/blogs articles.</p> <p>Raw data are contained in these CSV files:</p> <ul> <li>sources.csv</li> <li>articles.csv</li> <li>article_media.csv</li> <li>article_authors.csv</li> <li>discussion_posts.csv</li> <li>discussion_post_authors.csv</li> <li>fact_checking_articles.csv</li> <li>fact_checking_article_media.csv</li> <li>claims.csv</li> <li>feedback_facebook.csv</li> </ul> <p><em>Note: Personal information about discussion posts' authors (name, website, gravatar) are anonymised.</em></p> <p><br><em><strong>Annotations</strong></em></p> <p>Secondly, the dataset contains so called annotations. Entity annotations describe the individual raw data entities (e.g., article, source). Relation annotations describe relation between two of such entities.</p> <p>Each annotation is described by the following attributes:</p> <ol> <li>category of annotation (`annotation_category`). Possible values: label (annotation corresponds to ground truth, determined by human experts) and prediction (annotation was created by means of AI method).</li> <li>type of annotation (`annotation_type_id`). Example values: Source reliability (binary), Claim presence. The list of possible values can be obtained from enumeration in annotation_types.csv.</li> <li>method which created annotation (`method_id`). Example values: Expert-based source reliability evaluation, Fact-checking article to claim transformation method. The list of possible values can be obtained from enumeration methods.csv.</li> <li>its value (`value`). The value is stored in JSON format and its structure differs according to particular annotation type.</li> </ol> <p><br>At the same time, annotations are associated with a particular object identified by:</p> <ol> <li>entity type (parameter `entity_type` in case of entity annotations, or `source_entity_type` and `target_entity_type` in case of relation annotations). Possible values: sources, articles, fact-checking-articles.</li> <li>entity id (parameter `entity_id` in case of entity annotations, or `source_entity_id` and `target_entity_id` in case of relation annotations).</li> </ol> <p><br>The dataset provides specifically these entity annotations:</p> <ul> <li>Source reliability (binary). Determines validity of source (website) at a binary scale with two options: reliable source and unreliable source.</li> <li>Article veracity. Aggregated information about veracity from article-claim pairs.</li> </ul> <p>The dataset provides specifically these relation annotations:</p> <ul> <li>Fact-checking article to claim mapping. Determines mapping between fact-checking article and claim.</li> <li>Claim presence. Determines presence of claim in article.</li> <li>Claim stance. Determines stance of an article to a claim. </li> </ul> <p><br>Annotations are contained in these CSV files:</p> <ul> <li>entity_annotations.csv</li> <li>relation_annotations.csv</li> </ul> <p><em>Note: Identification of human annotators authors (email provided in the annotation app) is anonymised.</em></p> <p> </p> <p><em><strong>Enumerations</strong></em></p> <p>Finally, the dataset provides additional CSV files with enumerations:</p> <ul> <li>media_types.csv</li> <li>source_types.csv</li> <li>annotation_types.csv</li> <li>methods.csv</li> </ul>
Data for publication: Weighted manifold alignment using wave kernel signatures for aligning medical image datasets
<p>2D dynamic sagittal MR images, corresponding to volunteers E-H in the publication "Weighted manifold alignment using wave kernel signatures for aligning medical image datasets".</p>
Dataset from: "Articles Examining Medical Errors in Emergency Medicine Journals: A Comprehensive Visualization and Bibliometric Analysis"
Open the record for dataset details and reuse information.
Dataset related to Article "A SINGLE-GROUP STUDY ON THE EFFECT OF ONABOTULINUMTOXINA IN PATIENTS WITH CHRONIC MIGRAINE ASSOCIATED TO MEDICATION OVERUSE HEADACHE: PAIN CATASTROPHIZING PLAYS A ROLE"_ submitted
<p>THIS DATASET INCLUDES INFORMATION ON HEADACHE FREQUENCY, MEDICATION INTAKE, MIDAS SCORE, AVERAGE PAIN, HIT-6, ASC-12 AND PCS REFERRED TO PATIENTS RECEIVING ONABOTULINUMTOXINA AS PROPHYLAXIS FOR CM-MOH OVER 12 MONTHS.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.