Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.9.0
Dataset results
9 results for “multi label classification”
Dataset and Code for Manuscript "Multi-angle pulse shape detection of scattered light in flow cytometry for label-free cell cycle classification"
<p>Dataset of measurements for cell cycle analysis with description:</p> <ul> <li>ReadMe file with explanations on the data set and analysis</li> <li>exemplary Matlab script file for analysis</li> <li>binary data files conatining the pulse shapes in all channels</li> <li>FCS data files containing common flow cytometry parameters in each channel</li> </ul> <p>Data on unsorted HEK cells, HEK cells sorted for cell cycle phases, and unsorted Jurkat cell are included.</p>
The fully executable procedure of the U-Net model combined with the Multi-textRG algorithm to achieve fine ice-water classification ---- another 332 scenes of data-fused SIC labels.
<p>This data source is related to the manuscript titled "Combining the U-Net model and a Multi-textRG algorithm for fine SAR ice-water classification", which will be submitted to the journal---The Cryosphere. </p> <ul> <li>The"ready-to-train-fused_01.zip" to "ready-to-train-fused_10.zip" includes 200 scenes of data-fused SIC labels accessible with doi: 10.5281/zenodo.10973107, https://zenodo.org/records/10973107. </li> <li>The "ready-to-train-fused_11.zip" to "ready-to-train-fused_21.zip" includes another 332 scenes of data-fused SIC labels. </li> </ul>
Multi-label Datasets used in "Adapting Transformers for Multi-Label Text Classification"
<p>The three Multi-Label datasets used in the article "Adapting Transformers for Multi-Label Text Classification".</p> <p>- AAPD Dataset (ArXiv Academic Paper Dataset) [Yang et al. 2018]<sup>1</sup></p> <p>- Reuters-21578 Dataset: https://archive.ics.uci.edu/ml/datasets/reuters-21578+text+categorization+collection</p> <p>- MFHAD (Multilabel French HAL Abstracts Dataset)</p> <p> </p> <p><sup>1</sup>Pengcheng Yang, Xu Sun, Wei Li, Shuming Ma, Wei Wu, and Houfeng Wang. 2018.<br> SGM: Sequence Generation Model for Multi-label Classification. In Proceedings<br> of the 27th International Conference on Computational Linguistics. Association for<br> Computational Linguistics, Santa Fe, New Mexico, USA, 3915–3926.</p>
Metabolic pathway inference using multi-label classification with rich pathway features
<p>We include samples of various data types used in the work "Metabolic pathway inference using multi-label classification with rich pathway features"</p> <p>More information about the software package and instructions are provided in <a href="https://github.com/hallamlab/mlLGPR">hallamlab/mlLGPR</a></p>
MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer
<p>The dataset is published with:<br> <br> <em>MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer. Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Punta Cana, Dominican Republic.</em><br> <br> <strong>Documents: </strong>MultiEURLEX comprises 65k EU in 23 official EU languages. Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU. Each EUROVOC label ID is associated with a Label descriptor, e.g., [60, `agri-foodstuffs'], [6006, `plant product'], [1115, `fruit']. The descriptors are also available in 23 languages. Chalkidis et al. (2019) published a monolingual (English) version of this dataset, called EURLEX57K, comprising 57k EU laws with the originally assigned gold labels.</p> <p><strong>Languages: </strong>MultiEURLEX covers 23 languages from 7 families. EU laws are published in all official EU languages, except for Irish for resource-related reasons (Read more: https://europa.eu/european-union/about-eu/eu-languages_en). This wide coverage makes the dataset a valuable testbed for cross-lingual transfer. All languages use the Latin script, except for Bulgarian (Cyrillic script) and Greek.</p> <p><strong>Multi-granular Labeling: </strong>EUROVOC<strong> </strong>has eight levels of concepts. Each document is assigned one or more concepts (labels). If a document is assigned a concept, the ancestors and descendants of that concept are typically not assigned to the same document. The documents were originally annotated with concepts from levels 3 to 8. We created three alternative sets of labels per document, by replacing each assigned concept by its ancestor from levels 1, 2, or 3, respectively. Thus, we provide four sets of gold labels per document, one for each of the first three levels of the hierarchy, plus the original sparse label assignment.</p> <p><strong>Supported Tasks: </strong>Similarly to EURLEX (Chalkidis et al., 2019), MultiEURLEX can be used for legal topic classification, a multi-label classification task where legal documents need to be assigned concepts (in our case, from EUROVOC) reflecting their topics. Unlike EURLEX57K, however, MultiEURLEX supports labels from three different granularities (EUROVOC levels). More importantly, apart from monolingual (one-to-one) experiments, it can be used to study cross-lingual transfer scenarios, including one-to-many (systems trained in one language and used in other languages with no training data), and many-to-one or many-to-many (systems jointly trained in multiple languages and used in one or more other languages).</p> <p><strong>Data Split and Concept Drift: </strong>MultiEURLEX is chronologically split in training (55k, 1958-2010), development (5k, 2010-2012), test (5k, 2012-2016) subsets, using the English documents. The test subset contains the same 5k documents in all 23 languages. The development subset also contains the same 5k documents in 23 languages, except Croatian. Croatia is the most recent EU member (2013); older laws are gradually translated. For the official languages of the seven oldest member countries, the same 55k training documents are available; for the other languages, only a subset of the 55k training documents is available. Compared to EURLEX57K (Chalkidis et al., 2019), MultiEURLEX is not only larger (8k more documents) and multilingual; it is also more challenging, as the chronological split leads to temporal real-world concept drift across the training, development, test subsets, i.e., differences in label distribution and phrasing, representing a realistic temporal generalization problem (Huang and Paul, 2019; Lazaridou et al., 2021). Recently, Søgaard et al. (2021) showed this setup is more realistic, as it does not overestimate real performance, contrary to random splits (Gorman and Bedrick, 2019).</p>
MedDialog-FR: a French Version of the MedDialog Corpus for Multi-label Classification and Response Generation related to Women's Intimate Health
<h1>MedDialog-FR: a French Version of the MedDialog Corpus for Multi-label Classification and Response Generation related to Women's Intimate Health</h1> <p> </p> <div><strong>Contributors</strong>: Xingyu Liu, Vincent Segonne, Aidan Mannion, Didier Schwab, Lorraine Goeuriot, François Portet</div> <p> </p> <div><strong>Total Number of Single-Turn Dialogues</strong>: 16,149 dialogues of women's intimate health, 7,120 dialogues of general medicine</div> <p> </p> <div>Given the lack of French dialogue corpora for data-driven dialogue systems and the paucity of available information related to women's intimate health, MedDialog-FR is an annotated corpus of question-and-answer sessions between a patient and a doctor concerning women's intimate health. The corpus is composed of about 20,000 sessions automatically translated from the English version of MedDialog-EN. The corpus test set is composed of 1,400 sessions that have been manually post-edited and annotated with 22 categories from the UMLS ontology.</div> <p> </p> <h2>Overview of the dataset</h2> <p> </p> <div>To construct the French MedDialog Dataset (<em>MedDialog-FR</em>), we initially extracted from <em>MedDialog-EN</em> and automatically translated a total of 16,149 dialogues related to women's intimate health and an additional 7,120 dialogues related to general medicine. <em>MedDialog-EN</em> is composed of textual single-turn dialogues: a medical question by a patient and a response by a physician. From the translated dialogues, we randomly selected 900 dialogues on women's intimate health and 500 dialogues concerning general medicine to be post-edited. Subsequently, we performed multi-label annotation on the 900 questions extracted from these same dialogues focused on women's intimate health. In total, 1,286 labels were annotated, with 1.43 labels per instance in average.</div> <p> </p> <div>The summary of the statistics of the dataset:</div> <table> <tbody> <tr> <td><strong>Task</strong></td> <td><strong>Women</strong></td> <td><strong>General</strong></td> </tr> <tr> <td>Machine translation (#dialogs)</td> <td>16,149</td> <td>7,120</td> </tr> <tr> <td>Post-editing (#dialogs)</td> <td>900</td> <td>500</td> </tr> <tr> <td>Multi-label annotation (#questions)</td> <td>900</td> <td>-</td> </tr> </tbody> </table> <p> </p> <h2>Structure of the dataset</h2> <p> </p> <div>The dataset contains the following elements separated in general medicine domain (<em>MedDialog-FR-general</em>) and women's intimate health domain (<em>MedDialog-FR-women</em>):</div> <div>```</div> <div>├── MedDialog-FR-general/</div> <div>├──── machine_translation/meddialog-fr-general_machine_translation.csv</div> <div>├──── post-editing/meddialog-fr-general_post-editing.csv</div> <p> </p> <div>├── MedDialog-FR-women/</div> <div>├──── machine_translation/meddialog-fr-women_machine_translation.csv</div> <div>├──── post-editing/meddialog-fr-women_post-editing.csv</div> <div>├──── multilabel_annotation/dataset_multilabel_meddialog_22labels.csv</div> <div>├──── response_generation/dataset_response_generation_meddialog.csv</div> <div> </div> <div>```</div> <div>All the .csv files contain a column named id, which indicates the original file of *MedDialog-EN* with the id in that file. For example, hm3_96_q or hm3_96_a refers to the session with the id of 96 within the healthcaremaginc3 file. The suffix of '_q' and '_a' indicates question and answer</div> <p> </p> <h3>Machine translation</h3> <div>The .csv file contains 3 columns: id, en and fr</div> <div>- en: original question and answer in English</div> <div>- fr: translated question and answer in French</div> <p> </p> <div>Example lines:</div> <div>hm4_1121_q \t J'ai 52 ans, mes dernières règles remontent au 6 décembre, je pensais que c'était peut-être le début de la ménopause. J'ai fait un test d'urine pour la grossesse, qui s'est révélé positif, puis j'ai fait un test quantitatif de hcg 45343 (je suis infirmière et je l'ai fait au laboratoire de l'hôpital où je travaille). J'ai des crampes et des saignements (bruns) depuis 2 à 3 mois.</div> <p> </p> <div>hm4_1121_a \t Bonjour, j'ai compris votre préoccupation. Comme vous avez mentionné que le taux de bêta HCG est plus élevé, je vous suggère de faire une échographie. Cela confirmera l'âge gestationnel et la viabilité de la grossesse. Si vous tenez à poursuivre la grossesse, veuillez discuter des risques encourus avec votre gynécologue traitant. Vous pouvez également opter pour une interruption de grossesse avec des médicaments en toute sécurité jusqu'à 9 semaines de grossesse sous surveillance médicale. J'espère que cette réponse vous aidera.</div> <p> </p> <h3>Post-editing</h3> <div>The .csv file contains 3 columns: id, machine_translation and post-edited</div> <p> </p> <div>Example line:</div> <div>hm4_3334_q \t bonjour docteur je suis atteinte de pcos, je me suis mariée en novembre 2011.nous essayons d'avoir une grossesse depuis deux mois ... et ma question est comment savoir la sévérité du pcos et quel est le meilleur moment pour concv \t bonjour docteur, je suis atteinte de SOPK, je me suis mariée en novembre 2011. Nous essayons de concevoir depuis deux mois ... et ma question est comment savoir la sévérité du SOPK et quel est le meilleur moment pour concevoir.</div> <p> </p> <h3>Multi-label annotation</h3> <div>The .csv file contains 5 columns: id, source_file, labels and split.</div> <div>- source_file: the source file where the text content for classification can be found with id</div> <div>- labels: UMLS IDs representing expert-validated labels for classification</div> <div>- split: train, dev or test</div> <p> </p> <div>Example line:</div> <div>hm1_33568_q \t '../post-editing/meddialog-fr-women_post-editing.csv' \t ['C0700589', 'C0227791'] \t train</div> <p><br><br></p> <div><strong>Partitioning</strong>:</div> <div>We split the <em>MedDialog-FR-women</em> multi-label dataset into a training set of 500 instances, a validation set of 100 instances and a test set of 300 instances. The ratio was chosen to balance the need for maximizing the amount of fine-tuning data available while also ensuring that the test set is large enough for the results to be statistically significant, given the scarcity of some categories. The split statistics are summarized in the following table. To maintain consistent label distribution, we leveraged the iterative stratification algorithm during the data splitting process.</div> <table> <tbody> <tr> <td><strong>Split</strong></td> <td><strong>#Questions</strong></td> </tr> <tr> <td>Train</td> <td>500</td> </tr> <tr> <td>Validation</td> <td>100</td> </tr> <tr> <td>Test</td> <td>300</td> </tr> </tbody> </table> <p> </p> <h3>Response generation</h3> <div>The .csv file contains 3 columns: id, split and source_file</div> <div>- split: train, dev or test</div> <div>- source_file: the source file where the text content for response generation can be found with id</div> <div><strong>Partitioning:</strong></div> <div>The validation and test data contain the same session ID as the multi-label validation and test, but they include the corresponding answers. for the training set, we use the same ones as multi-label dataset plus the machine translated sessions.</div> <div> </div> <table> <tbody> <tr> <td><strong>Split</strong></td> <td><strong>#Dialogues</strong></td> </tr> <tr> <td>Train</td> <td>15,749</td> </tr> <tr> <td>Validation</td> <td>100</td> </tr> <tr> <td>Test</td> <td>300</td> </tr> </tbody> </table> <div> </div> <p><br><br></p> <h2>Corpus data cleaning</h2> <div>By examining the MedDialog-EN corpus, we identified data that could potentially leak personal information such as the first and last name, email address, URL, etc. In order to safeguard privacy, we conducted a series of data cleaning procedures, especially anonymization:</div> <div>1. replace URLs with #URL# (regex pattern: `https?:\/\/(www\.)?[-a-zA-Z0-9@:%._\+~#=]{1,256}\.[a-zA-Z0-9()]{1,6}\b([-a-zA-Z0-9()@:%_\+.~#?&//=]*)</div> <div>`)</div> <div>2. replace emails with #EMAIL# (regex pattern: `^[\w-\.]+@([\w-]+\.)+[\w-]{2,4}`)</div> <div>3. replace phone numbers with #TEL# (regex pattern: `^[\+]?[(]?[0-9]{3}[)]?[-\s\.]?[0-9]{3}[-\s\.]?[0-9]{4,6}`)</div> <div>4. replace dates with #DATE# (regex patterns: `\d{1,2}\/\d{1,2}\/\d{2,4}</div> <div>`; `(Jan(?:uary)?|Feb(?:ruary)?|Mar(?:ch)?|Apr(?:il)?|May|Jun(?:e)?|Jul(?:y)?|Aug(?:ust)?|Sep(?:tember)?|Oct(?:ober)?|Nov(?:ember)?|Dec(?:ember)?)\s+(\d{1,2})\s+(\d{4})`)</div> <div>5. replace hospital or clinic names with #HOSPITAL# (text patterns: `clinic`; `hospital`)</div> <div>6. replace names in questions with #Person1#, and names in answers with #Person2#. If there is a name in an answer identical to the name in its question, replace it with #Person1#. (text patterns: `I am`; `I'm`;`Dr`; `Doctor`)</div> <div>7. replace the names of data source forums with coded letters (text patterns: forum names)</div> <p><br><br></p> <h2>Ethics Statement and Limitations</h2> <div>Access to actual medical data is very restricted and protected in France. We thus used an already publicly available corpus in English. But we did not simply translate it. We first made sure that no personal information could be found in the data. This is why we replaced all names that could have been kept in the original data. We also performed post-edition after automatic translation to adapt the phrasing and medical term to the French culture. All people recruited for annotation were treated fairly. This includes, but is not limited to, compensating them fairly and ensuring that they were voluntary participants. We do not foresee any direct social consequences or ethical issues.</div> <p> </p> <div>Authors of MedDialog were warned at our project and answered our questions.</div> <p> </p> <div>Since the original corpus is derived from dialogues in the U.S.A., there might be some cultural differences with French-speaking countries in the way people interact with doctors and which treatments and medical advises can be provided.</div> <p> </p> <div>Answers to questions should not be applied for self-treatment.</div> <div> </div>
Pseudo-Label Generation for Multi-Label Text Classification
With the advent and expansion of social networking, the amount of generated text data has seen a sharp increase. In order to handle such a huge volume of text data, new and improved text mining techniques are a necessity. One of the characteristics of text data that makes text mining difficult, is multi-labelity. In order to build a robust and effective text classification method which is an integral part of text mining research, we must consider this property more closely. This kind of property is not unique to text data as it can be found in non-text (e.g., numeric) data as well. However, in text data, it is most prevalent. This property also puts the text classification problem in the domain of multi-label classification (MLC), where each instance is associated with a subset of class-labels instead of a single class, as in conventional classification. In this paper, we explore how the generation of pseudo labels (i.e., combinations of existing class labels) can help us in performing better text classification and under what kind of circumstances. During the classification, the high and sparse dimensionality of text data has also been considered. Although, here we are proposing and evaluating a text classification technique, our main focus is on the handling of the multi-labelity of text data while utilizing the correlation among multiple labels existing in the data set. Our text classification technique is called pseudo-LSC (pseudo-Label Based Subspace Clustering). It is a subspace clustering algorithm that considers the high and sparse dimensionality as well as the correlation among different class labels during the classification process to provide better performance than existing approaches. Results on three real world multi-label data sets provide us insight into how the multi-labelity is handled in our classification process and shows the effectiveness of our approach.
MULTI-LABEL ASRS DATASET CLASSIFICATION USING SEMI-SUPERVISED SUBSPACE CLUSTERING
MULTI-LABEL ASRS DATASET CLASSIFICATION USING SEMI-SUPERVISED SUBSPACE CLUSTERING MOHAMMAD SALIM AHMED, LATIFUR KHAN, NIKUNJ OZA, AND MANDAVA RAJESWARI Abstract. There has been a lot of research targeting text classification. Many of them focus on a particular characteristic of text data - multi-labelity. This arises due to the fact that a document may be associated with multiple classes at the same time. The consequence of such a characteristic is the low performance of traditional binary or multi-class classification techniques on multi-label text data. In this paper, we propose a text classification technique that considers this characteristic and provides very good performance. Our multi-label text classification approach is an extension of our previously formulated [3] multi-class text classification approach called SISC (Semi-supervised Impurity based Subspace Clustering). We call this new classification model as SISC-ML(SISC Multi-Label). Empirical evaluation on real world multi-label NASA ASRS (Aviation Safety Reporting System) data set reveals that our approach outperforms state-of-theart text classification as well as subspace clustering algorithms.
ANALYZING AVIATION SAFETY REPORTS: FROM TOPIC MODELING TO SCALABLE MULTI-LABEL CLASSIFICATION
ANALYZING AVIATION SAFETY REPORTS: FROM TOPIC MODELING TO SCALABLE MULTI-LABEL CLASSIFICATION AMRUDIN AGOVIC*, HANHUAI SHAN*, AND ARINDAM BANERJEE* Abstract. The Aviation Safety Reporting System (ASRS) is used to collect voluntarily submitted aviation safety reports from pilots, controllers and others. As such it is particularly useful in researching aviation safety deficiencies. In this paper we address two challenges related to the analysis of ASRS data: (1) the unsupervised extraction of meaningful and interpretable topics from ASRS reports and (2) multi-label classification of ASRS data based on a set of predefined categories. For topic modeling we investigate the practical usefulness of Latent Dirichlet Allocation (LDA) when it comes to modeling ASRS reports in terms of interpretable topics. We also utilize LDA to generate a more compact representation of ASRS reports to be used in multi-label classification. For multi-label classification we propose a novel and highly scalable multi-label classification algorithm based on multi-variate regression. Empirical results indicate that our approach is superior to several baseline and state-of-the-art approaches.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.