Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

22

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

22 results for “Multi-label”

Learn how ShareScore rates datasets ↗
zenodo44/100

Multi-label Tweet Dataset for Textual Propaganda Detection related to anti-CAA protest in India (2019-2021)

<p>This&nbsp; is a collection of English Tweets&nbsp; about the&nbsp;Citizenship (Amendment) Bill protests that occurred in India in&nbsp;2019-2021. The dataset contains&nbsp; tweet&nbsp;instances multiple labels for identified propaganda techniques. The data set consists of tweet ids, hashtags&nbsp;used, and corresponding propaganda techniques. Labels have&nbsp;been automatically generated using Weak Supervision.&nbsp;</p> <p>As of 2023, there are very limited textual propaganda detection dataset for Tweets. This&nbsp;dataset is released to facilitate future research&nbsp;as propaganda has become omnipresent in modern social media.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

ZeroCostDL4Mic / DeepBacs - Multi-label U-Net training dataset (Bacillus subtilis) and pretrained model

<p>Training and test images of live <em>B. subtilis </em>cells expressing FtsZ-GFP for the task of segmentation.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>The example shows the fluorescence widefield image of live <em>B. subtilis </em>cells expressing FtsZ-GFP, the manually annotated instance segmentation mask and the corresponding 2-label semantic segmentation mask used for model training.</p> <p>&nbsp;</p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired fluorescence and segmented mask images</p> <p><strong>Microscopy data type</strong>: 2D widefield images (fluorescence)&nbsp;</p> <p><strong>Microscope</strong>: Custom-built 100x inverted microscope bearing a 100x TIRF objective (Nikon CFI Apochromat TIRF 100XC Oil); images were captured on a Prime BSI sCMOS camera (Teledyne Photometrics)</p> <p><strong>Cell type</strong>: <em>B. subtilis</em> strain SH130 grown under agarose pads</p> <p><strong>File format</strong>: .tiff (8-bit)</p> <p><strong>Image size</strong>: 1024 x 1024 px&sup2; (Pixel size: 65 nm)</p> <p><strong>Image preprocessing</strong>: Images were denoised using PureDenoise and resulting 32-bit images were converted into 8-bit images after normalizing to 1% and 99.98% percentiles. Images were manually annotated using the Labkit Fiji plugin and mask images with labeled cytosol and cell boundaries were created using a custom Fiji macro (see our <a href="https://github.com/HenriquesLab/DeepBacs/tree/main/ImageJ-macros">github repository</a>).</p> <p>&nbsp;</p> <p><strong>Multi-label U-Net model</strong>:</p> <p>The U-Net (2D) multilabel model was generated using the ZeroCostDL4Mic platform (Chamier &amp; Laine et al., 2021). It was trained from scratch for 200 epochs on 733 paired image patches (image dimensions: (1024 x 1024 px&sup2;), patch size: (256 x 256 px&sup2;)) with a batch size of 8 and a categorical_crossentrop loss function, using the U-Net (2D) multilabel ZeroCostDL4Mic notebook (v 1) (Chamier &amp; Laine et al., 2021). Key python packages used include tensorflow (v 0.1.12), Keras (v 2.3.1), numpy (v 1.19.5), cuda (v 11.1.105). The training was accelerated using a Tesla P100GPU.</p> <p>&nbsp;</p> <p><strong>Author(s)</strong>: Mia Conduit<sup>1,2</sup>, S&eacute;amus Holden<sup>1,3</sup></p> <p><strong>Contact email</strong>: <a href="mailto:Seamus.Holden@newcastle.ac.uk">Seamus.Holden@newcastle.ac.uk</a></p> <p>&nbsp;</p> <p><strong>Affiliation</strong>:</p> <p>1) Centre for Bacterial Cell Biology, Biosciences Institute, Newcastle University, NE2 4AX UK</p> <p>2) ORCID: 0000-0002-7169-907X</p> <p>&nbsp;</p> <p>&nbsp;<strong>Associated publications</strong>: Whitley <em>et al</em>., 2021, Nature Communications, https://doi.org/10.15252/embj.201696235</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Dataset: A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern

<p>NORMAN Suspect List Exchange was used for the generation of the dataset. Datasets with clear label (LC or GC) were used. More specifically, we used S3 NORMANCT15, which contains a list of compounds that were detected in surface water from the Danube River in a pan-European collaborative trial employing both GC-HRMS and LC-HRMS. Moreover, the GC and LC target list were used by the following two institutes: National and Kapodistrian University of Athens (NKUA) and Helmholtz Centre for Environmental Research (UFZ). S21 UATHTARGETS is the LC target list of NKUA, S65 UATHTARGETSGC is the GC target list of NKUA and S53 UFZWANATARG contains the LC and GC target list of UFZ. Finally, two GC target lists (S51 WRIGCHRMS and S70 EISUSGCEIMS) were used. These lists contain GC substance lists and were provided by two Slovak institutes, the Water Research Institute (WRI) and Environmental Institute. The aforementioned compound lists were merged together to form a labelled dataset. The SMILES were used to calculate 1446 molecular descriptors. 1446 descriptors were produced by PaDEL-descriptor, logP was produced by JRgui and boiling point by USEPA ECOSAR.</p> <p>The dataset is used in the publication:</p> <p>&quot;A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern&quot; authored by</p> <p>Nikiforos Alygizakis, Vasileios Konstantakos, Grigoris Bouziotopoulos , Evangelos Kormentzas, Jaroslav Slobodnik and Nikolaos S. Thomaidis</p> <p>Github repository:&nbsp;https://github.com/nalygizakis/LCvsGC</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Topic Modeling for Multi-label Research Articles

<p>The abstract and title for a set of research articles, and their assigned topics.&nbsp;The research article abstracts and titles are sourced from the following 6 topics:</p> <ol> <li> <p>Computer Science</p> </li> <li> <p>Physics</p> </li> <li> <p>Mathematics</p> </li> <li> <p>Statistics</p> </li> <li> <p>Quantitative Biology</p> </li> <li> <p>Quantitative Finance</p> </li> </ol> <p>Note that a research article can possibly have more than 1 topic. For more details check out the paper of :</p> <p>&quot;Evaluation of SVM Transformations for Multi-Label Research Article Classification&quot;</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

MORFITT : A multi-label corpus of French scientific articles in the biomedical domain

<p>This article presents MORFITT, the first multi-label corpus in French annotated in specialties in the medical field. MORFITT is composed of 3~624 abstracts of scientific articles from PubMed, annotated in 12 specialties for a total of 5,116 annotations. We detail the corpus, the experiments and the preliminary results obtained using a classifier based on the pre-trained language model CamemBERT. These preliminary results demonstrate the difficulty of the task, with a weighted average F1-score of 61.78%.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

EPIC Aff: A multi-label segmentation visual affordances dataset.

<p>This is the EPIC-Aff&nbsp;dataset introduced on the ICCV 2023 paper &quot;Multi-label affordance mapping from egocentric vision&quot;.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

MiL-FISH: Multi-labelled oligonucleotides for fluorescence in situ hybridisation improve visualization of bacterial cells

<p>Comparison of mono-, MiL- and CARD-FISH on LR-White embedded <em>Olavius algarvensis</em> cross sections A: Ethanol preserved specimen, cu = cuticle, sym = symbionts, sep = septum, epi = epidermis, mu = muscle, vbv = ventral blood vessel &amp; nerve chord, chl = chloragogen tissue, grid square = example of region shown in images B, C, D and E.</p> <p>Probes for Gammaproteobacteria (Gam42a; green) and a subgroup of sulfate-reducing Deltaproteobacteria (DSS658; red) on B) mono-FISH on ethanol preserved specimen, 19 hour hybridisation. C) mono-FISH on ethanol preserved specimen, 3 hour hybridisation and D) CARD-FISH on Carnoy&rsquo;s / PFA fixed specimen, 3-hour hybridisation. E) MiL-FISH on Carnoy&rsquo;s / PFA fixed specimen, 3 hour hybridisation. Scale bar = 5 &micro;m.</p>

opencc-by-4.0Nov 2015View details →
zenodo36/100

MiL-FISH: Multi-labelled oligonucleotides for fluorescence in situ hybridisation improve visualization of bacterial cells

<p>Sections of LR-White embedded <em>Olavius algarvensis </em>eggs. A) DAPI stained overview of egg after first cleavage, grid square = region in B, cle = cleavage. B) Gamma- (ii) and Delta- (i) proteobacteria hybridised with 4x labelled Gam42a &amp; 4x labelled DSS658 probe respectively. Autofluorescence of egg yolk is overcome and bacteria are seen to closely associated with the developing embryo. y = egg yolk C) DAPI stained overview of juvenile worm in egg, grid square = region in D. D) Gamma 1 symbiont phylotype (iii) hybridised with 16S rRNA specific probe labelled with 2x FITC and 2x Cy3 to produce yellow in the overlay. Symbiont cells are incorporated between cuticle and epidermis and in close proximity to egg yolk. y = egg yolk, c = cuticle. Scale: A &amp; C = 50 &micro;m, B &amp; D = 5 &micro;m.</p>

opencc-zeroNov 2015View details →
zenodo36/100

MiL-FISH: Multi-labelled oligonucleotides for fluorescence in situ hybridisation improve visualization of bacterial cells

<p>Epifluorescence images of MiL-FISH labelled microorganisms: A) <em>Beggiatoa sp. </em>hybridised with Gam42a, B) <em>Desulfococcus biacutus </em>with DSS658, C) <em>Roseobacter sp. </em>with Ros537, D) <em>Sulfurimonas denitrificans </em>with EPSY914, E) <em>Rhodopirellula sp. </em>SH1 T with PLA46, F) <em>Gramella forsetii </em>with CF319a, G) <em>Metallosphaera sedula </em>with Arch915, H) Composite image of all seven microbial partners in an artificial mix. Letters correlate to images of individual organisms A-G. Scale bar: A &amp; H = 10 &micro;m, B-G = 5 &micro;m.</p> <p>From:</p> <p>Schimak MP, Kleiner M, Wetzel S, Liebeke M, Dubilier N, Fuchs BM. 2015. MiL-FISH: Multi-labelled oligonucleotides for fluorescence in situ hybridisation improve visualization of bacterial cells. Applied and Environmental Microbiology. Accepted manuscript posted online. doi:10.1128/AEM.02776-15</p>

opencc-by-4.0Nov 2015View details →
zenodo36/100

Supplementary datasets for the paper of "Multi-resBind: a residual network-based multi-label classifier for in vivo RNA binding prediction and preference visualization"

<p>There are two eCLIP datasets (cell lines of K562 and HepG2).&nbsp;The eCLIP datasets were then divided into five categories for each cell line: low, medium 1, medium 2, high 1 and high 2 with peaks of &gt;1,000 but &lt;2,000, &gt;2,000 but &lt;4,000, &gt;4,000 but &lt;7,000, &gt;7,000 but &lt;10,000 and &gt;10,000, respectively.&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Multi-label Datasets used in "Adapting Transformers for Multi-Label Text Classification"

<p>The three Multi-Label&nbsp;datasets used in the article &quot;Adapting Transformers for Multi-Label Text Classification&quot;.</p> <p>- AAPD Dataset&nbsp;&nbsp;(ArXiv Academic Paper Dataset) [Yang et al. 2018]<sup>1</sup></p> <p>- Reuters-21578 Dataset:&nbsp;https://archive.ics.uci.edu/ml/datasets/reuters-21578+text+categorization+collection</p> <p>- MFHAD (Multilabel French HAL Abstracts Dataset)</p> <p>&nbsp;</p> <p><sup>1</sup>Pengcheng Yang, Xu Sun, Wei Li, Shuming Ma, Wei Wu, and Houfeng Wang. 2018.<br> SGM: Sequence Generation Model for Multi-label Classification. In Proceedings<br> of the 27th International Conference on Computational Linguistics. Association for<br> Computational Linguistics, Santa Fe, New Mexico, USA, 3915&ndash;3926.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Super Resolved and Multi-Label Segmented PEMFC

<p>Super Resolved and Multi-Label Segmented PEMFC with Configuration Files for; Large-scale Physically Accurate Modelling of Real Proton Exchange Membrane Fuel Cell with Deep Learning, Nature Communications.</p> <table> <tbody> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/ffov_crop_origsize.tiff">ffov_crop_origsize.tiff </a>- LOW RESOLUTION REGISTERED TRAINING BLOCK<br> md5:0b8293b1bf6ded48df1cdd6a36b22e61</td> <td>8.8 MB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/inputFile.db">inputFile.db </a>- LBPM INPUT FILE<br> md5:542d244a9552d6b4b1b7ecfa8fd1c153</td> <td>1 kB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/LRTest.tif">LRTest.tif </a>- LOW RESOLUTION FULL FOV DOMAIN<br> md5:85f2ccaff7b96188062efd6ab241630b</td> <td>1.3 GB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/PEFC_hres_0p7um.tiff">PEFC_hres_0p7um.tiff </a>- HIGH RESOLUTION REGISTERED TRAINING BLOCK<br> md5:ec918d150a6bb6ee5af576cb0bc55c03</td> <td>1.7 GB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/SRSegFinal_Processed_oneWeave_Opened.tif">SRSegFinal_Processed_oneWeave_Opened.tif </a>- SUPER RESOLVED AND MULTI-LABEL SEGMENTED&nbsp;FULL FOV DOMAIN<br> md5:20c7aea8c3b832b400811958b8ea243a</td> <td>2.2 GB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/SRSegWithGC.tif">SRSegWithGC.tif </a>- SUPER RESOLVED AND MULTI-LABEL SEGMENTED&nbsp;FULL FOV DOMAIN WITH GAS CHANNELS<br> md5:8a6cdd51d7b1b55f1ce9e0beaaa0aaed</td> <td>2.5 GB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td><a href="https://zenodo.org/api/files/19f99ad1-b88b-45aa-8731-a88d2753fe3c/SRVis_1100_4000_8000.raw">SRVis_1100_4000_8000.raw </a>- SUPER RESOLVED&nbsp;FULL FOV DOMAIN<br> md5:3ba78b624e60845c09f1d34820eb68e5</td> <td>35.2 GB</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> </tbody> </table> <table> <tbody> <tr> <th>&nbsp;</th> <th>&nbsp;</th> <th>&nbsp;</th> <th>&nbsp;</th> </tr> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> </tr> </tbody> </table>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Multi-Label Datasets with Missing Values

<p>Consisting of six multi-label datasets&nbsp;from the <a href="https://archive.ics.uci.edu">UCI Machine Learning repository.</a></p> <p>Each dataset contains missing values which have been artificially added at the following rates: 5, 10, 15, 20, 25, and<br> 30%.&nbsp;The &ldquo;amputation&rdquo; was performed using the &ldquo;Missing Completely at Random&rdquo; mechanism.</p> <p>File names are represented as follows:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; amp_<em>DB</em>_<em>MR</em>.arff</p> <p>where:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; <em>DB&nbsp;</em>= original dataset;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<em>MR </em>= missing rate.</p> <p>For more details, please read:</p> <p>IEEE Access article (in review process)</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Metabolic pathway inference using multi-label classification with rich pathway features

<p>We include samples of various data types used in the work &quot;Metabolic pathway inference using multi-label classification with rich pathway features&quot;</p> <p>More information about the software package and instructions are provided in&nbsp;<a href="https://github.com/hallamlab/mlLGPR">hallamlab/mlLGPR</a></p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Multi-label Pathway Prediction based on Active Dataset Subsampling

<p>We include samples of various data types used in the work &quot;Multi-label Pathway Prediction based on Active Dataset Subsampling&quot; (under-review)</p> <p>More information about the software package and instructions are provided in&nbsp;<a href="https://github.com/hallamlab/leADS">hallamlab/leADS</a></p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

DeepMRG: a multi-label deep learning classifier for predicting bacterial metal resistance genes

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo32/100

MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer

<p>The dataset is published with:<br> <br> <em>MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer. Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. Proceedings of the&nbsp;2021&nbsp;Conference on&nbsp;Empirical Methods in Natural Language Processing. 2021. Punta Cana, Dominican Republic.</em><br> <br> <strong>Documents: </strong>MultiEURLEX&nbsp;comprises 65k EU&nbsp;in 23 official EU languages. Each EU&nbsp;law has been annotated with EUROVOC&nbsp;concepts (labels) by the Publication Office of EU. Each EUROVOC&nbsp;label ID is associated with a Label descriptor, e.g., [60, `agri-foodstuffs&#39;], &nbsp;[6006, `plant product&#39;], [1115, `fruit&#39;]. The descriptors are also available in 23 languages. Chalkidis et al. (2019) published a&nbsp;monolingual&nbsp;(English) version of this dataset, called EURLEX57K, comprising 57k EU&nbsp;laws with the originally assigned gold labels.</p> <p><strong>Languages: </strong>MultiEURLEX&nbsp;covers 23 languages from 7 families. EU&nbsp;laws are published in all official EU&nbsp;languages, except for Irish for resource-related reasons&nbsp;(Read more:&nbsp;https://europa.eu/european-union/about-eu/eu-languages_en).&nbsp;This wide coverage makes the dataset a valuable testbed for cross-lingual transfer. All languages use the Latin script, except for Bulgarian (Cyrillic script) and Greek.</p> <p><strong>Multi-granular Labeling: </strong>EUROVOC<strong>&nbsp;</strong>has eight levels of concepts. Each document is assigned one or more concepts (labels). If a document is assigned a concept, the ancestors and descendants of that concept are typically not assigned to the same document. The documents were originally annotated with concepts from levels 3 to 8. &nbsp;We created three alternative sets of labels per document, by replacing each assigned concept by its ancestor from levels 1, 2, or 3, respectively. Thus, we provide four sets of gold labels per document, one for each of the first three levels of the hierarchy, plus the original sparse label assignment.</p> <p><strong>Supported Tasks:&nbsp;</strong>Similarly to EURLEX&nbsp;(Chalkidis et al., 2019), MultiEURLEX&nbsp;can be used for legal topic classification, a multi-label classification task where legal documents need to be assigned concepts (in our case, from EUROVOC) reflecting their topics. Unlike EURLEX57K, however, MultiEURLEX&nbsp;supports labels from three different granularities (EUROVOC&nbsp;levels). More importantly, apart from monolingual (one-to-one) experiments, it can be used to study cross-lingual transfer scenarios, including one-to-many&nbsp;(systems trained in one language and used in other languages with no training data), and many-to-one&nbsp;or many-to-many&nbsp;(systems jointly trained in multiple languages and used in one or more other languages).</p> <p><strong>Data Split and Concept Drift:&nbsp;</strong>MultiEURLEX&nbsp;is chronologically&nbsp;split in training (55k, 1958-2010), development (5k, 2010-2012), test (5k, 2012-2016) subsets, using the English documents. The test subset contains the same 5k documents in all 23 languages. The development subset also contains the same 5k documents in 23 languages, except Croatian. Croatia is the most recent EU&nbsp;member (2013); older laws are gradually translated.&nbsp;For the official languages of the seven oldest member countries, the same 55k training documents are available; for the other languages, only a subset of the 55k training documents is available.&nbsp;Compared to EURLEX57K&nbsp;(Chalkidis et al., 2019), MultiEURLEX&nbsp;is not only larger (8k more documents) and multilingual; it is also more challenging, as the chronological split leads to temporal real-world concept drift&nbsp;across the training, development, test subsets, i.e., differences in label distribution and phrasing, representing a realistic temporal generalization&nbsp;problem (Huang and Paul, 2019; Lazaridou et al., 2021). Recently, S&oslash;gaard et al. (2021) showed this setup is more realistic, as it does not overestimate real performance, contrary to random splits (Gorman and Bedrick, 2019).</p>

opencc-by-4.0Aug 2021View details →
zenodo24/100

MedDialog-FR: a French Version of the MedDialog Corpus for Multi-label Classification and Response Generation related to Women's Intimate Health

<h1>MedDialog-FR: a French Version of the MedDialog Corpus for Multi-label Classification and Response Generation related to Women's Intimate Health</h1> <p>&nbsp;</p> <div><strong>Contributors</strong>: Xingyu Liu, Vincent Segonne, Aidan Mannion, Didier Schwab, Lorraine Goeuriot, Fran&ccedil;ois Portet</div> <p>&nbsp;</p> <div><strong>Total Number of Single-Turn Dialogues</strong>: 16,149 dialogues of women's intimate health, 7,120 dialogues of general medicine</div> <p>&nbsp;</p> <div>Given the lack of French dialogue corpora for data-driven dialogue systems and the paucity of available information related to women's intimate health, MedDialog-FR is an annotated corpus of question-and-answer sessions between a patient and a doctor concerning women's intimate health. The corpus is composed of about 20,000 sessions automatically translated from the English version of MedDialog-EN. The corpus test set is composed of 1,400 sessions that have been manually post-edited and annotated with 22 categories from the UMLS ontology.</div> <p>&nbsp;</p> <h2>Overview of the dataset</h2> <p>&nbsp;</p> <div>To construct the French MedDialog Dataset (<em>MedDialog-FR</em>), we initially extracted from <em>MedDialog-EN</em> and automatically translated a total of 16,149 dialogues related to women's intimate health and an additional 7,120 dialogues related to general medicine. <em>MedDialog-EN</em> is composed of textual single-turn dialogues: a medical question by a patient and a response by a physician. From the translated dialogues, we randomly selected 900 dialogues on women's intimate health and 500 dialogues concerning general medicine to be post-edited. Subsequently, we performed multi-label annotation on the 900 questions extracted from these same dialogues focused on women's intimate health. In total, 1,286 labels were annotated, with 1.43 labels per instance in average.</div> <p>&nbsp;</p> <div>The summary of the statistics of the dataset:</div> <table> <tbody> <tr> <td><strong>Task</strong></td> <td><strong>Women</strong></td> <td><strong>General</strong></td> </tr> <tr> <td>Machine translation (#dialogs)</td> <td>16,149</td> <td>7,120</td> </tr> <tr> <td>Post-editing (#dialogs)</td> <td>900</td> <td>500</td> </tr> <tr> <td>Multi-label annotation (#questions)</td> <td>900</td> <td>-</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2>Structure of the dataset</h2> <p>&nbsp;</p> <div>The dataset contains the following elements separated in general medicine domain (<em>MedDialog-FR-general</em>) and women's intimate health domain (<em>MedDialog-FR-women</em>):</div> <div>```</div> <div>├── MedDialog-FR-general/</div> <div>├──── machine_translation/meddialog-fr-general_machine_translation.csv</div> <div>├──── post-editing/meddialog-fr-general_post-editing.csv</div> <p>&nbsp;</p> <div>├── MedDialog-FR-women/</div> <div>├──── machine_translation/meddialog-fr-women_machine_translation.csv</div> <div>├──── post-editing/meddialog-fr-women_post-editing.csv</div> <div>├──── multilabel_annotation/dataset_multilabel_meddialog_22labels.csv</div> <div>├──── response_generation/dataset_response_generation_meddialog.csv</div> <div>&nbsp;</div> <div>```</div> <div>All the .csv files contain a column named id, which indicates the original file of *MedDialog-EN* with the id in that file. For example, hm3_96_q or hm3_96_a refers to the session with the id of 96 within the healthcaremaginc3 file. The suffix of '_q' and '_a' indicates question and answer</div> <p>&nbsp;</p> <h3>Machine translation</h3> <div>The .csv file contains 3 columns: id, en and fr</div> <div>- en: original question and answer in English</div> <div>- fr: translated question and answer in French</div> <p>&nbsp;</p> <div>Example lines:</div> <div>hm4_1121_q \t J'ai 52 ans, mes derni&egrave;res r&egrave;gles remontent au 6 d&eacute;cembre, je pensais que c'&eacute;tait peut-&ecirc;tre le d&eacute;but de la m&eacute;nopause. J'ai fait un test d'urine pour la grossesse, qui s'est r&eacute;v&eacute;l&eacute; positif, puis j'ai fait un test quantitatif de hcg 45343 (je suis infirmi&egrave;re et je l'ai fait au laboratoire de l'h&ocirc;pital o&ugrave; je travaille). J'ai des crampes et des saignements (bruns) depuis 2 &agrave; 3 mois.</div> <p>&nbsp;</p> <div>hm4_1121_a \t Bonjour, j'ai compris votre pr&eacute;occupation. Comme vous avez mentionn&eacute; que le taux de b&ecirc;ta HCG est plus &eacute;lev&eacute;, je vous sugg&egrave;re de faire une &eacute;chographie. Cela confirmera l'&acirc;ge gestationnel et la viabilit&eacute; de la grossesse. Si vous tenez &agrave; poursuivre la grossesse, veuillez discuter des risques encourus avec votre gyn&eacute;cologue traitant. Vous pouvez &eacute;galement opter pour une interruption de grossesse avec des m&eacute;dicaments en toute s&eacute;curit&eacute; jusqu'&agrave; 9 semaines de grossesse sous surveillance m&eacute;dicale. J'esp&egrave;re que cette r&eacute;ponse vous aidera.</div> <p>&nbsp;</p> <h3>Post-editing</h3> <div>The .csv file contains 3 columns: id, machine_translation and post-edited</div> <p>&nbsp;</p> <div>Example line:</div> <div>hm4_3334_q \t bonjour docteur je suis atteinte de pcos, je me suis mari&eacute;e en novembre 2011.nous essayons d'avoir une grossesse depuis deux mois ... et ma question est comment savoir la s&eacute;v&eacute;rit&eacute; du pcos et quel est le meilleur moment pour concv \t bonjour docteur, je suis atteinte de SOPK, je me suis mari&eacute;e en novembre 2011. Nous essayons de concevoir depuis deux mois ... et ma question est comment savoir la s&eacute;v&eacute;rit&eacute; du SOPK et quel est le meilleur moment pour concevoir.</div> <p>&nbsp;</p> <h3>Multi-label annotation</h3> <div>The .csv file contains 5 columns: id, source_file, labels and split.</div> <div>- source_file: the source file where the text content for classification can be found with id</div> <div>- labels: UMLS IDs representing expert-validated labels for classification</div> <div>- split: train, dev or test</div> <p>&nbsp;</p> <div>Example line:</div> <div>hm1_33568_q \t '../post-editing/meddialog-fr-women_post-editing.csv' \t ['C0700589', 'C0227791'] \t train</div> <p><br><br></p> <div><strong>Partitioning</strong>:</div> <div>We split the&nbsp;<em>MedDialog-FR-women</em> multi-label dataset into a training set of 500 instances, a validation set of 100 instances and a test set of 300 instances. The ratio was chosen to balance the need for maximizing the amount of fine-tuning data available while also ensuring that the test set is large enough for the results to be statistically significant, given the scarcity of some categories. The split statistics are summarized in the following table. To maintain consistent label distribution, we leveraged the iterative stratification algorithm during the data splitting process.</div> <table> <tbody> <tr> <td><strong>Split</strong></td> <td><strong>#Questions</strong></td> </tr> <tr> <td>Train</td> <td>500</td> </tr> <tr> <td>Validation</td> <td>100</td> </tr> <tr> <td>Test</td> <td>300</td> </tr> </tbody> </table> <p>&nbsp;</p> <h3>Response generation</h3> <div>The .csv file contains 3 columns: id, split and source_file</div> <div>- split: train, dev or test</div> <div>- source_file: the source file where the text content for response generation can be found with id</div> <div><strong>Partitioning:</strong></div> <div>The validation and test data contain the same session ID as the multi-label validation and test, but they include the corresponding answers. for the training set, we use the same ones as multi-label dataset plus the machine translated sessions.</div> <div>&nbsp;</div> <table> <tbody> <tr> <td><strong>Split</strong></td> <td><strong>#Dialogues</strong></td> </tr> <tr> <td>Train</td> <td>15,749</td> </tr> <tr> <td>Validation</td> <td>100</td> </tr> <tr> <td>Test</td> <td>300</td> </tr> </tbody> </table> <div>&nbsp;</div> <p><br><br></p> <h2>Corpus data cleaning</h2> <div>By examining the MedDialog-EN corpus, we identified data that could potentially leak personal information such as the first and last name, email address, URL, etc. In order to safeguard privacy, we conducted a series of data cleaning procedures, especially anonymization:</div> <div>1. replace URLs with #URL# (regex pattern: `https?:\/\/(www\.)?[-a-zA-Z0-9@:%._\+~#=]{1,256}\.[a-zA-Z0-9()]{1,6}\b([-a-zA-Z0-9()@:%_\+.~#?&amp;//=]*)</div> <div>`)</div> <div>2. replace emails with #EMAIL# (regex pattern: `^[\w-\.]+@([\w-]+\.)+[\w-]{2,4}`)</div> <div>3. replace phone numbers with #TEL# (regex pattern: `^[\+]?[(]?[0-9]{3}[)]?[-\s\.]?[0-9]{3}[-\s\.]?[0-9]{4,6}`)</div> <div>4. replace dates with #DATE# (regex patterns: `\d{1,2}\/\d{1,2}\/\d{2,4}</div> <div>`; `(Jan(?:uary)?|Feb(?:ruary)?|Mar(?:ch)?|Apr(?:il)?|May|Jun(?:e)?|Jul(?:y)?|Aug(?:ust)?|Sep(?:tember)?|Oct(?:ober)?|Nov(?:ember)?|Dec(?:ember)?)\s+(\d{1,2})\s+(\d{4})`)</div> <div>5. replace hospital or clinic names with #HOSPITAL# (text patterns: `clinic`; `hospital`)</div> <div>6. replace names in questions with #Person1#, and names in answers with #Person2#. If there is a name in an answer identical to the name in its question, replace it with #Person1#. (text patterns: `I am`; `I'm`;`Dr`; `Doctor`)</div> <div>7. replace the names of data source forums with coded letters (text patterns: forum names)</div> <p><br><br></p> <h2>Ethics Statement and Limitations</h2> <div>Access to actual medical data is very restricted and protected in France. We thus used an already publicly available corpus in English. But we did not simply translate it. We first made sure that no personal information could be found in the data. This is why we replaced all names that could have been kept in the original data. We also performed post-edition after automatic translation to adapt the phrasing and medical term to the French culture. All people recruited for annotation were treated fairly. This includes, but is not limited to, compensating them fairly and ensuring that they were voluntary participants. We do not foresee any direct social consequences or ethical issues.</div> <p>&nbsp;</p> <div>Authors of MedDialog were warned at our project and answered our questions.</div> <p>&nbsp;</p> <div>Since the original corpus is derived from dialogues in the U.S.A., there might be some cultural differences with French-speaking countries in the way people interact with doctors and which treatments and medical advises can be provided.</div> <p>&nbsp;</p> <div>Answers to questions should not be applied for self-treatment.</div> <div>&nbsp;</div>

restrictedcc-by-4.0Mar 2024View details →
nasa20/100

Pseudo-Label Generation for Multi-Label Text Classification

With the advent and expansion of social networking, the amount of generated text data has seen a sharp increase. In order to handle such a huge volume of text data, new and improved text mining techniques are a necessity. One of the characteristics of text data that makes text mining difficult, is multi-labelity. In order to build a robust and effective text classification method which is an integral part of text mining research, we must consider this property more closely. This kind of property is not unique to text data as it can be found in non-text (e.g., numeric) data as well. However, in text data, it is most prevalent. This property also puts the text classification problem in the domain of multi-label classification (MLC), where each instance is associated with a subset of class-labels instead of a single class, as in conventional classification. In this paper, we explore how the generation of pseudo labels (i.e., combinations of existing class labels) can help us in performing better text classification and under what kind of circumstances. During the classification, the high and sparse dimensionality of text data has also been considered. Although, here we are proposing and evaluating a text classification technique, our main focus is on the handling of the multi-labelity of text data while utilizing the correlation among multiple labels existing in the data set. Our text classification technique is called pseudo-LSC (pseudo-Label Based Subspace Clustering). It is a subspace clustering algorithm that considers the high and sparse dimensionality as well as the correlation among different class labels during the classification process to provide better performance than existing approaches. Results on three real world multi-label data sets provide us insight into how the multi-labelity is handled in our classification process and shows the effectiveness of our approach.

restrictednotspecifiedMar 2025View details →
nasa20/100

MULTI-LABEL ASRS DATASET CLASSIFICATION USING SEMI-SUPERVISED SUBSPACE CLUSTERING

MULTI-LABEL ASRS DATASET CLASSIFICATION USING SEMI-SUPERVISED SUBSPACE CLUSTERING MOHAMMAD SALIM AHMED, LATIFUR KHAN, NIKUNJ OZA, AND MANDAVA RAJESWARI Abstract. There has been a lot of research targeting text classification. Many of them focus on a particular characteristic of text data - multi-labelity. This arises due to the fact that a document may be associated with multiple classes at the same time. The consequence of such a characteristic is the low performance of traditional binary or multi-class classification techniques on multi-label text data. In this paper, we propose a text classification technique that considers this characteristic and provides very good performance. Our multi-label text classification approach is an extension of our previously formulated [3] multi-class text classification approach called SISC (Semi-supervised Impurity based Subspace Clustering). We call this new classification model as SISC-ML(SISC Multi-Label). Empirical evaluation on real world multi-label NASA ASRS (Aviation Safety Reporting System) data set reveals that our approach outperforms state-of-theart text classification as well as subspace clustering algorithms.

restrictednotspecifiedMar 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record