Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “Relation corpus”
CrowdTruth Corpus for Open Domain Relation Extraction from Sentences
<p>This repository contains a ground truth corpus for open domain relation extraction from sentences, acquired with crowdsourcing and processed with <strong><a href="http://crowdtruth.org/">CrowdTruth</a></strong> metrics that capture ambiguity in annotations by measuring inter-annotator disagreement.</p> <p>The dataset contains annotations for 4,100 sentences sampled from Angeli et al. (1) and Riedel et al. (2), over 16 relations, with each sentence annotated by 15 workers. The sentences have been pre-processed with Distant Supervision (3) using the Freebase knowledge base, in order to identify the term pairs in each sentence that are likely to express a relation. The crowdsourced data was collected from <a href="http://figure-eight.com/">Figure Eight</a> and <a href="https://www.mturk.com/">Amazon Mechanical Turk</a>.</p> <p>This corpus has been discussed in the following papers:</p> <ul> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1809.00537">Crowdsourcing Semantic Label Propagation in Relation Classification</a></strong>. <a href="http://fever.ai/">FEVER</a> Workshop at <a href="http://emnlp2018.org/">EMNLP 2018</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1711.05186">False Positive and Cross-relation Signals in Distant Supervision Data</a></strong>. <a href="http://www.akbc.ws/">AKBC</a> Workshop at <a href="http://nips.cc/">NIPS 2017</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="http://crowdtruth.org/wp-content/uploads/2017/03/collint17-open-domain.pdf">Disagreement in Crowdsourcing and Active Learning for Better Distant Supervision Quality</a></strong>. <a href="http://collectiveintelligenceconference.org/">Collective Intelligence 2017</a>.</li> </ul> <p>Sentence-level data is available in file: <code>|--data/output/aggregated_sentences.csv</code></p> <p>Worker-level data is available in file: <code>|--data/output/aggregated_workers.csv</code></p> <p>Raw crowdsourcig data is available in folder: <code>|--data/input/</code></p> <p>Results of the relation classification model are available in folder: <code>|--data/model_results/</code></p> <p> </p> <p>References</p> <p>(1) Angeli, Gabor, et al. "Combining distant and partial supervision for relation extraction." Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 2014.</p> <p>(2) Riedel, Sebastian, et al. "Relation extraction with matrix factorization and universal schemas." Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2013.</p> <p>(3) Mintz, Mike, et al. "Distant supervision for relation extraction without labeled data." Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2-Volume 2. Association for Computational Linguistics, 2009.</p>
MiRoR11 - P2 - Annotated corpus for the relation between reported outcomes and their significance levels
<p>Corpus of relations between outcomes and significance levels</p> <p>This dataset contains annotations of the relations between reported outcomes and their significance levels.<br> Tab-separated format is used. The file contains the following comumns:<br> filename, sentence text, outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_sig_rel contains the dataset splits for 10-fold cross-validation.</p>
PretoxTM Corpus: a gold standard corpus of preclinical treatment-related findings annotated from toxicology reports
<p>The PretoxTM Corpus is a gold standard corpus of preclinical treatment-related findings annotated from toxicology reports.</p> <p>Example documents annotated by domain experts, also known as a gold standard corpus, are needed in order to develop, train, and validate text mining tools. To this aim we designed and performed an annotation activity for the development of the corpus of treatment-related findings: the PretoxTM corpus.</p> <p>A treatment-related finding expression enclose several named entities; the most relevant one is the abnormal effect detected; which depending on the study domain of the finding, can be given by a measurement, test, or examination named Study Test and an abnormal Manifestation result obtained for that study test; or by an abnormal Finding in study domains where there is no associated test or measurement (e.g., clinical, macroscopic and microscopic). Other related named entities that could be present to complete the treatment-related finding are; the Specimen of the abnormal observation, the Sex of the subject, the Group of subjects in which the observation was detected and the Dose level administration of the compound. Examples of sentences with treatment-related findings are: “The decrease in food consumption and body weight of the animals from the mid dose onwards is regarded as evidence of general toxicity.” and "At dose level 3, absolute and relative liver weights were increased in male rats.”.</p> <p>Contributions: The PretoxTM corpus was developed by BSC, with the contribution of IMIM and a team of experts from the eTRANSAFE EFPIA partners.</p> <p>License: Creative Commons Attribution-ShareAlike 4.0 International (cc by sa 4.0).</p> <p>The PretoxTM resources have been developed as part of the eTRANSAFE project.</p> <p>For more information about PretoxTM please visit:</p> <p>PretoxTM Corpus Gitlab: <a href="https://gitlab.bsc.es/inb/etransafe/preclinical-toxicological-corpus/">https://gitlab.bsc.es/inb/etransafe/preclinical-toxicological-corpus/</a></p> <p>PretoxTM central documentation: <a href="https://pretoxtm.gitlab.io/documentation/">https://pretoxtm.gitlab.io/documentation/</a></p>
A crowdsourced chemical-induced disease relation corpus
<p>A sentence-bound chemical-induced disease relation corpus was produced with crowdsourcing as part of the BioCreative V challenge.</p>
silverCID — a silver standard corpus for Chemical-induced Diseases relation extraction — Edit
<p>silverCID — a silver standard corpus for Chemical-induced Diseases relation extraction — Edit</p>
An annotated corpus of clinical trial publications supporting schema-based relational information extraction
<p>Repository of an annotated corpus of clinical trial abstracts supporting schema-based relational information extraction and the code for the inter-annotation agreement calculation and the baseline information extraction method.</p>
MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction
<p><strong>MEDDOPLACE</strong> stands for MEDical DOcument PLAce-related Content Extraction. It is a shared task and set of resources focused on the detection, normalization (entity linking/toponym resolution) and classification of different kinds of places, as well as related types of information such as clinical departments, nationalities or patient movements, in medical documents in Spanish.</p> <p>This repository includes the corpus' <strong>train and test sets</strong> in multiple formats, as well as the <strong>SNOMED gazetteer</strong>, <strong>cross-mapping</strong> between SNOMED and MeSH and the <strong>multilingual silver standard in 8 languages </strong>(Catalan, English, French, Italian, Dutch, Portuguese, Romanian and Swedish). For more information, please check the attached README file.</p> <p>MEDDOPLACE was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of IberLEF 2023. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a>.</p> <p> </p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-López, Eulàlia Farré-Maduell, Antonio Miranda-Escalada, Vicent Brivá-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoplace, title={MEDDOPLACE Shared Task overview: recognition, normalization and classification of locations and patient movement in clinical texts}, author={Lima-López, Salvador and Farré-Maduell, Eulàlia and Brivá-Iglesias, Vicent and Gasco-Sanchez, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {71}, year={2023}, issn = {1135-5948},<br>DOI = {10.26342/2023-71-23}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561/3961}, pages = {301--311} }</code></pre> <p><strong>Related Links:</strong></p> <p>- MEDDOPLACE website: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a></p> <p>- MEDDOPLACE overview paper: <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561">http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561</a></p> <p>- Annotation Guidelines (Spanish): <a href="https://doi.org/10.5281/zenodo.7775234">https://doi.org/10.5281/zenodo.7775234</a></p> <p>- Annotation Guidelines (English): <a href="https://doi.org/10.5281/zenodo.7928145">https://doi.org/10.5281/zenodo.7928145</a></p> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p>
Biomedical Entities and Relations on Spanish Clinical Case Corpus: BERSCCC
<p>This first version of a spanish corpus contains 200 clinical reports annotated with biomedical entities and semantic relations. <br> These reports belong to the Spanish Clinical Case Corpus (SPACCC) (https://doi.org/10.5281/zenodo.2560316)<br> and each of them has been annotated by three persons that work in the medicine, biomolecular or pharmaceutic area.</p> <p>The annotators had to identify the following thirteen types of entities in the spanish lenguage: Enfermedad/Síndrome, Gen, Parte del cuerpo/Órgano, Glúcido, Procedimiento de Diagnóstico, Proteína, Procedimiento Terapeútico, Síntoma/Signo, Sustancia Farmacológica, Lípido, Organismo, Químico Orgánico and Abreviatura/Sigla/Alias.<br> And the next eight semantic relations: Analiza, Altera, Causa, Diagnostica, Manifestación de, Produce, Trata and Refiere a. </p> <p>Finally there were identified 6,636 biomedical entities (37,081 mentions) and 4,864 semantic relations (7,622 mentions).</p> <p>These resources are freely distributed under a Creative Commons Attribution 4.0 International License.<br> The scripts used to create this corpus can be found at: https://github.com/drugs4covid/bio-corpora</p> <p>Author: Lucía Sánchez González, Ontology Engineering Group, Universidad Politécnica de Madrid.</p> <p>Supervisors: </p> <p>- Carlos Badenes Olmedo, Ontology Engineering Group, Universidad Politécnica de Madrid.</p> <p>- María Poveda Villalón, Ontology Engineering Group, Universidad Politécnica de Madrid.</p> <p>Project Member: Óscar Corcho García, Ontology Engineering Group, Universidad Politécnica de Madrid.</p> <p> Contact:</p> <p>Lucía Sánchez González at lu.sanchez@alumnos.upm.es or lusangonz99@gmail.com</p> <p>Acknowledgments to Project DRUGS4COVID++: Servicios de Inteligencia Artificial para la<br> creación de un grafo de conocimientos sobre fármacos usados en el control clínico de la<br> enfermedad, a partir de la explotación de grandes corpus de documentación científica sobre<br> SARS-COV-2 y COVID-19-AYUDAS FUNDACIÓN BBVA A EQUIPOS DE<br> INVESTIGACIÓN CIENTÍFICA SARS-CoV-2 y COVID-19.</p> <p> </p>
WikiDBs 10k - A Corpus Of Relational Databases From Wikidata
<p>WikiDBs-10k (<a href="https://wikidbs.github.io/" target="_blank" rel="noopener">https://wikidbs.github.io/</a>) is a corpus of relational databases built from Wikidata (<a href="https://www.wikidata.org/" target="_blank" rel="noopener">https://www.wikidata.org/</a>). This is the preliminary 10k version, the newer version of 100k databases (<a href="https://zenodo.org/records/11559814">https://zenodo.org/records/11559814</a>) includes more coherent databases and more diverse table and column names.</p> <p>The WikiDBs-10k corpus consists of 10,000 databases, for more details read our paper: <a href="https://ceur-ws.org/Vol-3462/TADA3.pdf" target="_blank" rel="noopener">https://ceur-ws.org/Vol-3462/TADA3.pdf</a> (TaDA@VLDB'23)</p> <p>Each database is saved in a sub-folder, the table files are provided as csv files and the database schema as a json file.</p> <p>We thank Till Döhmen and Madelon Hulsebos for generously providing the table statistics from their <a href="https://github.com/tdoehmen/gitschemas">GitSchemas dataset</a> and Jan-Micha Bodensohn for converting the dataset to SQLite files. This work has been supported by the BMBF and the state of Hesse as part of the NHR Program and the BMBF project KompAKI (grant number 02L19C150), as well as the HMWK cluster project 3AI. Finally, we want to thank hessian.AI, and DFKI Darmstadt for their support.</p>
Data for "RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature"
<div> <p><strong>RegulaTome corpus</strong>: this <a href="../api/records/10808330/files/RegulaTome-corpus.tar.gz/content" target="_blank" rel="noopener">file</a> contains the RegulaTome corpus in <a href="https://brat.nlplab.org/">BRAT</a> format. The directory <strong>"splits" </strong>has the corpus split based on the train/dev/test used for the training of the relation extraction system</p> <p><strong>RegulaTome annodoc</strong>: The annotation guidelines along with the annotation configuration files for BRAT are provided in <a href="../api/records/10808330/files/annodoc+config.tar.gz/content" target="_blank" rel="noopener">annodoc+config.tar.gz</a>. The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/regulatome-annodoc/ </a></p> <p>The tagger software can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a>. The command used to run tagger before large-scale execution of the RE system is:</p> <p><code>gzip -cd `ls -1 pmc/*.en.merged.filtered.tsv.gz` `ls -1r pubmed/*.tsv.gz` | cat dictionary/excluded_documents.txt - | tagger/tagcorpus --threads=16 --autodetect --types=dictionary/curated_types.tsv --entities=dictionary/all_entities.tsv --names=dictionary/all_names_textmining.tsv --groups=dictionary/all_groups.tsv --stopwords=dictionary/all_global.tsv --local-stopwords=dictionary/all_local.tsv --type-pairs=dictionary/all_type_pairs.tsv --out-matches=all_matches.tsv</code></p> <p><strong>Input documents </strong>for large-scale execution, which is done on entire <a href="https://a3s.fi/March-2024-PubMed/PubMed_20230314.tar.gz" target="_blank" rel="noopener">PubMed</a> (as of March 2024) and <a href="https://a3s.fi/Jan-2024-documents/PMC_Nov_23.tar.gz" target="_blank" rel="noopener">PMC Open Access</a> (as of November 2023) articles in BioC format. The files are converted to a <a href="https://a3s.fi/March-2024-PubMed/all_documents.tsv" target="_blank" rel="noopener">tab-delimited format </a>to be compatible with the RE system input (see below).</p> <p><strong>Input dictionary files</strong>: all the files necessary to execute the command above are available in <a href="../api/records/10808330/files/tagger_dictionary_files.tar.gz/content" target="_blank" rel="noopener">tagger_dictionary_files.tar.gz </a></p> <p><strong>Tagger output</strong>: we filter the results of the tagger run down to gene/protein hits, and documents with more than 1 hit (since we are doing relation extraction) before feeding it to our RE system. The filtered output is available in <a href="../api/records/10808330/files/tagger_matches_ggp_only_gt_1_hit.tsv.gz/content" target="_blank" rel="noopener">tagger_matches_ggp_only_gt_1_hit.tsv.gz</a></p> <p><strong>Relation extraction system input</strong>: <a href="../api/records/10808330/files/combined_input_for_re.tar.gz/content" target="_blank" rel="noopener">combined_input_for_re.tar.gz</a>: these are the directories with all the .ann and .txt files used as input for the large-scale execution of the relation extraction pipeline. The files are generated from the tagger tsv output (see above, <a href="../api/records/10808330/files/tagger_matches_ggp_only_gt_1_hit.tsv.gz/content" target="_blank" rel="noopener">tagger_matches_ggp_only_gt_1_hit.tsv.gz</a>) using the <a href="https://github.com/spyysalo/string-db-tools/blob/main/tagger2standoff.py">tagger2standoff.py</a> script from the <a href="https://github.com/spyysalo/string-db-tools/">string-db-tools</a> repository.</p> <p><strong>Relation extraction models</strong>. The Transformer-based model used for large-scale relation extraction and prediction on the test set is at <a href="../api/records/10808330/files/relation_extraction_multi-label-best_model.tar.gz/content" target="_blank" rel="noopener">relation_extraction_multi-label-best_model.tar.gz</a></p> <p>The pre-trained RoBERTa model on PubMed and PMC and MIMIC-III with a BPE Vocab learned from PubMed (RoBERTa-large-PM-M3-Voc), which is used by our system is available <a href="https://github.com/facebookresearch/bio-lm/blob/main/README.md">here</a>.</p> <p><strong>Relation extraction system output</strong>: the tab-delimited outputs of the relation extraction system are found at <a href="https://a3s.fi/regulatome-ls/large_scale_relation_extraction_results.tar.gz" target="_blank" rel="noopener">large_scale_relation_extraction_results.tar.gz </a><strong>!!!ATTENTION this file is approximately 1TB in size, so make sure you have enough space to download it on your machine!!!</strong></p> <p>The relation extraction system output files have 86 columns: PMID, Entity BRAT ID1, Entity BRAT ID2, and scores per class produced by the relation extraction model. Each file has a header to denote which score is in which column.</p> </div>
Hackathon - TF-TG relation annotation Gold Standard Corpus
<p>TF-TG relation annotation Gold Standard Corpus.</p> <p>The file contains 130 PMIDs annotated in the following way:</p> <ul> <li>PMID.txt: the original TXT file with the abstract.</li> <li>PMID.ann: the file with the BRAT annotations. Each annotation file contains terms (e.g. T1, T2, ...) and relations (e.g. R1, R2, ...), where each relation is from one term (e.g. T1) to another term (e.g. T2).</li> </ul>
WikiDBs - A Large-Scale Corpus Of Relational Databases From Wikidata
<p><a href="https://wikidbs.github.io/" target="_blank" rel="noopener">WikiDBs</a> is an open-source corpus of 100,000 relational databases. We aim to support research on tabular representation learning on multi-table data. The corpus is based on <a href="https://www.wikidata.org/" target="_blank" rel="noopener">Wikidata</a> and aims to follow certain characteristics of real-world databases.</p> <p>WikiDBs was published as a <a href="https://openreview.net/pdf?id=abXaOcvujs" target="_blank" rel="noopener">spotlight paper</a> at the Dataset & Benchmarks track at NeurIPS 2024.</p> <p>WikiDBs contains the database schemas, as well as table contents. The database tables are provided as CSV files, and each database schema as JSON. The 100,000 databases are available in five splits, containing 20k databases each. In total, around 165 GB of disk space are needed for the full corpus. We also provide a script to convert the databases into SQLite.</p>
Effect of lactation and location relative to the corpus luteum on the transcriptome of the bovine oviduct epithelium
GEO Series GSE124110. Bos taurus. 56 samples. Type: Expression profiling by high throughput sequencing.
MedDialog-FR: a French Version of the MedDialog Corpus for Multi-label Classification and Response Generation related to Women's Intimate Health
<h1>MedDialog-FR: a French Version of the MedDialog Corpus for Multi-label Classification and Response Generation related to Women's Intimate Health</h1> <p> </p> <div><strong>Contributors</strong>: Xingyu Liu, Vincent Segonne, Aidan Mannion, Didier Schwab, Lorraine Goeuriot, François Portet</div> <p> </p> <div><strong>Total Number of Single-Turn Dialogues</strong>: 16,149 dialogues of women's intimate health, 7,120 dialogues of general medicine</div> <p> </p> <div>Given the lack of French dialogue corpora for data-driven dialogue systems and the paucity of available information related to women's intimate health, MedDialog-FR is an annotated corpus of question-and-answer sessions between a patient and a doctor concerning women's intimate health. The corpus is composed of about 20,000 sessions automatically translated from the English version of MedDialog-EN. The corpus test set is composed of 1,400 sessions that have been manually post-edited and annotated with 22 categories from the UMLS ontology.</div> <p> </p> <h2>Overview of the dataset</h2> <p> </p> <div>To construct the French MedDialog Dataset (<em>MedDialog-FR</em>), we initially extracted from <em>MedDialog-EN</em> and automatically translated a total of 16,149 dialogues related to women's intimate health and an additional 7,120 dialogues related to general medicine. <em>MedDialog-EN</em> is composed of textual single-turn dialogues: a medical question by a patient and a response by a physician. From the translated dialogues, we randomly selected 900 dialogues on women's intimate health and 500 dialogues concerning general medicine to be post-edited. Subsequently, we performed multi-label annotation on the 900 questions extracted from these same dialogues focused on women's intimate health. In total, 1,286 labels were annotated, with 1.43 labels per instance in average.</div> <p> </p> <div>The summary of the statistics of the dataset:</div> <table> <tbody> <tr> <td><strong>Task</strong></td> <td><strong>Women</strong></td> <td><strong>General</strong></td> </tr> <tr> <td>Machine translation (#dialogs)</td> <td>16,149</td> <td>7,120</td> </tr> <tr> <td>Post-editing (#dialogs)</td> <td>900</td> <td>500</td> </tr> <tr> <td>Multi-label annotation (#questions)</td> <td>900</td> <td>-</td> </tr> </tbody> </table> <p> </p> <h2>Structure of the dataset</h2> <p> </p> <div>The dataset contains the following elements separated in general medicine domain (<em>MedDialog-FR-general</em>) and women's intimate health domain (<em>MedDialog-FR-women</em>):</div> <div>```</div> <div>├── MedDialog-FR-general/</div> <div>├──── machine_translation/meddialog-fr-general_machine_translation.csv</div> <div>├──── post-editing/meddialog-fr-general_post-editing.csv</div> <p> </p> <div>├── MedDialog-FR-women/</div> <div>├──── machine_translation/meddialog-fr-women_machine_translation.csv</div> <div>├──── post-editing/meddialog-fr-women_post-editing.csv</div> <div>├──── multilabel_annotation/dataset_multilabel_meddialog_22labels.csv</div> <div>├──── response_generation/dataset_response_generation_meddialog.csv</div> <div> </div> <div>```</div> <div>All the .csv files contain a column named id, which indicates the original file of *MedDialog-EN* with the id in that file. For example, hm3_96_q or hm3_96_a refers to the session with the id of 96 within the healthcaremaginc3 file. The suffix of '_q' and '_a' indicates question and answer</div> <p> </p> <h3>Machine translation</h3> <div>The .csv file contains 3 columns: id, en and fr</div> <div>- en: original question and answer in English</div> <div>- fr: translated question and answer in French</div> <p> </p> <div>Example lines:</div> <div>hm4_1121_q \t J'ai 52 ans, mes dernières règles remontent au 6 décembre, je pensais que c'était peut-être le début de la ménopause. J'ai fait un test d'urine pour la grossesse, qui s'est révélé positif, puis j'ai fait un test quantitatif de hcg 45343 (je suis infirmière et je l'ai fait au laboratoire de l'hôpital où je travaille). J'ai des crampes et des saignements (bruns) depuis 2 à 3 mois.</div> <p> </p> <div>hm4_1121_a \t Bonjour, j'ai compris votre préoccupation. Comme vous avez mentionné que le taux de bêta HCG est plus élevé, je vous suggère de faire une échographie. Cela confirmera l'âge gestationnel et la viabilité de la grossesse. Si vous tenez à poursuivre la grossesse, veuillez discuter des risques encourus avec votre gynécologue traitant. Vous pouvez également opter pour une interruption de grossesse avec des médicaments en toute sécurité jusqu'à 9 semaines de grossesse sous surveillance médicale. J'espère que cette réponse vous aidera.</div> <p> </p> <h3>Post-editing</h3> <div>The .csv file contains 3 columns: id, machine_translation and post-edited</div> <p> </p> <div>Example line:</div> <div>hm4_3334_q \t bonjour docteur je suis atteinte de pcos, je me suis mariée en novembre 2011.nous essayons d'avoir une grossesse depuis deux mois ... et ma question est comment savoir la sévérité du pcos et quel est le meilleur moment pour concv \t bonjour docteur, je suis atteinte de SOPK, je me suis mariée en novembre 2011. Nous essayons de concevoir depuis deux mois ... et ma question est comment savoir la sévérité du SOPK et quel est le meilleur moment pour concevoir.</div> <p> </p> <h3>Multi-label annotation</h3> <div>The .csv file contains 5 columns: id, source_file, labels and split.</div> <div>- source_file: the source file where the text content for classification can be found with id</div> <div>- labels: UMLS IDs representing expert-validated labels for classification</div> <div>- split: train, dev or test</div> <p> </p> <div>Example line:</div> <div>hm1_33568_q \t '../post-editing/meddialog-fr-women_post-editing.csv' \t ['C0700589', 'C0227791'] \t train</div> <p><br><br></p> <div><strong>Partitioning</strong>:</div> <div>We split the <em>MedDialog-FR-women</em> multi-label dataset into a training set of 500 instances, a validation set of 100 instances and a test set of 300 instances. The ratio was chosen to balance the need for maximizing the amount of fine-tuning data available while also ensuring that the test set is large enough for the results to be statistically significant, given the scarcity of some categories. The split statistics are summarized in the following table. To maintain consistent label distribution, we leveraged the iterative stratification algorithm during the data splitting process.</div> <table> <tbody> <tr> <td><strong>Split</strong></td> <td><strong>#Questions</strong></td> </tr> <tr> <td>Train</td> <td>500</td> </tr> <tr> <td>Validation</td> <td>100</td> </tr> <tr> <td>Test</td> <td>300</td> </tr> </tbody> </table> <p> </p> <h3>Response generation</h3> <div>The .csv file contains 3 columns: id, split and source_file</div> <div>- split: train, dev or test</div> <div>- source_file: the source file where the text content for response generation can be found with id</div> <div><strong>Partitioning:</strong></div> <div>The validation and test data contain the same session ID as the multi-label validation and test, but they include the corresponding answers. for the training set, we use the same ones as multi-label dataset plus the machine translated sessions.</div> <div> </div> <table> <tbody> <tr> <td><strong>Split</strong></td> <td><strong>#Dialogues</strong></td> </tr> <tr> <td>Train</td> <td>15,749</td> </tr> <tr> <td>Validation</td> <td>100</td> </tr> <tr> <td>Test</td> <td>300</td> </tr> </tbody> </table> <div> </div> <p><br><br></p> <h2>Corpus data cleaning</h2> <div>By examining the MedDialog-EN corpus, we identified data that could potentially leak personal information such as the first and last name, email address, URL, etc. In order to safeguard privacy, we conducted a series of data cleaning procedures, especially anonymization:</div> <div>1. replace URLs with #URL# (regex pattern: `https?:\/\/(www\.)?[-a-zA-Z0-9@:%._\+~#=]{1,256}\.[a-zA-Z0-9()]{1,6}\b([-a-zA-Z0-9()@:%_\+.~#?&//=]*)</div> <div>`)</div> <div>2. replace emails with #EMAIL# (regex pattern: `^[\w-\.]+@([\w-]+\.)+[\w-]{2,4}`)</div> <div>3. replace phone numbers with #TEL# (regex pattern: `^[\+]?[(]?[0-9]{3}[)]?[-\s\.]?[0-9]{3}[-\s\.]?[0-9]{4,6}`)</div> <div>4. replace dates with #DATE# (regex patterns: `\d{1,2}\/\d{1,2}\/\d{2,4}</div> <div>`; `(Jan(?:uary)?|Feb(?:ruary)?|Mar(?:ch)?|Apr(?:il)?|May|Jun(?:e)?|Jul(?:y)?|Aug(?:ust)?|Sep(?:tember)?|Oct(?:ober)?|Nov(?:ember)?|Dec(?:ember)?)\s+(\d{1,2})\s+(\d{4})`)</div> <div>5. replace hospital or clinic names with #HOSPITAL# (text patterns: `clinic`; `hospital`)</div> <div>6. replace names in questions with #Person1#, and names in answers with #Person2#. If there is a name in an answer identical to the name in its question, replace it with #Person1#. (text patterns: `I am`; `I'm`;`Dr`; `Doctor`)</div> <div>7. replace the names of data source forums with coded letters (text patterns: forum names)</div> <p><br><br></p> <h2>Ethics Statement and Limitations</h2> <div>Access to actual medical data is very restricted and protected in France. We thus used an already publicly available corpus in English. But we did not simply translate it. We first made sure that no personal information could be found in the data. This is why we replaced all names that could have been kept in the original data. We also performed post-edition after automatic translation to adapt the phrasing and medical term to the French culture. All people recruited for annotation were treated fairly. This includes, but is not limited to, compensating them fairly and ensuring that they were voluntary participants. We do not foresee any direct social consequences or ethical issues.</div> <p> </p> <div>Authors of MedDialog were warned at our project and answered our questions.</div> <p> </p> <div>Since the original corpus is derived from dialogues in the U.S.A., there might be some cultural differences with French-speaking countries in the way people interact with doctors and which treatments and medical advises can be provided.</div> <p> </p> <div>Answers to questions should not be applied for self-treatment.</div> <div> </div>
FRACAS: FRench Annotated Corpus of Attribution relations in newS
<p>A human-annotated corpus for French quotation extraction<strong> </strong>containing 1676 newswire texts with 10 965 annotated attribution relations (quotes attributed to its speaker).</p> <p><strong>Data</strong>: 1676 newswire texts in French from Reuters annotated with 10 965 attribution relations</p> <p><strong>Date: </strong>April 1995 to April 1996</p> <p><strong>Data structure</strong>:</p> <p>{</p> <p>"text": <em>text of the newswire,</em></p> <p>"entities": <em>a list of each entity in the following format</em> ["id": <em>unique_id</em>, "text": <em>text of entity</em>, "label": <em>entity label</em>, <em>"</em>gender": <em>gender (if labelled</em>), "char_span": <em>a list of character index span</em>]</p> <p>"relations"<em>: </em><em>a list of each relation in the following format</em> [<em>id of relation</em>, <em>label of relation</em>, <em>id of first entity</em>, <em>id of second entity</em></p> <p>}</p> <p><strong>Labels:</strong></p> <ul> <li>Entities: <ul> <li><strong>Quotation</strong> (Direct, Indirect or Mixed)</li> <li><strong>Speaker</strong> (Agent, Organization, Group of People, Source Pronoun)</li> <li><strong>Cue</strong></li> </ul> </li> <li>Attributes: <ul> <li>Speaker Gender (Male, Female, Mixed, Unknown, Other)</li> </ul> </li> <li>Relations: <ul> <li>Speaker <strong>Quoted in </strong>Quotation</li> <li>Cue <strong>Indicates </strong>Quotation</li> <li>Source Pronoun <strong>Refers</strong> to Speaker</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.