Datasets for Out-of-KB Mention Discovery with Entity Linking
<p>The repository contains datasets for out-of-KB mention discovery from texts, documented in the work, <em>Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking</em>, on arXiv: <a href="https://arxiv.org/abs/2302.07189">https://arxiv.org/abs/2302.07189</a> (CIKM 2023).</p> <p>Each data setting (as a sub-folder) contains train, valid, and test files and also 100 random sample files for each data split for debugging.</p> <p>Data folder names with “syn_full” at the end are synonym augmented data (each synonym as an entity) for the setting.</p> <p>Ontology .jsonl files have two versions for each, "syn_attr" setting treats synonyms are attributes, "syn_full" setting treats synonyms as entities.</p> <p> </p> <p>Data scripts are available at <a href="https://github.com/KRR-Oxford/BLINKout#data-scripts">https://github.com/KRR-Oxford/BLINKout#data-scripts</a></p> <p> </p> <p>Acknowledgement of the data sources below:</p> <p>ShARe/CLEF 2013 dataset is from <a href="https://physionet.org/content/shareclefehealth2013/1.0/">https://physionet.org/content/shareclefehealth2013/1.0/</a></p> <p>MedMention dataset is from <a href="https://github.com/chanzuckerberg/MedMentions">https://github.com/chanzuckerberg/MedMentions</a></p> <p>UMLS (versions 2012AB, 2014AB, 2017AA) is from <a href="https://www.nlm.nih.gov/research/umls/index.html">https://www.nlm.nih.gov/research/umls/index.html</a></p> <p>SNOMED CT (corresponding versions) is from <a href="https://www.nlm.nih.gov/healthit/snomedct/index.html">https://www.nlm.nih.gov/healthit/snomedct/index.html</a></p> <p>NILK dataset is from <a href="https://zenodo.org/record/6607514">https://zenodo.org/record/6607514</a></p> <p>WikiData 2017 dump is from <a href="https://archive.org/download/enwiki-20170220/enwiki-20170220-pages-articles.xml.bz2">https://archive.org/download/enwiki-20170220/enwiki-20170220-pages-articles.xml.bz2</a></p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0