Skip to main content
zenodoopen

DrugProt Complete PubMed Knowledge Graph

<p><strong>DrugProt Complete PubMed Knowledge Graph</strong></p><p><strong>Please cite if you use any DrugProt resource:</strong></p><p>Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of&nbsp;DrugProt task at BioCreative VII: data and&nbsp;methods for&nbsp;large-scale text mining and&nbsp;knowledge graph generation of&nbsp;heterogenous chemical–protein relations, <i>Database</i>, Volume 2023, 2023, baad080</p><blockquote><p><i>@article{miranda2023overview, &nbsp;title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, &nbsp;author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, &nbsp;journal={Database}, &nbsp;volume={2023}, &nbsp;pages={baad080}, &nbsp;year={2023}, &nbsp;publisher={Oxford University Press UK} }</i></p></blockquote><p><strong>Description</strong></p><p>This dataset contains a knowledge graph built from PubMed dump abstracts (December 2021). A NER system has been applied to each of the abstracts to extract mentions of type "CHEMICAL" and "GENE", as well as a RE system to detect existing relations between these mentions such as ACTIVATOR, INHIBITOR, AGONIST or PRODUCT_OF, among others (see article for a full list of relations considered).</p><p>Given the volume of the dataset, the repository is divided into 1114 folders. Each of these folders contains a chunk of PubMed abstracts, entities and relationships, divided into the following 3 files:</p><ul><li><i>abstracts.tsv</i>: Tabular file in which each line represents a pubmed document. The file has 3 columns:<ul><li>Pubmed_id: Numerical identifier of the document in PubMed</li><li>Title: Title of the document</li><li>Abstract: Abstract text.</li></ul></li><li>entities.tsv: List of the entities extracted from the abstracts. Each line represents an extracted entity, and has 5 columns:<ul><li>Pubmed_id: Numerical identifier of the document in PubMed</li><li>Mention_id: Numerical identifier of the mention in the document.</li><li>Entity_type: Type of mention. It can be CHEMICAL or GENE.</li><li>Span_ini: Index of the first character of the annotated span in the text</li><li>Span_end: Index of the first character after the annotated span.</li><li>Span: Text span of the annotation</li></ul></li><li>relations.tsv: File of existing relations between entities. Each line represents a relationship, and has the following fields:<ul><li>Pubmed_id:&nbsp;Numerical identifier of the document in PubMed</li><li>Relation_type:&nbsp;DrugProt relation type among entities/arguments.</li><li>Arg1: Mention of CHEMICAL</li><li>Arg2: Mention of GENE</li></ul></li></ul><p>&nbsp;</p><p><strong>Files:</strong></p><ul><li>drugprot-silver-standard-kg.zip : Folders with the files previously explained</li></ul><p>&nbsp;</p><p><strong>Related resources:</strong></p><ul><li><a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">Web</a></li><li><a href="https://doi.org/10.5281/zenodo.4955410">DrugProt corpus</a></li><li><a href="https://github.com/tonifuc3m/drugprot-evaluation-library">Evaluation library</a></li><li><a href="https://codalab.lisn.upsaclay.fr/competitions/8293">Online evaluation (CodaLab)</a></li><li><a href="https://doi.org/10.5281/zenodo.4957137">Relation annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957576">Gene and protein annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957518">Chemicals and drugs annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.5042178">FAQ</a></li><li><a href="https://doi.org/10.5281/zenodo.5119878">DrugProt Large Scale Additional SubTrack</a></li><li><a href="https://doi.org/10.5281/zenodo.5656991">DrugProt Large Scale document collection protocol</a></li><li><a href="https://doi.org/10.5281/zenodo.8246229">DrugProt Complete PubMed Knowledge Graph</a><br>&nbsp;</li></ul>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4

Topics