Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1 result for “biocreative”

Learn how ShareScore rates datasets ↗
zenodo40/100

DrugProt corpus: Biocreative VII Track 1 - Text mining drug and chemical-protein interactions

<p>Gold Standard annotations of the DrugProt corpus (training and development sets). Also, test and background sets.</p><p>&nbsp;</p><p><strong>Please cite if you use any DrugProt resource:</strong></p><p>Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of&nbsp;DrugProt task at BioCreative VII: data and&nbsp;methods for&nbsp;large-scale text mining and&nbsp;knowledge graph generation of&nbsp;heterogenous chemical–protein relations, <i>Database</i>, Volume 2023, 2023, baad080</p><blockquote><p><i>@article{miranda2023overview, &nbsp;title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, &nbsp;author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, &nbsp;journal={Database}, &nbsp;volume={2023}, &nbsp;pages={baad080}, &nbsp;year={2023}, &nbsp;publisher={Oxford University Press UK} }</i></p></blockquote><p>Miranda, Antonio, et al. "Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations."&nbsp;<i>Proceedings of the seventh BioCreative challenge evaluation workshop</i>. 2021.</p><blockquote><p><i>@inproceedings{miranda2021overview, &nbsp;title={Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations}, &nbsp;author={Miranda, Antonio and Mehryary, Farrokh and Luoma, Jouni and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, &nbsp;booktitle={Proceedings of the seventh BioCreative challenge evaluation workshop}, &nbsp;year={2021} }</i></p></blockquote><p>&nbsp;</p><p><strong>Introduction</strong></p><p>The aim of the DrugProt track (similar to the previous CHEMPROT task of BioCreative VI) is to promote the development and evaluation of systems that are able to automatically detect in relations between chemical compounds/drug and genes/proteins. We have therefore generated a manually annotated corpus, the&nbsp;<i>DrugProt corpus</i>, where domain experts have exhaustively labeled:(a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (<i>DrugProt relation classes</i>). There is also an increasing interested in the integration of chemical and biomedical data understood as curation of relationships between biological and chemical entities from text and storing such information in form of structured annotation databases. Such databases are of key relevance not only for biological but also for pharmacological and clinical research. A range of different types chemical-protein/gene interactions are of key relevance for biology, including metabolic relations (e.g. substrates, products) inhibition, binding or induction associations.</p><p>The DrugProt track aims to address these needs and to promote the development of systems able to extract chemical-protein interactions that might be of relevance for precision medicine as well as for drug discovery and basic biomedical research.</p><p>The DrugProt track in BioCreative VII (BC VII) will explore recognition of chemical-protein entity relations from abstracts.</p><p>Teams participating in this track are provided with:</p><ul><li>PubMed abstracts</li><li>Manually annotated chemical compound mentions</li><li>Manually annotated gene/protein mentions</li><li>Manually annotated chemical compound-protein relations</li></ul><p>&nbsp;</p><p><strong>Zip structure:</strong></p><ul><li>Training set folder with<ul><li>drugprot_training_abstracts.tsv:&nbsp;PubMed records</li><li>drugprot_training_entities.tsv:&nbsp;manually labeled mention annotations of chemical compounds and genes/proteins</li><li>drugprot_training_relations.tsv: chemical-­protein relation annotations</li></ul></li><li>Development set folder with<ul><li>drugprot_development_abstracts.tsv</li><li>drugprot_development_entities.tsv</li><li>drugprot_development_relations.tsv</li></ul></li><li>Test+background set folder with<ul><li>test_background_abstracts.tsv</li><li>test_background_entities.tsv</li></ul></li></ul><p>&nbsp;</p><p><strong>Data format&nbsp;description</strong></p><p>The <strong>input text files</strong> for the DrugProt track are plain-text, UTF8-encoded PubMed records in a tab-separated format with the following three columns:</p><ol><li>Article identifier (PMID, PubMed identifier)</li><li>Title of the article</li><li>Abstract of the article</li></ol><p>&nbsp;</p><p>DrugProt <strong>entity mention annotation files</strong>&nbsp;contain manually labeled mention annotations of chemical compounds and genes/proteins. Such files consist of tab-separated fields containing the following six columns:</p><ol><li>Article identifier (PMID)</li><li>Term number (for this record)</li><li>Type of entity mention (CHEMICAL, GENE-Y, GENE-N)</li><li>Start character offset of the entity mention</li><li>End character offset of the entity mention</li><li>Text string of the entity mention</li></ol><p>Each line contains one entity, and <i>each entity is uniquely identified by its PMID and the Term Number</i>. Besides, each annotation contains an annotation type, the start-offset -the index of the first character of the annotated span in the text-, the end-offset -the index of the first character after the annotated span- and the text spanned&nbsp;by the annotation.</p><p>Example DrugProt <i>training</i> entity mention annotations:</p><p>11808879 T1 GENE-Y 1860 1866 KIR6.2 11808879 T2 GENE-N 1993 2016 glutamate dehydrogenase 11808879 T3 GENE-Y 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</p><p>&nbsp;</p><p>Example DrugProt <i>development</i> entity mention annotations (no distinction between GENE-Y and GENE-N):</p><p>11808879 T1 GENE 1860 1866 KIR6.2 11808879 T2 GENE 1993 2016 glutamate dehydrogenase 11808879 T3 GENE 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</p><p><br>DrugProt <strong>relation annotations</strong> are distributed as a file that contains the detailed chemical-protein relation annotations prepared for the DrugProt track. There are no relation annotations for the test+background set (the goal of the task is to predict them). It consists of tab-separated columns containing:</p><ol><li>Article identifier (PMID)</li><li>DrugProt relation</li><li>Interactor argument 1 (<i>of type CHEMICAL</i>)</li><li>Interactor argument 2 (<i>of type GENE</i>)</li></ol><p>Each line contains one relation, and <i>each relation is identified by the PMID, the relation type and the two related entities</i>. In the below example, to find the entities involved in the first relation, you must find the entities with Term Identifier T1 and T52 <i>within the PMID&nbsp;12488248.</i></p><p>Example DrugProt relation&nbsp;annotations:</p><p>12488248 INHIBITOR Arg1:T1 Arg2:T52 12488248 INHIBITOR Arg1:T2 Arg2:T52 23220562 ACTIVATOR Arg1:T12 Arg2:T42 23220562 ACTIVATOR Arg1:T12 Arg2:T43 23220562 INDIRECT-DOWNREGULATOR Arg1:T1 Arg2:T14</p><p>&nbsp;</p><p>Please, cite:</p><p>@inproceedings{krallinger2017overview,&nbsp;title={Overview of the BioCreative VI chemical-protein interaction Track},&nbsp;author={Krallinger, Martin and Rabal, Obdulia and Akhondi, Saber A and P{\'e}rez, Mart{\i}n P{\'e}rez and Santamar{\'\i}a, Jes{\'u}s and Rodr{\'\i}guez, Gael P{\'e}rez and others},&nbsp;booktitle={Proceedings of the sixth BioCreative challenge evaluation workshop},&nbsp;volume={1},&nbsp;pages={141--146},&nbsp;year={2017}}</p><p>&nbsp;</p><p><strong>Summary statistics:</strong></p><p>Training set Development set Documents 3500 750 Tokens 1001168 199620 Annotated Entities 89529 18858 Annotated Relations 17288 3765</p><p>&nbsp;</p><p>Annotated Entities:</p><p>Training Entities Development Entities CHEMICAL 46274 9853 GENE-Y [Normalizable] 28421 - GENE-N [Non-Normalizable] 14834 - Gene Total (N+Y) 43255 9005 Total 89529 18858</p><p>&nbsp;</p><p>Annotated Relations:</p><p>Training Relations Development Relations INDIRECT-DOWNREGULATOR 1330 332 INDIRECT-UPREGULATOR 1379 302 DIRECT-REGULATOR 2250 458 ACTIVATOR 1429 246 INHIBITOR 5392 1152 AGONIST 659 131 AGONIST-ACTIVATOR 29 10 AGONIST-INHIBITOR 13 2 ANTAGONIST 972 218 PRODUCT-OF 921 158 SUBSTRATE 2003 495 SUBSTRATE_PRODUCT-OF 25 3 PART-OF 886 258 Total 17288 3765</p><p>&nbsp;</p><p>For further information, please visit&nbsp;<a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/</a> or email us at krallinger.martin@gmail.com and antoniomiresc@gmail.com</p><p>&nbsp;</p><p><strong>Related resources:</strong></p><ul><li><a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">Web</a></li><li><a href="https://github.com/tonifuc3m/drugprot-evaluation-library">Evaluation library</a></li><li><a href="https://codalab.lisn.upsaclay.fr/competitions/8293">Online evaluation (CodaLab)</a></li><li><a href="https://doi.org/10.5281/zenodo.4957137">Relation annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957576">Gene and protein annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957518">Chemicals and drugs annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.7252201">DrugProt Silver Standard Knowledge Graph</a></li><li><a href="https://doi.org/10.5281/zenodo.5042178">FAQ</a></li><li><a href="https://doi.org/10.5281/zenodo.5119878">DrugProt Large Scale Additional SubTrack</a></li><li><a href="https://doi.org/10.5281/zenodo.5656991">DrugProt Large Scale document collection protocol</a></li><li><a href="https://doi.org/10.5281/zenodo.8246229">DrugProt Complete PubMed Knowledge Graph</a><br>&nbsp;</li></ul>

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record