Hackathon - TF-TG literature triage unlabelled data
<p>Once literature triage system is ready it is time to actually try to apply if to records that do not have any label in order to find the subset that does describe TF-TG interactions (are relevant). This is the corpus that has to be labeled by the systems created (hopefully) during the hackathon. To make the results more useful we have pre-selected records that do mention TFs by exploiting either automatic human TF mention recognition or external references from databases that have manually curated information on transcription factors (from GeneRif or UniProt). This means that these abstracts should be enriched with TF relevant records. This record has the same format as the training data except that the last column with the class label is missing.</p> <p>It contains PMIDs and Abstracts.</p> <ul> <li> <p>Name: <a href="https://zenodo.org/record/2562913/files/greekc_triage_unlabelled_v01.tsv?download=1">greekc_triage_unlabelled_v01.tsv</a></p> </li> <li> <p>Example:</p> </li> <li> <p>Format: tsv-separated columns (PMID, PubAnnotation JSON formated results of Pubtator for this record together with the automatically detected gene mentions using GnormPlus providing the Entrez Gene Identifiers together with the mention offsets, i.e. start and end character positions</p> </li> <li> <p>PubAnnotation format description: <a href="http://www.pubannotation.org/docs/annotation-format/">http://www.pubannotation.org/docs/annotation-format/</a></p> </li> <li> <p>PubTator record retrieval description:</p> </li> </ul> <p>https://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/tmTools/curl.html</p> <p><strong>Warning:</strong> This file is quite big!</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0