Skip to main content
zenodoopen

tBiodiv: Semantic Table Annotations Benchmark for Biodiversity Domain

<p><strong>tBiodiv </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiodiv </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p>We updated this repository&nbsp; with full verion of the dataset, we will update it again with the test ground truth (gt) data in the future.</p> <p>The supported tasks for semantic table annotations are:&nbsp;</p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics