Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “Semantic Table Annotation”
tFoodL: Larger Semantic Table Annotations Benchmark for Food Domain
<p><strong>tFoodL</strong> is the successor work of <a href="https://zenodo.org/records/10048187">tFood</a> that is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using 10 levels of a recursive hierarchy of related concepts in Wikidata.</p><p>Similar to tFood, it is a dataset for tabular data to knowledge graph matching. It is derived for the Food domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tFoodL</strong> contains 43,255 entity and horizontal tables, while this repository contains only the validation fold (10%) of the entire benchmark with its ground truth data (gt). </p><p>The supported tasks for semantic table annotations are: </p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>
tBiodivL: Larger Semantic Table Annotations Benchmark for Biodiversity Domain
<p><strong>tBiodivL</strong> is a dataset for tabular data to knowledge graph matching. It is derived from the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiodivL</strong> is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using 10 levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283015">tBiodiv</a></p><p><strong>tBiodivL </strong>contains <strong>222,353</strong> entity and horizontal tables, while this repository contains only a sample of <strong>1% of the total generated tables</strong> of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>312 GB</strong>. We will update this repository with the full dataset in the Future.</p><p>Please get in touch if you are interested in the full dataset, </p><p>The supported tasks for semantic table annotations are: </p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>
tBiomedL: Larger Semantic Table Annotations Benchmark for Biomedical Domain
<p><strong>tBiomedL </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiomedL </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using five levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283103">tBiomed</a></p><p><strong>tBiomedL </strong>contains <strong>860,479</strong> entity and horizontal tables, while this repository contains only <strong>a sample of 1%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>27</strong> <strong>GB</strong>. We will update this repository with the full dataset, including the test fold with its ground truth data in the Future.</p><p>Please get in touch if you are interested in the full dataset, </p><p>The supported tasks for semantic table annotations are: </p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>
tBiomed: Semantic Table Annotations Benchmark for Biomedical Domain
<p><strong>tBiomed </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiomed </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p><strong>tBiomed </strong>contains <strong>26,778</strong> entity and horizontal tables, while this repository contains only a <strong>validation fold</strong> of the original data representing <strong>20%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>1</strong> <strong>GB</strong>.</p> <p>We included the full version of the dataset. We will update this repository ground truth data of the test set in the Future.</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>
SemTab 24: Semantic Table Annotations Benchmark for LLM-based approaches
<p><strong>SuperSemtab24 </strong>is a dataset for tabular data to knowledge graph matching.</p> <p>The dataset is divided into training and validation sets. The dataset includes general-purpose tables and intentionally misspelled entities to evaluate the model's robustness. Participants must annotate the entity mentions in the validation set and submit their annotations (following a target file).</p> <p>The repository contains the full version of the dataset; the ground truth (GT) of the test set will be uploaded in the future.</p>
tFood: Semantic Table Annotations Benchmark for Food Domain
<p>tFood is a dataset for tabular data to knowledge graph matching. It is derived for the Food domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables </strong>are where each of which represents a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> </ol> <p>This dataset version will be used during SemTab 2023 - Round 1. So, the ground truth data for the test set is currently hidden. We will add such ground truth after the conclusion of the challenge. </p> <p> </p> <p> </p>
tBiodiv: Semantic Table Annotations Benchmark for Biodiversity Domain
<p><strong>tBiodiv </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiodiv </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p>We updated this repository with full verion of the dataset, we will update it again with the test ground truth (gt) data in the future.</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.