GitTables benchmark - column type detection
<p><strong>Note: the download page of the entire GitTables corpus is here: <a href="https://zenodo.org/record/4943312">https://zenodo.org/record/4943312</a>.</strong></p> <p>This dataset represents a <strong>small subset</strong> of tables from <a href="https://gittables.github.io">GitTables</a> curated for benchmarking column type detection methods. This benchmark evaluates systems that match table columns to semantic types from the <a href="http://wikidata.dbpedia.org/services-resources/ontology">DBpedia</a> and <a href="https://schema.org/docs/schemas.html">Schema.org</a> ontologies. It is featured in the <a href="https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2021/index.html">SemTab 2021 challenge</a> (CTA task).</p> <p>This dataset consists of the following files:</p> <ul> <li>“tables.zip”: directory with a sample of 1101 tables from GitTables. Filenames correspond to table IDs, the first column (without column name) corresponds to row indices, column names are replaced with "col_0", "col_1", etc. which match to the targets and labels (semantic types).</li> <li>“<ontology>_targets.csv”: target columns per table ID, 1 file per ontology (DBpedia or Schema.org): columns are “table_id” (ignore the "_<ontology>" suffix) and “target_column” (i.e. the column that should be annotated).</li> <li>“<ontology>_gt.csv”: ground truth column annotations per table ID, 1 file per ontology: columns are “table_id” (ignore the _<ontology> suffix), “target_column”, “annotation_id”, “annotation_label”.</li> <li>“<ontology>_labels.csv”: unique labels present in the annotated tables, 1 file per ontology: columns are “annotation_id” and “annotation_label”.</li> </ul> <p>The labels (semantic types) from each ontology come from:</p> <ul> <li>DBpedia: only properties from the <a href="http://wikidata.dbpedia.org/services-resources/ontology">DBpedia ontology</a>.</li> <li>Schema.org: properties and types from <a href="https://schema.org/docs/schemas.html">Schema.org</a>.</li> </ul> <p>For the entire GitTables corpus, please refer to <a href="https://zenodo.org/record/4943312#.YZQUWr3ML6Y.">this dataset</a>. Visit <a href="https://gittables.github.io">https://gittables.github.io</a> for more background and contact details.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0