Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “Semantic annotations”
MiRoR11 - P2 - Annotated corpus for semantic similarity of clinical trial outcomes
<p>Outcome similarity corpus</p> <p>This dataset contains annotations of semantic similarity for pairs of primary and reported outcomes.<br> Tab-separated format is used. The files contain the following columns:<br> filename, sentence pair ID, sentence pair text, primary outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_relations_split contains the dataset splits for 10-fold cross-validation.</p>
Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10
<p>Semantic Annotation for Tabular Data with DBpedia: Adapted SemTab 2019 with DBpedia 2016-10</p> <p>Github: https://github.com/phucty/mtab4dbpedia<br> ---------------------------------------------------------------------------------------------------------------------------------------</p> <p>CEA: </p> <ul> <li> <p>Keep only valid entities in DBpedia 2016-10</p> </li> <li> <p>Resolve percentage encoding</p> </li> <li> <p>Add missing redirect entities</p> </li> </ul> <p>CTA: </p> <ul> <li> <p>Keep only valid types</p> </li> <li> <p>Resolve transitive types (parents and equivalent types of the specific type) with DBpedia ontology 2016-10</p> </li> </ul> <p>CPA:</p> <ul> <li> <p>Add equivalent properties</p> </li> </ul> <p>Statistic of Adapted Tabular data SemTab 2019</p> <pre><code>| | CEA | | | CPA | | | CTA | | | |---------|:--------:|:-------:|:------:|:--------:|:-------:|:------:|:--------:|---------|--------| | | Orginal | Adapted | Change | Orginal | Adapted | Change | Orginal | Adapted | Change | | Round 1 | 8418 | 8406 | -0.14% | 116 | 116 | 0.00% | 120 | 120 | 0.00% | | Round 2 | 463796 | 457567 | -1.34% | 6762 | 6762 | 0.00% | 14780 | 14333 | -3.02% | | Round 3 | 406827 | 406820 | 0.00% | 7575 | 7575 | 0.00% | 5762 | 5673 | -1.54% | | Round 4 | 107352 | 107351 | 0.00% | 2747 | 2747 | 0.00% | 1732 | 1717 | -0.87% |</code></pre> <p> </p> <p>---------------------------------------------------------------------------------------------------------------------------------------<br> DBpedia 2016-10 extra resources: (Original dataset http://downloads.dbpedia.org/2016-10/)</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_classes_2016-10.csv</p> <p>Information: DBpedia classes and parents: (We remove the abstract types: Agent, Thing)</p> <p>Total: 759 classes</p> <p>Structure: [class, parents (separate with space)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "City","Location Place PopulatedPlace Settlement"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_properties_2016-10.csv</p> <p>Information: DBpedia properties and these equivalents</p> <p>Total: 2865 properties</p> <p>Structure: [property, it’s equivalent properties] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "restingDate","deathDate"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_domains_2016-10.csv</p> <p>Information: DBpedia properties and these domain types</p> <p>Total: 2421 properties (have types as their domain)</p> <p>Structure: [property, type (domain)] (without prefix dbo: or http://dbpedia.org/ontology/)</p> <p>Example: "deathDate","Person"</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>File: _dbpedia_entities_2016-10.jsonl.bz2 </p> <p>Information: DBpedia entity dump</p> <p>Format: json list bz2 (bz2 Compressed json list)</p> <p>Source: DBpedia dump 2016-10 core</p> <p>Total: 5,289,577 entities (No disambiguation entities)</p> <p>Structure:</p> <p>An entity: for example “Tokyo”: (datatype: dictionary),</p> <p>{</p> <p>'wd': 'Q1322032', (Wikidata ID, datatype: string)</p> <p>'wp': 'Tokyo', (Wikipedia ID, add prefix <a href="https://en.wikipedia.org/wiki/">https://en.wikipedia.org/wiki/</a> + wp to get the Wikipedia URL, datatype: string)</p> <p>'dp': 'Tokyo', (DBpedia ID, add prefix <a href="http://dbpedia.org/resource/">http://dbpedia.org/resource/</a> + dp to get the DBpedia URL, datatype: string)</p> <p>'label': 'Tokyo', (Entity label, datatype: string)</p> <p>'aliases': ['To-kyo', 'Tôkyô Prefecture', ..], (Other entity names, datatype: list) </p> <p>'aliases_multilingual': ['东京小子', 'طوكيو', ...], (Other entity names in multilingual, datatype: list)</p> <p>'types_specific': 'City', (Entity direct type, datatype: string) </p> <p>'types_transitive': ['Human settlement', 'City', 'PopulatedPlace', 'Location', 'Place', 'Settlement'], (Entity transitive types, datatype: list)</p> <p>'claims_entity': { (entity statements, datatype: dictionary. Keys: properties, Values: list of tail entities)</p> <p>'governingBody': ['Tokyo Metropolitan Government'], </p> <p> 'subdivision': ['Honshu', 'Kantō region'],</p> <p>...</p> <p>},</p> <p>'claims_literal': {</p> <p>'string': { (String literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>'postalCode': ['JP-13'], </p> <p>'utcOffset': ['+09:00', '+9'],</p> <p>…</p> <p>}</p> <p>'time': { (Time literal: datatype: dictionary. Keys: properties, Values: list of date time</p> <p>'populationAsOf': ['2016-07-31'], </p> <p>...</p> <p>}), </p> <p>'quantity': { (Numerical literal: datatype: dictionary. Keys: properties, Values: list of values</p> <p>populationDesity: [6224.66, 6349.0], </p> <p>'maximumElevation': [2017], </p> <p>...</p> <p>},</p> <p>'pagerank': 2.2167366040153352e-06 (Entity page rank score calculated on DBpedia Graph)</p> <p>}</p> <p>---------------------------------------------------------------------------------------------------------------------------------------</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish
<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets). </p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p> </p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud, the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others. </p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then, for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L – Scientific Literature: </strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials </strong></strong>contains records from <a href="https://reec.aemps.es/reec/public/web.html">Registro Español de Estudios Clínicos (REEC)</a>. REEC doesn't provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC <a href="https://github.com/luisgasco/REECapi">API</a>. </li> <li><strong>[Subtrack 3 corpus] MESINESP-P – Patents: </strong>This corpus includes patents in Spanish extracted from Google Patents which have the IPC code “A61P” and “A61K31”.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants' predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p> </p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized (folder <em>separated</em>). This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>. </p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train: <ul> <li>training_set_track1_all.json: Full training set for subtrack 1. </li> <li>training_set_track1_only_articles.json: Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json: </li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1. </li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json: Manually annotated development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json: Manually annotated development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip </strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020 set)</li> <li>List of synonyms (the descriptors and synonyms from Latin Spanish DeCS 2020 set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo </strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p> </p> <p><strong>Data format description</strong></p> <p>The <strong>input text files</strong> for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial >140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial >130/80 y <140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP <strong>entity mention files</strong> contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p> </p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L – Scientific Literature </strong>: <ul> <li><em><strong>Training set: </strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish. We have filtered out empty abstracts and non-Spanish abstracts. We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database. We distribute two different datasets: <ul> <li><strong>Articles training set: </strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set: </strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: <ul> <li><strong>Training set: </strong>The training dataset contains records from <a href="https://reec.aemps.es/reec/public/web.html">Registro Español de Estudios Clínicos (REEC)</a>. REEC doesn't provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC <a href="https://github.com/luisgasco/REECapi">API</a>. Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a <a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team. </li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set: </strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P – Patents: </strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code “A61P” and “A61K31”. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set: </strong>We provide a <strong>test set</strong> containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes “A61P” and “A61K31”.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li> We provide this information to the participants as additional data in the “Additional Data” folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p> </p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td> </td> <td> </td> <td> </td> <td> </td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td> </td> <td> </td> <td> </td> <td> </td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p> </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2 Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p> </p> <p>For further information, please email us at luis.gasco@bsc.es</p>
CafeteriaSA corpus: Scientific abstracts annotated across different food semantic resources
<p>In the last decades, a great amount of work has been done in predictive modeling of issues related to human and environmental health. Resolution of issues related to healthcare is made possible by the existence of several biomedical vocabularies and standards, which play a crucial role in understanding health information, together with a large amount of health-related data. However, despite the large number of available resources and work done in the health and environmental domains, there is a lack of semantic resources that can be utilized in the food and nutrition domain, as well as their interconnections. For this purpose, in an European Food Safety Authority-funded project CAFETERIA, we have developed the first annotated corpus of 500 scientific abstracts that consists of 6,407 annotated food entities with regard to Hansard taxonomy, 4,299 for FoodOn, and 3,623 for SNOMED-CT. The CafeteriaSA corpus will enable further development of natural language processing methods for food information extraction from textual data that will allow extracting of food information from scientific textual data.</p>
CafeteriaFCD corpus: Food consumption data annotated with regard to different food semantic resources
<p>The FoodBase curated version which contains 1,000 manually evaluated recipes, annotated with the appropriate semantic tags from the Hansard Taxonomy, FoodON and SNOMED-CT.</p>
Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis
<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>
Semantic annotation of a part of the Italian Copyright Legislation
<p>The dataset is a structured JSONL file focusing on copyright law. Each entry contains key fields that annotate legal texts, mainly in Italian. These fields include:</p> <p>1. ID: A unique numerical identifier.</p> <p>2. Text: Contains the actual legal provisions.</p> <p>3. Chapter ID & Heading: Identifiers and titles for chapters, categorizing the legal text.</p> <p>4. Article and Paragraph ID: Further break down of the text into articles and paragraphs.</p> <p>5. Insertions: Highlights inserted text fragments in the legal text.</p> <p>6. References: Cites external references with URLs and descriptions.</p> <p>7. Entities: Labels sections of the text, identifying their beginning and ending offsets.</p> <p>8. Relations: Intended to describe relationships between entities, although this field is empty in the sample.</p> <p>9. Comments: A field for comments, also empty in the sample.</p> <p> </p>
TweetsCOV19 - A Semantically Annotated Corpus of Tweets About the COVID-19 Pandemic (Part 1, October 2019 - April 2020)
<p><strong><a href="https://data.gesis.org/tweetscov19/">TweetsCOV19</a></strong><strong> </strong>is a semantically annotated corpus of Tweets about the COVID-19 pandemic. It is a subset of <a href="https://data.gesis.org/tweetskb">TweetsKB</a> and aims at capturing online discourse about various aspects of the pandemic and its societal impact. <strong>Metadata</strong> information about the tweets as well as extracted <strong>entities</strong>, <strong>sentiments</strong>, <strong>hashtags</strong>, <strong>user mentions</strong>, and <strong>resolved URLs </strong>are exposed in RDF using established RDF/S vocabularies*.</p> <p>We also provide a <em><strong>tab-separated values (tsv)</strong></em> version of the dataset. Each line contains features of a tweet instance. Features are separated by tab character ("\t"). The following list indicate the feature indices:</p> <ol> <li>Tweet Id: Long.</li> <li>Username: String. Encrypted for privacy issues*.</li> <li>Timestamp: Format ( "EEE MMM dd HH:mm:ss Z yyyy" ).</li> <li>#Followers: Integer.</li> <li>#Friends: Integer.</li> <li>#Retweets: Integer.</li> <li>#Favorites: Integer.</li> <li>Entities: String. For each entity, we aggregated the original text, the annotated entity and the produced score from <a href="https://github.com/yahoo/FEL">FEL</a> library. Each entity is separated from another entity by char ";". Also, each entity is separated by char ":" in order to store "original_text:annotated_entity:score;". If FEL did not find any entities, we have stored "null;".</li> <li>Sentiment: String. <a href="http://sentistrength.wlv.ac.uk/">SentiStrength</a> produces a score for positive (1 to 5) and negative (-1 to -5) sentiment. We splitted these two numbers by whitespace char " ". Positive sentiment was stored first and then negative sentiment (i.e. "2 -1").</li> <li>Mentions: String. If the tweet contains mentions, we remove the char "@" and concatenate the mentions with whitespace char " ". If no mentions appear, we have stored "null;".</li> <li>Hashtags: String. If the tweet contains hashtags, we remove the char "#" and concatenate the hashtags with whitespace char " ". If no hashtags appear, we have stored "null;".</li> <li>URLs: String: If the tweet contains URLs, we concatenate the URLs using ":-: ". If no URLs appear, we have stored "null;"</li> </ol> <p>This dataset consists of <strong>8,151,524 tweets</strong> in total, posted by <strong>3,664,518 users</strong> and reflects the societal discourse about COVID-19 on Twitter in the period of October 2019 until April 2020.</p> <p>To extract the dataset from <a href="https://data.gesis.org/tweetskb">TweetsKB</a>, we compiled a seed list of 268 COVID-19-related <a href="https://data.gesis.org/tweetscov19/keywords.txt">keywords</a>.</p> <p><em>* For the sake of privacy, we anonymize user IDs and we do not provide the text of the tweets.</em></p> <p> </p>
Semantic annotation of PLoS journal citation contexts
<p>Dataset </p>
tFoodL: Larger Semantic Table Annotations Benchmark for Food Domain
<p><strong>tFoodL</strong> is the successor work of <a href="https://zenodo.org/records/10048187">tFood</a> that is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using 10 levels of a recursive hierarchy of related concepts in Wikidata.</p><p>Similar to tFood, it is a dataset for tabular data to knowledge graph matching. It is derived for the Food domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tFoodL</strong> contains 43,255 entity and horizontal tables, while this repository contains only the validation fold (10%) of the entire benchmark with its ground truth data (gt). </p><p>The supported tasks for semantic table annotations are: </p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>
tBiodivL: Larger Semantic Table Annotations Benchmark for Biodiversity Domain
<p><strong>tBiodivL</strong> is a dataset for tabular data to knowledge graph matching. It is derived from the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiodivL</strong> is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using 10 levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283015">tBiodiv</a></p><p><strong>tBiodivL </strong>contains <strong>222,353</strong> entity and horizontal tables, while this repository contains only a sample of <strong>1% of the total generated tables</strong> of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>312 GB</strong>. We will update this repository with the full dataset in the Future.</p><p>Please get in touch if you are interested in the full dataset, </p><p>The supported tasks for semantic table annotations are: </p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>
tBiomedL: Larger Semantic Table Annotations Benchmark for Biomedical Domain
<p><strong>tBiomedL </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiomedL </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using five levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283103">tBiomed</a></p><p><strong>tBiomedL </strong>contains <strong>860,479</strong> entity and horizontal tables, while this repository contains only <strong>a sample of 1%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>27</strong> <strong>GB</strong>. We will update this repository with the full dataset, including the test fold with its ground truth data in the Future.</p><p>Please get in touch if you are interested in the full dataset, </p><p>The supported tasks for semantic table annotations are: </p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>
tBiomed: Semantic Table Annotations Benchmark for Biomedical Domain
<p><strong>tBiomed </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiomed </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p><strong>tBiomed </strong>contains <strong>26,778</strong> entity and horizontal tables, while this repository contains only a <strong>validation fold</strong> of the original data representing <strong>20%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>1</strong> <strong>GB</strong>.</p> <p>We included the full version of the dataset. We will update this repository ground truth data of the test set in the Future.</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>
SemTab 24: Semantic Table Annotations Benchmark for LLM-based approaches
<p><strong>SuperSemtab24 </strong>is a dataset for tabular data to knowledge graph matching.</p> <p>The dataset is divided into training and validation sets. The dataset includes general-purpose tables and intentionally misspelled entities to evaluate the model's robustness. Participants must annotate the entity mentions in the validation set and submit their annotations (following a target file).</p> <p>The repository contains the full version of the dataset; the ground truth (GT) of the test set will be uploaded in the future.</p>
FN-RE: A Corpus of Requirements Documents Enriched with Semantic Frame Annotations
<p>FN-RE is a human-labelled dataset using FrameNet scheme. The dataset is distributed and can be viewed using a web-index page. For further details about the annotation procedures, please refer to the annotation guidelines included in the folder.</p>
A Semantically Annotated 15-Class Ground Truth Dataset for Substation Equipment
<p>This dataset contains 1660 images of electric substations with 50705 annotated objects. The images were obtained using different cameras, including cameras mounted on Autonomous Guided Vehicles (AGVs), fixed location cameras and those captured by humans using a variety of cameras. A total of 15 classes of objects were identified in this dataset, and the number of instances for each class is provided in the following table:</p> <table align="center"> <caption>Object classes and how many times they appear in the dataset.</caption> <thead> <tr> <th scope="col">Class</th> <th scope="col">Instances</th> </tr> </thead> <tbody> <tr> <td>Open blade disconnect</td> <td>310</td> </tr> <tr> <td>Closed blade disconnect switch</td> <td>5243</td> </tr> <tr> <td>Open tandem disconnect switch</td> <td>1599</td> </tr> <tr> <td>Closed tandem disconnect switch</td> <td>966</td> </tr> <tr> <td>Breaker</td> <td>980</td> </tr> <tr> <td>Fuse disconnect switch</td> <td>355</td> </tr> <tr> <td>Glass disc insulator</td> <td>3185</td> </tr> <tr> <td>Porcelain pin insulator</td> <td>26499</td> </tr> <tr> <td>Muffle</td> <td>1354</td> </tr> <tr> <td>Lightning arrester</td> <td>1976</td> </tr> <tr> <td>Recloser</td> <td>2331</td> </tr> <tr> <td>Power transformer</td> <td>768</td> </tr> <tr> <td>Current transformer</td> <td>2136</td> </tr> <tr> <td>Potential transformer</td> <td>654</td> </tr> <tr> <td>Tripolar disconnect switch</td> <td>2349</td> </tr> </tbody> </table> <p>All images in this dataset were collected from a single electrical distribution substation in Brazil over a period of two years. The images were captured at various times of the day and under different weather and seasonal conditions, ensuring a diverse range of lighting conditions for the depicted objects. A team of experts in Electrical Engineering curated all the images to ensure that the angles and distances depicted in the images are suitable for automating inspections in an electrical substation.</p> <p>The file structure of this dataset contains the following directories and files:</p> <p> images: This directory contains 1660 electrical substation images in JPEG format.</p> <p>images: This directory contains 1660 electrical substation images in JPEG format.</p> <ul> <li><strong>labels_json: </strong>This directory contains JSON files annotated in the VOC-style polygonal format. Each file shares the same filename as its respective image in the images directory.</li> <li><strong>15_masks:</strong> This directory contains PNG segmentation masks for all 15 classes, including the porcelain pin insulator class. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>14_masks:</strong> This directory contains PNG segmentation masks for all classes except the porcelain pin insulator. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>porcelain_masks:</strong> This directory contains PNG segmentation masks for the porcelain pin insulator class. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>classes.txt:</strong> This text file lists the 15 classes plus the background class used in LabelMe.</li> <li><strong>json2png.py:</strong> This Python script can be used to generate segmentation masks using the VOC-style polygonal JSON annotations.</li> </ul> <p>The dataset aims to support the development of computer vision techniques and deep learning algorithms for automating the inspection process of electrical substations. The dataset is expected to be useful for researchers, practitioners, and engineers interested in developing and testing object detection and segmentation models for automating inspection and maintenance activities in electrical substations.</p> <p>The authors would like to thank UTFPR for the support and infrastructure made available for the development of this research and COPEL-DIS for the support through project PD-2866-0528/2020—Development of a Methodology for Automatic Analysis of Thermal Images. We also would like to express our deepest appreciation to the team of annotators who worked diligently to produce the semantic labels for our dataset. Their hard work, dedication and attention to detail were critical to the success of this project.</p>
Taxonomies for Semantic Research Data Annotation
<p>This dataset contains 35 of 39 taxonomies that were the result of a systematic review. The systematic review was conducted with the goal of identifying taxonomies suitable for semantically annotating research data. A special focus was set on research data from the hybrid societies domain.</p> <p>The following taxonomies were identified as part of the systematic review:</p> <table> <tbody> <tr> <td> <p><strong>Filename</strong></p> </td> <td> <p><strong>Taxonomy Title</strong></p> </td> </tr> <tr> <td> <p>acm_ccs</p> </td> <td> <p>ACM Computing Classification System [1]</p> </td> </tr> <tr> <td> <p>amec</p> </td> <td> <p>A Taxonomy of Evaluation Towards Standards [2]</p> </td> </tr> <tr> <td> <p>bibo</p> </td> <td> <p>A BIBO Ontology Extension for Evaluation of Scientific Research Results [3]</p> </td> </tr> <tr> <td> <p>cdt</p> </td> <td> <p>Cross-Device Taxonomy [4]</p> </td> </tr> <tr> <td> <p>cso</p> </td> <td> <p>Computer Science Ontology [5]</p> </td> </tr> <tr> <td> <p>ddbm</p> </td> <td> <p>What Makes a Data-driven Business Model? A Consolidated Taxonomy [6]</p> </td> </tr> <tr> <td> <p>ddi_am</p> </td> <td> <p>DDI Aggregation Method [7]</p> </td> </tr> <tr> <td> <p>ddi_moc</p> </td> <td> <p>DDI Mode of Collection [8]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>DemoVoc [9]</p> </td> </tr> <tr> <td> <p>discretization</p> </td> <td> <p>Building a New Taxonomy for Data Discretization Techniques [10]</p> </td> </tr> <tr> <td> <p>dp</p> </td> <td> <p>Demopaedia [11]</p> </td> </tr> <tr> <td> <p>dsg</p> </td> <td> <p>Data Science Glossary [12]</p> </td> </tr> <tr> <td> <p>ease</p> </td> <td> <p>A Taxonomy of Evaluation Approaches in Software Engineering [13]</p> </td> </tr> <tr> <td> <p>eco</p> </td> <td> <p>Evidence & Conclusion Ontology [14]</p> </td> </tr> <tr> <td> <p>edam</p> </td> <td> <p>EDAM: The Bioscientific Data Analysis Ontology [15]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>European Language Social Science Thesaurus [16]</p> </td> </tr> <tr> <td> <p>et</p> </td> <td> <p>Evaluation Thesaurus [17]</p> </td> </tr> <tr> <td> <p>glos_hci</p> </td> <td> <p>The Glossary of Human Computer Interaction [18]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>Humanities and Social Science Electronic Thesaurus [19]</p> </td> </tr> <tr> <td> <p>hcio</p> </td> <td> <p>A Core Ontology on the Human-Computer Interaction Phenomenon [20]</p> </td> </tr> <tr> <td> <p>hft</p> </td> <td> <p>Human-Factors Taxonomy [21]</p> </td> </tr> <tr> <td> <p>hri</p> </td> <td> <p>A Taxonomy to Structure and Analyze Human–Robot Interaction [22]</p> </td> </tr> <tr> <td> <p>iim</p> </td> <td> <p>A Taxonomy of Interaction for Instructional Multimedia [23]</p> </td> </tr> <tr> <td> <p>interrogation</p> </td> <td> <p>A Taxonomy of Interrogation Methods [24]</p> </td> </tr> <tr> <td> <p>iot</p> </td> <td> <p>Design Vocabulary for Human–IoT Systems Communication [25]</p> </td> </tr> <tr> <td> <p>kinect</p> </td> <td> <p>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors [26]</p> </td> </tr> <tr> <td> <p>maco</p> </td> <td> <p>Thesaurus Mass Communication [27]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>Thesaurus Cognitive Psychology of Human Memory [28]</p> </td> </tr> <tr> <td> <p>mixed_initiative</p> </td> <td> <p>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey [29]</p> </td> </tr> <tr> <td> <p>qos_qoe</p> </td> <td> <p>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction [30]</p> </td> </tr> <tr> <td> <p>ro</p> </td> <td> <p>The Research Object Ontology [31]</p> </td> </tr> <tr> <td> <p>senses_sensors</p> </td> <td> <p>A Human-Centered Taxonomy of Interaction Modalities and Devices [32]</p> </td> </tr> <tr> <td> <p>sipat</p> </td> <td> <p>A Taxonomy of Spatial Interaction Patterns and Techniques [33]</p> </td> </tr> <tr> <td> <p>social_errors</p> </td> <td> <p>A Taxonomy of Social Errors in Human-Robot Interaction [34]</p> </td> </tr> <tr> <td> <p>sosa</p> </td> <td> <p>Semantic Sensor Network Ontology [35]</p> </td> </tr> <tr> <td> <p>swo</p> </td> <td> <p>The Software Ontology [36]</p> </td> </tr> <tr> <td> <p>tadirah</p> </td> <td> <p>Taxonomy of Digital Research Activities in the Humanities [37]</p> </td> </tr> <tr> <td> <p>vrs</p> </td> <td> <p>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions [38]</p> </td> </tr> <tr> <td> <p>xdi</p> </td> <td> <p>Cross-Device Interaction [39]</p> </td> </tr> </tbody> </table> <p><br> We converted the taxonomies into SKOS (Simple Knowledge Organisation System) representation. The following 4 taxonomies were not converted as they were already available in SKOS and were for this reason excluded from this dataset:</p> <p>1) DemoVoc, cf. <a href="http://thesaurus.web.ined.fr/navigateur/">http://thesaurus.web.ined.fr/navigateur/</a><br> available at <a href="https://thesaurus.web.ined.fr/exports/demovoc/demovoc.rdf">https://thesaurus.web.ined.fr/exports/demovoc/demovoc.rdf</a></p> <p>2) European Language Social Science Thesaurus, cf. <a href="https://thesauri.cessda.eu/elsst/en/">https://thesauri.cessda.eu/elsst/en/</a><br> available at <a href="https://zenodo.org/record/5506929">https://zenodo.org/record/5506929</a></p> <p>3) Humanities and Social Science Electronic Thesaurus, cf. <a href="https://hasset.ukdataservice.ac.uk/hasset/en/">https://hasset.ukdataservice.ac.uk/hasset/en/</a><br> available at <a href="https://zenodo.org/record/7568355">https://zenodo.org/record/7568355</a></p> <p>4) Thesaurus Cognitive Psychology of Human Memory, cf. <a href="https://www.loterre.fr/presentation/">https://www.loterre.fr/presentation/</a><br> available at <a href="https://skosmos.loterre.fr/P66/en/">https://skosmos.loterre.fr/P66/en/</a></p> <p> </p> <p><strong>References</strong></p> <p>[1] “The 2012 ACM Computing Classification System,” <em>ACM Digital Library</em>, 2012. <a href="https://dl.acm.org/ccs">https://dl.acm.org/ccs</a> (accessed May 08, 2023).</p> <p>[2] AMEC, “A Taxonomy of Evaluation Towards Standards.” Aug. 31, 2016. Accessed: May 08, 2023. [Online]. Available: <a href="https://amecorg.com/amecframework/home/supporting-material/taxonomy/">https://amecorg.com/amecframework/home/supporting-material/taxonomy/</a></p> <p>[3] B. Dimić Surla, M. Segedinac, and D. Ivanović, “A BIBO ontology extension for evaluation of scientific research results,” in <em>Proceedings of the Fifth Balkan Conference in Informatics</em>, in BCI ’12. New York, NY, USA: Association for Computing Machinery, Sep. 2012, pp. 275–278. doi: <a href="https://doi.org/10.1145/2371316.2371376">10.1145/2371316.2371376</a>.</p> <p>[4] F. Brudy <em>et al.</em>, “Cross-Device Taxonomy: Survey, Opportunities and Challenges of Interactions Spanning Across Multiple Devices,” in <em>Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems</em>, in CHI ’19. New York, NY, USA: Association for Computing Machinery, Mai 2019, pp. 1–28. doi: <a href="https://doi.org/10.1145/3290605.3300792">10.1145/3290605.3300792</a>.</p> <p>[5] A. A. Salatino, T. Thanapalasingam, A. Mannocci, F. Osborne, and E. Motta, “The Computer Science Ontology: A Large-Scale Taxonomy of Research Areas,” in <em>Lecture Notes in Computer Science 1137</em>, D. Vrandečić, K. Bontcheva, M. C. Suárez-Figueroa, V. Presutti, I. Celino, M. Sabou, L.-A. Kaffee, and E. Simperl, Eds., Monterey, California, USA: Springer, Oct. 2018, pp. 187–205. Accessed: May 08, 2023. [Online]. Available: <a href="http://oro.open.ac.uk/55484/">http://oro.open.ac.uk/55484/</a></p> <p>[6] M. Dehnert, A. Gleiss, and F. Reiss, “What makes a data-driven business model? A consolidated taxonomy,” presented at the European Conference on Information Systems, 2021.</p> <p>[7] DDI Alliance, “DDI Controlled Vocabulary for Aggregation Method,” 2014. <a href="https://ddialliance.org/Specification/DDI-CV/AggregationMethod_1.0.html">https://ddialliance.org/Specification/DDI-CV/AggregationMethod_1.0.html</a> (accessed May 08, 2023).</p> <p>[8] DDI Alliance, “DDI Controlled Vocabulary for Mode Of Collection,” 2015. <a href="https://ddialliance.org/Specification/DDI-CV/ModeOfCollection_2.0.html">https://ddialliance.org/Specification/DDI-CV/ModeOfCollection_2.0.html</a> (accessed May 08, 2023).</p> <p>[9] INED - French Institute for Demographic Studies, “Thésaurus DemoVoc,” Feb. 26, 2020. <a href="https://thesaurus.web.ined.fr/navigateur/en/about">https://thesaurus.web.ined.fr/navigateur/en/about</a> (accessed May 08, 2023).</p> <p>[10] A. A. Bakar, Z. A. Othman, and N. L. M. Shuib, “Building a new taxonomy for data discretization techniques,” in <em>2009 2nd Conference on Data Mining and Optimization</em>, Oct. 2009, pp. 132–140. doi: <a href="https://doi.org/10.1109/DMO.2009.5341896">10.1109/DMO.2009.5341896</a>.</p> <p>[11] N. Brouard and C. Giudici, “Unified second edition of the Multilingual Demographic Dictionary (Demopaedia.org project),” presented at the 2017 International Population Conference, IUSSP, Oct. 2017. Accessed: May 08, 2023. [Online]. Available: <a href="https://iussp.confex.com/iussp/ipc2017/meetingapp.cgi/Paper/5713">https://iussp.confex.com/iussp/ipc2017/meetingapp.cgi/Paper/5713</a></p> <p>[12] DuCharme, Bob, “Data Science Glossary.” https://www.datascienceglossary.org/ (accessed May 08, 2023).</p> <p>[13] A. Chatzigeorgiou, T. Chaikalis, G. Paschalidou, N. Vesyropoulos, C. K. Georgiadis, and E. Stiakakis, “A Taxonomy of Evaluation Approaches in Software Engineering,” in <em>Proceedings of the 7th Balkan Conference on Informatics Conference</em>, in BCI ’15. New York, NY, USA: Association for Computing Machinery, Sep. 2015, pp. 1–8. doi: <a href="https://doi.org/10.1145/2801081.2801084">10.1145/2801081.2801084</a>.</p> <p>[14] M. C. Chibucos, D. A. Siegele, J. C. Hu, and M. Giglio, “The Evidence and Conclusion Ontology (ECO): Supporting GO Annotations,” in <em>The Gene Ontology Handbook</em>, C. Dessimoz and N. Škunca, Eds., in Methods in Molecular Biology. New York, NY: Springer, 2017, pp. 245–259. doi: <a href="https://doi.org/10.1007/978-1-4939-3743-1_18">10.1007/978-1-4939-3743-1_18</a>.</p> <p>[15] M. Black <em>et al.</em>, “EDAM: the bioscientific data analysis ontology,” <em>F1000Research</em>, vol. 11, Jan. 2021, doi: <a href="https://doi.org/10.7490/f1000research.1118900.1">10.7490/f1000research.1118900.1</a>.</p> <p>[16] Council of European Social Science Data Archives (CESSDA), “European Language Social Science Thesaurus ELSST,” 2021. <a href="https://thesauri.cessda.eu/en/">https://thesauri.cessda.eu/en/</a> (accessed May 08, 2023).</p> <p>[17] M. Scriven, <em>Evaluation Thesaurus</em>, 3rd Edition. Edgepress, 1981. Accessed: May 08, 2023. [Online]. Available: <a href="https://us.sagepub.com/en-us/nam/evaluation-thesaurus/book3562">https://us.sagepub.com/en-us/nam/evaluation-thesaurus/book3562</a></p> <p>[18] Papantoniou, Bill <em>et al.</em>, <em>The Glossary of Human Computer Interaction</em>. Interaction Design Foundation. Accessed: May 08, 2023. [Online]. Available: <a href="https://www.interaction-design.org/literature/book/the-glossary-of-human-computer-interaction">https://www.interaction-design.org/literature/book/the-glossary-of-human-computer-interaction</a></p> <p>[19] “UK Data Service Vocabularies: HASSET Thesaurus.” <a href="https://hasset.ukdataservice.ac.uk/hasset/en/">https://hasset.ukdataservice.ac.uk/hasset/en/</a> (accessed May 08, 2023).</p> <p>[20] S. D. Costa, M. P. Barcellos, R. de A. Falbo, T. Conte, and K. M. de Oliveira, “A core ontology on the Human–Computer Interaction phenomenon,” <em>Data Knowl. Eng.</em>, vol. 138, p. 101977, Mar. 2022, doi: <a href="https://doi.org/10.1016/j.datak.2021.101977">10.1016/j.datak.2021.101977</a>.</p> <p>[21] V. J. Gawron <em>et al.</em>, “Human Factors Taxonomy,” <em>Proc. Hum. Factors Soc. Annu. Meet.</em>, vol. 35, no. 18, pp. 1284–1287, Sep. 1991, doi: <a href="https://doi.org/10.1177/154193129103501807">10.1177/154193129103501807</a>.</p> <p>[22] L. Onnasch and E. Roesler, “A Taxonomy to Structure and Analyze Human–Robot Interaction,” <em>Int. J. Soc. Robot.</em>, vol. 13, no. 4, pp. 833–849, Jul. 2021, doi: <a href="https://doi.org/10.1007/s12369-020-00666-5">10.1007/s12369-020-00666-5</a>.</p> <p>[23] R. A. Schwier, “A Taxonomy of Interaction for Instructional Multimedia.” Sep. 28, 1992. Accessed: May 09, 2023. [Online]. Available: <a href="https://eric.ed.gov/?id=ED352044">https://eric.ed.gov/?id=ED352044</a></p> <p>[24] C. Kelly, J. Miller, A. Redlich, and S. Kleinman, “A Taxonomy of Interrogation Methods,” <em>Psychol. Public Policy Law</em>, vol. 19, p. 165, May 2013, doi: <a href="https://doi.org/10.1037/a0030310">10.1037/a0030310</a>.</p> <p>[25] Y. Chuang, L.-L. Chen, and Y. Liu, “Design Vocabulary for Human-IoT Systems Communication,” in <em>Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems</em>, Montreal QC Canada: ACM, Apr. 2018, pp. 1–11. doi: <a href="https://doi.org/10.1145/3173574.3173848">10.1145/3173574.3173848</a>.</p> <p>[26] N. Díaz Rodríguez, R. Wikström, J. Lilius, M. P. Cuéllar, and M. Delgado Calvo Flores, “Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors,” in <em>Ubiquitous Computing and Ambient Intelligence. Context-Awareness and Context-Driven Interaction</em>, G. Urzaiz, S. F. Ochoa, J. Bravo, L. L. Chen, and J. Oliveira, Eds., in Lecture Notes in Computer Science, vol. 8276. Cham: Springer International Publishing, 2013, pp. 254–261. doi: <a href="https://doi.org/10.1007/978-3-319-03176-7_33">10.1007/978-3-319-03176-7_33</a>.</p> <p>[27] “Thesaurus: mass communication - UNESCO Digital Library.” <a href="https://unesdoc.unesco.org/ark:/48223/pf0000015031">https://unesdoc.unesco.org/ark:/48223/pf0000015031</a> (accessed May 08, 2023).</p> <p>[28] Institute for Scientific and Technical Information, <em>Thesaurus Cognitive Psychology of Human Memory</em>, Version 2.0. 2021. Accessed: May 08, 2023. [Online]. Available: <a href="https://fairsharing.org/FAIRsharing.LcyXdU">https://fairsharing.org/FAIRsharing.LcyXdU</a></p> <p>[29] S. Jiang and R. C. Arkin, “Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey,” in <em>2015 IEEE International Conference on Systems, Man, and Cybernetics</em>, Oct. 2015, pp. 954–961. doi: <a href="https://doi.org/10.1109/SMC.2015.174">10.1109/SMC.2015.174</a>.</p> <p>[30] S. Moller, K.-P. Engelbrecht, C. Kuhnel, I. Wechsung, and B. Weiss, “A taxonomy of quality of service and Quality of Experience of multimodal human-machine interaction,” in <em>2009 International Workshop on Quality of Multimedia Experience</em>, Jul. 2009, pp. 7–12. doi: <a href="https://doi.org/10.1109/QOMEX.2009.5246986">10.1109/QOMEX.2009.5246986</a>.</p> <p>[31] K. Belhajjame <em>et al.</em>, “Using a suite of ontologies for preserving workflow-centric research objects,” <em>J. Web Semant.</em>, vol. 32, pp. 16–42, May 2015, doi: <a href="https://doi.org/10.1016/j.websem.2015.01.003">10.1016/j.websem.2015.01.003</a>.</p> <p>[32] M. Augstein and T. Neumayr, “A Human-Centered Taxonomy of Interaction Modalities and Devices,” <em>Interact. Comput.</em>, vol. 31, no. 1, pp. 27–58, Jan. 2019, doi: <a href="https://doi.org/10.1093/iwc/iwz003">10.1093/iwc/iwz003</a>.</p> <p>[33] J. Jerald, “A Taxonomy of Spatial Interaction Patterns and Techniques,” <em>IEEE Comput. Graph. Appl.</em>, vol. 38, no. 1, pp. 11–19, Jan. 2018, doi: <a href="https://doi.org/10.1109/MCG.2018.011461524">10.1109/MCG.2018.011461524</a>.</p> <p>[34] L. Tian and S. Oviatt, “A Taxonomy of Social Errors in Human-Robot Interaction,” <em>ACM Trans. Hum.-Robot Interact.</em>, vol. 10, no. 2, pp. 1–32, Jun. 2021, doi: <a href="https://doi.org/10.1145/3439720">10.1145/3439720</a>.</p> <p>[35] A. Haller, K. Janowicz, S. Cox, D. Phuoc, K. Taylor, and M. Lefrançois, <em>Semantic Sensor Network Ontology</em>. 2017.</p> <p>[36] J. Malone <em>et al.</em>, “The Software Ontology (SWO): a resource for reproducibility in biomedical data analysis, curation and digital preservation,” <em>J. Biomed. Semant.</em>, vol. 5, no. 1, p. 25, Jun. 2014, doi: <a href="https://doi.org/10.1186/2041-1480-5-25">10.1186/2041-1480-5-25</a>.</p> <p>[37] L. Borek, Q. Dombrowski, J. Perkins, and C. Schöch, “TaDiRAH: a Case Study in Pragmatic Classification,” <em>Digit. Humanit. Q.</em>, vol. 010, no. 1, Feb. 2016.</p> <p>[38] M. A. Muhanna, “Virtual reality and the CAVE: Taxonomy, interaction challenges and research directions,” <em>J. King Saud Univ. - Comput. Inf. Sci.</em>, vol. 27, no. 3, pp. 344–361, Jul. 2015, doi: <a href="https://doi.org/10.1016/j.jksuci.2014.03.023">10.1016/j.jksuci.2014.03.023</a>.</p> <p>[39] F. Scharf, C. Wolters, M. Herczeg, and J. Cassens, “Cross-Device Interaction: Definition, Taxonomy and Application,” presented at the AMBIENT 2013 : The Third International Conference on Ambient Computing, Applications, Services and Technologies, Porto, Portugal: IARIA, 2013, pp. 35–41. Accessed: May 08, 2023. [Online]. Available: <a href="https://www.imis.uni-luebeck.de/de/forschung/publikationen/6380">https://www.imis.uni-luebeck.de/de/forschung/publikationen/6380</a></p>
tFood: Semantic Table Annotations Benchmark for Food Domain
<p>tFood is a dataset for tabular data to knowledge graph matching. It is derived for the Food domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables </strong>are where each of which represents a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> </ol> <p>This dataset version will be used during SemTab 2023 - Round 1. So, the ground truth data for the test set is currently hidden. We will add such ground truth after the conclusion of the challenge. </p> <p> </p> <p> </p>
tBiodiv: Semantic Table Annotations Benchmark for Biodiversity Domain
<p><strong>tBiodiv </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiodiv </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p>We updated this repository with full verion of the dataset, we will update it again with the test ground truth (gt) data in the future.</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>
Collaborative annotation and semantic enrichment of 3D media: Demo of a new FOSS toolchain
<p>A suite of tools for collaborative annotation and semantic enrichment of 3D cultural artefacts is being developed as part of the <a href="https://nfdi4culture.de/">NFDI4Culture</a> project across several partner organisations (led by the <a href="https://www.tib.eu/de/forschung-entwicklung/forschungsgruppen-und-labs/open-science">Open Science Lab at TIB, Hannover</a>). Operating within Task area 1: Data capture and enrichment, the proposed toolchain focuses on the annotation of 3D data within an open knowledge graph environment, so that 3D objects’ metadata and related annotations can be linked to various resources, part of the semantic web and all data is searchable via a public SPARQL endpoint. </p> <p>This short video presents the core steps in the media file upload and annotation workflow facilitated by the toolchain. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.