Skip to main content
zenodoopen

MUHAI Benchmark : Task 3 (Understanding Complex Concepts)

<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 3 Understanding complex concepts</strong></p> <p>&nbsp;</p> <p>This dataset helps investigating whether&nbsp;symbolic reasoning can help statistical models tro understand complex concepts. Complex concepts are expressed in the form of Image Schemas (i.e.,&nbsp;mental templates that&nbsp;summarise human&nbsp;experiences in the form of patterns of object relations and actions).<br> The submission includes the the ImageSchema dataset with ground truth labels :<br> 1.&nbsp;A question to be asked<br> 2. the type of Image schema (class)<br> 3.&nbsp;the type of phrasing (literal , metaphoric, a distracting sentence)<br> 4. the type of questioning (one referring to the image schema by name, and another describing its content)&nbsp;<br> 5. a question indicating whether it is a yes or no answer<br> 6. the image schema&nbsp;variables identified<br> <br> Each sample in the datasets starts with a question about the presence of the given schema in the following sentence, and follows with a single sentence to be classified as either &quot;yes&quot;&nbsp;or &quot;no&quot;&nbsp;(presence or absence of a schema).<br> <br> This&nbsp;can be used by a system (eg a language model, a symbolic system, a neuro-symbolic approach) to&nbsp;identify image schemas. The&nbsp;file &quot;language-models.csv&quot; includes&nbsp;the results of&nbsp;two language models that were tested (T0pp, GPT-3).</p> <p>Metrics used to evaluate:<br> 1.&nbsp;Accuracy&nbsp;: no. correct&nbsp;predictions&nbsp; / no. of total sentences&nbsp;(TP + TN / P + N)<br> 2. Precision: no. correct&nbsp;image schema predictions&nbsp; /&nbsp;total correct image schema&nbsp;predictions (TP / TP + FP)<br> 3. Recall :&nbsp; &nbsp;: no. correct&nbsp;image schema predictions&nbsp; /&nbsp;total predicted image schema (TP / TP + FN)<br> 4. F1 : harmonic mean of Precision and Recall</p> <p>Code :&nbsp;https://github.com/kmitd/muhai-EPL&nbsp;</p>

ShareScore

48/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
8

Topics