Skip to main content
zenodoopen

Leaderboard Spanish Language Benchmark for Artificial Intelligence Models (TELEIA)

<h1>TELEIA Datasets Leaderboard</h1> <p>These dataset contains the answers of different LLMs to the <a href="../doi/10.5281/zenodo.12571762">TELEIA (Spanish Language Benchmark for Artificial Intelligence Models)</a> dataset.<br><br><span>LLMs evaluated:</span></p> <ul> <li>Yi-6B-Chat</li> <li>Meta-Llama-3-8B-Instruct</li> <li>Llama-2-7b-chat-hf</li> <li>gemma-7b-it</li> <li>Mistral-7B-Instruct-v0.1</li> <li>occiglot-7b-es-en-instruct</li> <li>GPT3.5</li> <li>GPT4</li> </ul> <p><span>Files:</span></p> <ul> <li><em>TELEIA_Cervantes_AVE_results.xlsx: </em>vocabulary and grammatical structures, following the format of the Cervantes AVE exam</li> <li><em>TELEIA_PCE_results.xlsx: </em>test on morphology and semantics resembling the style of the PCE exam, consisting of short questions or sentences to be completed</li> <li><em>TELEIA_SIELE_results.xlsx:&nbsp;</em>different texts with questions related to them, based on the reading comprehension task of the SIELE exam</li> </ul> <p>Each .xlsx contains a sheet with the results of each model and the following columns:</p> <ul> <li><em>question:&nbsp;</em>question from TELEIA</li> <li><em>option_a:</em> possible answer from TELEIA&nbsp; &nbsp;&nbsp;</li> <li><em>option_b:</em> possible answer from TELEIA<em> &nbsp; &nbsp; &nbsp;&nbsp;</em></li> <li><em>option_c: </em>possible answer from TELEIA <em>&nbsp; &nbsp; &nbsp; &nbsp;</em></li> <li><em>option_d: </em>possible answer from TELEIA<em> &nbsp; &nbsp; &nbsp; &nbsp;</em></li> <li><em>correct_answer:</em> correct answer form TELEIA&nbsp;&nbsp;</li> <li><em>llm_question:</em> complete question made to the LLM&nbsp; &nbsp;&nbsp;</li> <li><em>tokens_in:</em> list of tokens that compound the question&nbsp; &nbsp;&nbsp;</li> <li><em>tokens_in_count:</em> number of tokens that compound the question&nbsp; &nbsp;&nbsp;</li> <li><em>llm_answer:</em> raw answer from the LLM&nbsp; &nbsp;&nbsp;</li> <li><em>llm_answer_filtered:</em> answer in format {A,B,C,D} from the LLM&nbsp; &nbsp;&nbsp;</li> <li><em>tokens_out :</em> list of tokens that compound the raw answer&nbsp; &nbsp;&nbsp;</li> <li><em>tokens_out_count:</em> number of tokens that compound the raw answer &nbsp; &nbsp;</li> <li><em>word_count :</em> &nbsp;number of words that compound the raw answer &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</li> </ul> <p>&nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
16
Reuse readiness
8
Engagement
4