Evaluation Results of English / Japanese LLMs Using Swallow-Evaluation ver.202407
<h1>Evaluation Results of English / Japanese LLMs Using Swallow-Evaluation ver.202407</h1> <p>This dataset is the source material of our observational analysis paper, "Significance and Effectiveness of Training LLM with Japanese Texts [Saito+, 2024]." It includes the evaluation results of 35 LLMs on 19 Japanese and English tasks, using <a href="https://github.com/swallow-llm/swallow-evaluation">Swallow-evaluation</a> ver.202407.</p> <p>As part of the Swallow project at <a href="https://www.titech.ac.jp/english">Tokyo Institute of Technology</a>, this dataset was developed to enable rigorous and comprehensive comparison of Japanese and English LLMs developed in Japan and worldwide.</p> <h2>Details</h2> <h3>Tasks</h3> <p>Evaluation experiments are conducted on LLMs using 10 datasets for Japanese language understanding and generation tasks, and 9 datasets for English language understanding and generation tasks. All evaluation scores are normalized within a range from 0 (lowest) to 1 (highest). Refer to the reference [Saito+, 2024] for the complete list of evaluation tasks and datasets, evaluation metrics, and task configurations.</p> <h3>Environment</h3> <p>The evaluaitons were primarily conducted on A100 nodes (AIST), using Python as the programming language.</p> <h3>Limitation</h3> <p>While efforts were made to evaluate under fair conditions, considering the unique specifications of each LLM (such as tokenization and system prompts), minor differences in evaluation specifics (like prompt formatting and dependence on eval. environment) may cause task evaluation scores to change independently of the LLM’s performance.</p> <h2>Reference</h2> <p>```<br>@techreport{<br> weko_238505_1,<br> author = "齋藤,幸史郎 and 水木,栄 and 大井,聖也 and 中村,泰士 and 塩谷,泰平 and 前田,航希 and Youmi,Ma and 服部,翔 and 藤井,一喜 and 岡本,拓己 and 石田,茂樹 and 高村,大也 and 横田,理央 and 岡崎,直観",<br> title = "LLMに日本語テキストを学習させる意義",<br> booktitle = "研究報告自然言語処理(NL)",<br> pages = "1--15",<br> year = "2024",<br> institution = "東京工業大学, 東京工業大学/産業技術総合研究所, 東京工業大学, 東京工業大学, 東京工業大学, 東京工業大学, 東京工業大学, 東京工業大学, 東京工業大学, 東京工業大学, 東京工業大学, 産業技術総合研究所, 東京工業大学, 東京工業大学",<br> number = "12",<br> month = "aug" <br>}<br>```</p> <h2>License</h2> <p>This dataset is licensed under a <a href="http://creativecommons.org/licenses/by-sa/4.0/" rel="nofollow">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <h2>Creators</h2> <p>Swallow LLM (<a href="https://github.com/swallow-llm">GitHub</a>, <a href="https://swallow-llm.github.io/index.en.html">Official Web Page</a>)</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 0