Text generated by OPUS-MT and T5 models with single-bit errors in the parameters
<h2>Description</h2> <p>The dataset contains text generated using T5 and OPUS-MT model with and with single-bit errors in the parameters of the LLM. The T5 LLM used the <a href="https://huggingface.co/datasets/cnn_dailymail/viewer/3.0.0/test">CNN Daily Mail</a> dataset for summarization and OPUS-MT used the <a href="https://aclanthology.org/2017.iwslt-1.1/">IWSLT2017</a> dataset for Chinese-to-English translation.</p> <p> </p> <p>Folders:</p> <ul> <li>t5_fp32: T5 model with a quantified version of FP32</li> <li>t5_fp16: T5 model with a quantified version of FP16</li> <li>opus_fp32: OPUS-MT model with a quantified version of FP32</li> <li>opus_fp16: OPUS-MT model with a quantified version of FP16</li> </ul> <p>Files:</p> <ul> <li><strong>{cnn/iwslt2017}_input_text.txt</strong>: Input text, that is, text to summarize (cnn and T5) or Chinese text to translate (iwslt2017 and OPUS-MT). For each dataset in total there are <em>number_input_texts.</em></li> <li><strong>{cnn/iwslt2017}_output_reference.txt:</strong> Example of result expected for CNN (T5) and IWSLT2017 (OPUS-MT). For each dataset in total there are <em>number_input_texts.</em></li> <li><strong>{cnn/iwslt2017}_output_predict_fault_free:</strong> Example of predictions without single-bit errors. For each dataset in total there are <em>number_input_texts.</em></li> <li><strong>{cnn/iwslt2017}_output_predict_single_fi_bit_100times:</strong> Example of predictions with 100 different single-bit error. In each dataset in total there are <em>100*number input texts</em>.</li> </ul> <h2>Paper</h2> <ul> <li>Paper: <a href="https://doi.org/10.48550/arXiv.2403.16393">Concurrent Linguistic Error Detection (CLED) for Large Language Models</a></li> <li>Cite:</li> </ul> <p><code>@misc{zhu2024concurrent,</code><br><code> title={Concurrent Linguistic Error Detection (CLED) for Large Language Models}, </code><br><code> author={Jinhua Zhu and Javier Conde and Zhen Gao and Pedro Reviriego and Shanshan Liu and Fabrizio Lombardi},</code><br><code> year={2024},</code><br><code> eprint={2403.16393},</code><br><code> archivePrefix={arXiv},</code><br><code> primaryClass={cs.AI}</code><br><code>}</code></p> <p> </p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0