Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

296

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

296 results for “language models”

Learn how ShareScore rates datasets ↗
zenodo48/100

Language-enhanced cognitive skills model

<p>This is a model of cognitive skills required in the workplace which enhance previous models by including a more detailed measurement of linguistic skills. Linguistic skills are defined as the set of abilities, competencies and knowledge which principally involve the use of linguistic code. More specifically, the linguistic items used are reading and writing competencies, ability to speak, listen or communicate, as well as knowledge of second languages (as a whole). These variables were factorialised together with a list of competencies from previous models of cognitive skills. Principal component analysis (PCA) with equamax rotation was applied to reduce the dimensionality of all items to a few interpretable dimensions according to the correlations between them.&nbsp;The result is nine factors with similar variances among at least three express linguistic-related skills: The first factor expresses the demand for scientific and engineering knowledge. The second refers to a collection of competencies which could be called verbal-reasoning. These include deductive and inductive reasoning skills or those of identifying and solving complex problems. Some linguistic competencies relating to the level of oral and written comprehension and expression are also relevant in this factor. The third factor expresses numerical or quantitative competencies. The fourth expresses the demand for communicative competencies, composed of variables related to efficient communication goals such as clarity of speech, active listening or speaking. The fifth factor expresses creative abilities. The sixth, competencies and knowledge linked to electronics and computers. The seventh expresses managerial competencies. The eighth expresses nurturing competencies and the ninth factor basically expresses knowledge of foreign languages.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Dataset for "Large Language Models as molecular design engines"

<ol> <li><strong>claude-gpt-paper.zip :</strong><br><br>This dataset contains data and results associated with the paper "Large Language Models as molecular design<br>engines" The paper investigates the use of large language models, specifically Claude 3 Opus, for generating and analyzing chemical structures based on various prompts from A-H (as mentioned in the manuscript), and guided design related to electron-withdrawing groups (EWG), electron-donating groups (EDG).</li> </ol> <p>The dataset includes:</p> <ol> <li>PM7 MOPAC energy calculations for generated molecules, along with their SMILES representations and molecule IDs.</li> <li>PM7-calculated charges for the generated molecules.</li> <li>Output files from the Claude 3 Opus language model for each prompt category along.</li> <li>Original dataset (subset of ZINC database) used to build common keys and the initial design space.</li> <li>JSON file containing common keys for featurizing unknown SMILES.</li> <li>PCA object to convert molecule embeddings to 3-dimensional embeddings.</li> </ol> <p>The data is organized into the following folders:</p> <ul> <li><code>pm7_charge_results</code>: Contains HOMO-LUMO energy differences for plotting.</li> <li><code>pm7_charge_calculation</code>: Contains PM7 MOPAC energy calculations and charges.</li> <li><code>out</code>: Contains output files from the Claude 3 Opus language model.</li> <li><code>fact-dropbox</code>: Contains the original dataset, common keys, and PCA object file.</li> </ul> <p>The data can be used to reproduce the results presented in the paper and serve as a foundation for further research in this area.</p> <p>For a detailed description of the folder structure and contents, please refer to the File_descriptions.md file included in the dataset.<br><br><br>2. llm-visulizer-dashapp.zip<br><br>This is the code for the visualizer app for viewing the molecules generated by the LLM. The README.md file has details about running the app.</p> <p>3. claude-gpt-paper-codes.zip&nbsp;</p> <p>This contains the notebook GPT_modification_just_plots.ipynb for plotting, and other codes. The README.md file has details about running the main notebook for getting the plots.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

TDA4ContextualEmbeddings - Public - Debug Data for the codebase of the publication "Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction"

<p>Debug dataset for testing the <a href="https://gitlab.cs.uni-duesseldorf.de/general/dsml/tda4contextualembeddings-public">codebase</a> of the paper <a href="https://doi.org/10.18653/v1/2024.sigdial-1.31">&ldquo;Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction&rdquo;</a> published at the 25th Meeting of the Special Interest Group on Discourse and Dialogue, Kyoto, Japan (SIGDIAL 2024).</p>

openapache2.0Nov 2024View details →
zenodo48/100

Dataset for : A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification

<p>We present&nbsp;a novel solution combining Large Language Model (LLM) capabilities with Formal Verification strategies to falsify and automatically repair software vulnerabilities. Initially, we employ Bounded Model Checking (BMC) to locate the software vulnerability and derive a counterexample. Relying on mathematical proofs, counterexamples provide evidence that the system behaves incorrectly or contains a vulnerability, thereby preventing the generation of false positive alerts. The counterexample that has been detected, along with the source code, are provided to the LLM engine. Our approach involves establishing a specialized prompt language for conducting code debugging and generation to understand the vulnerability&#39;s root cause and repair the code. Finally, we use BMC to verify the corrected version of the code generated by the LLM. As a proof of concept, we create \esbmcai based on the Efficient SMT-based Context-Bounded Model Checker (ESBMC) and a pre-trained Transformer model, specifically gpt-3.5-turbo, to detect and fix errors in C programs. We generated a dataset comprising $1{,}000$ C code samples, each consisting of $20$ to $50$ lines of C code. Experimental results show that our proposed method achieved an impressive success rate of up to $80$\% in repairing vulnerable code, encompassing buffer overflow, arithmetic overflow, and pointer dereference failures. To our knowledge, \esbmcai represents the first proposal for a pioneering initiative to integrate a Large Language Model (LLM) with software model checking. We advocate that this automated approach has the potential to incorporate into the software development lifecycle&#39;s continuous integration and deployment (CI/CD) process.&nbsp;</p> <p>&nbsp;</p> <p>The uploaded&nbsp;dataset contains 1000 codes,&nbsp; each comprising 20&nbsp;to 50&nbsp;lines of C code generated with gpt-3.5-turbo. The material also consists of a version of ESBMC statically compiled with all dependencies, a classifier script, and the output file.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Phlorest phylogeny derived from Chacon & List 2015 'Improved computational models of sound change shed light on the history of the Tukanoan languages'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Chacon TC, List J-M (2015) Improved computational models of sound change shed light on the history of the Tukanoan languages. Journal of Language Relationship, 3:177–203.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

FoodSky: A Food-oriented Large Language Model, and FoodEarth: A Foundamental Food Corpus and Instruction Dataset

<p>Food is the cornerstone of both survival and social life. With the increasing complexity of global dietary needs and preferences, there is a growing demand for food intelligence to enable tasks like recipe recommendation and diet-disease correlation discovery. To address this, we introduce the Food-oriented Large Language Model (LLM) FoodSky, which offers fine-grained perception and reasoning of food data. We constructed a food corpus, FoodEarth, from various authoritative sources to enhance FoodSky's knowledge. We also developed the Topic-based Selective State Space Model and Hierarchical Topic Retrieval Augmented Generation algorithms to improve FoodSky's ability to capture fine-grained food semantics and generate context-aware food-relevant text. Extensive experiments show that FoodSky outperforms general-purpose LLMs on the Chinese National Chef Exam and Dietetic Exam, achieving accuracies of 67.2% and 66.4%, respectively. FoodSky not only enhances culinary creativity and promotes healthier eating patterns but also establishes a new standard for domain-specific LLMs tackling real-world food-related issues.</p>

opencc-zeroSep 2024View details →
zenodo44/100

Replication Package for the paper "Conversing with business process-aware Large Language Models: the BPLLM framework"

<p>Replication Package for the research paper "<em>Conversing with business process-aware Large Language Models: the BPLLM framework</em>".</p> <p>The package includes the process models, the questions (and expected answers), the results of the qualitative evaluation, and the Hugging Face links to the fine-tuned versions of Llama 3.1 8B employed in the quantitative evaluation of the framework.</p> <p>In particular, the process models are:</p> <ul> <li>The natural language Directly-follows graph (DFG) of the Food Delivery process: <em>food_delivery_activities.txt</em> for the definition of the activities and <em>food_delivery_flow.txt</em> for the sequence flow.</li> <li>The BPMN model of the Food Delivery, E-commerce, and Reimbursement processes: <em>ecommerce.bpmn</em>, <em>food_delivery.bpmn</em>, and <em>reimbursement.bpmn</em>.</li> </ul> <p>The datasets with the questions and the expected answers are:</p> <ul> <li><em>1_questions_answers_not_refined_for_DFG.csv</em> ;</li> <li><em>1.1_questions_answers_refined_for_DFG.csv</em> ;</li> <li><em>2_questions_answers_not_refined.csv</em> ;</li> <li><em>3_questions_answers_refined.csv</em> ;</li> <li><em>4_questions_answers_different_processes.csv</em> ;</li> <li><em>5_questions_answers_similar_processes.csv</em> ;</li> <li><em>6_questions_answers_refined_ft.csv</em> .</li> </ul> <p>The complete results of the qualitative evaluation are contained in the file <em>qualitative_experiments_results.pdf</em>.</p> <p>The Hugging Face links to the fine-tuned versions of Llama 3.1 8B are reported in <em>hf_links_finetuned_models.pdf</em>.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Data for "SeaMoon: from protein language models to continuous structural heterogeneity"

<p>Datasets used for development of SeaMoon:&nbsp;<br><a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</p> <p>This upload contains the following data:</p> <ul> <li><strong>precomputed_emb.tar.gz</strong> is a compressed archive containing the precomputed data used for training and testing the models of the SeaMoon method, in Torch <strong>.pt </strong>format.&nbsp;<br>The file prefixes consist of two IDs, "ID1_ID2_", identifying the <a href="https://github.com/PhyloSofS-Team/DANCE">DANCE</a> [1] protein conformational collection used for its generation. "ID1" represents the first member of the collection in alphabetical order, while "ID2" is the reference conformation for the structural alignment. The "ESM_data" or "ProstT5_data" suffixes designate the type of embeddings, generated by either ESM2 [2] or ProstT5 [3].<br>The dictionnary contains the following keys: <ul> <li><strong>emb:</strong> The per-residue embedding.</li> <li><strong>data: </strong>A tuple containing "ID2" (the reference), the amino acid sequence, and the coverage of the positions in the original DANCE collection.</li> <li><strong>eigvect:</strong> The eigenvectors of the covariance matrix of the "ID1_ID2" collection, centered on reference conformaton "D2".</li> <li><strong>eigval:&nbsp;</strong>The associated eigenvalues.</li> <li><strong>ref:</strong> The coordinates of the C-alpha atoms of the reference conformaton "ID2".</li> </ul> </li> <li><strong>train_list.txt, train_list_5ref.txt, val_list.txt </strong>and<strong> test_list.txt</strong> contain the identifiers of the samples used for training and evaluating the SeaMoon models. In the "5ref" setting, we used up to 5 reference conformations per collection.&nbsp;</li> </ul> <p>For details on SeaMoon see:</p> <div> <div>SeaMoon: Prediction of molecular motions based on language models</div> </div> <div>Valentin Lombard, Dan Timsit, Sergei Grudinin, Elodie Laine</div> <div>bioRxiv 2024.09.23.614585; doi: https://doi.org/10.1101/2024.09.23.614585</div> <div>&nbsp;</div> <div>For more information on data usage and generation please see <a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</div> <div>&nbsp;</div> <div>Abstract:</div> <p>How protein move and deform determines their interactions with the environment and is thus of utmost importance for cellular functioning. Following the revolution in single protein 3D structure prediction, researchers have focused on repurposing or developing deep learning models for sampling alternative protein conformations. In this work, we explored whether continuous compact representations of protein motions could be predicted directly from protein sequences, without exploiting nor sampling protein structures. Our approach, called SeaMoon, leverages protein Language Model (pLM) embeddings as input to a lightweight (~1M trainable parameters) convolutional neural network. SeaMoon achieves a success rate of up to 40% when assessed against ~1,000 collections of experimental conformations exhibiting a wide range of motions. SeaMoon capture motions not accessible to the normal mode analysis, an unsupervised physics-based method relying solely on a protein structure's 3D geometry, and generalises to proteins that do not have any detectable sequence similarity to the training set. SeaMoon is easily retrainable with novel or updated pLMs.&nbsp;</p> <p>&nbsp;</p> <p>[1] Lombard, V.; Grudinin, S.; Laine, E. Explaining Conformational Diversity in Protein Families through Molecular Motions. Scientific Data 2024, 11, 752.</p> <p>[2] Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.; Lu, W.; Smetanin, N.; Verkuil, R.; Kabeli, O.; Shmueli, Y.; Dos Santos Costa, A.; Fazel-Zarandi, M.; Sercu, T.; Candido, S.; Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123&ndash;1130.</p> <p>[3] Heinzinger, M.; Weissenow, K.; Sanchez, J. G.; Henkel, A.; Steinegger, M.; Rost, B. ProstT5: Bilingual language model for protein sequence and structure. bioRxiv 2023, 2023&ndash;07.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Chemical Language Model Linker datasets and models

<p><strong>Dataset</strong></p> <p>The unfiltered version of the PubChem dataset used for evaluation in the Chemical Language Model Linker (ChemLML) manuscript. The original dataset comes from the PubChem database. If you use this dataset, please see the PubChem <a href="https://pubchem.ncbi.nlm.nih.gov/docs/downloads">download policies</a> and <a href="https://pubchem.ncbi.nlm.nih.gov/docs/citation-guidelines">citation guidelines</a>.</p> <p>There are entires for 257,619 chemicals, each with the fields:</p> <ul> <li>description</li> <li>Name</li> <li>CID</li> <li>ANID</li> <li>SMILES</li> <li>SELFIES</li> </ul> <p>The PubChem dataset is available under the Creative Commons Zero v1.0 Universal license.</p> <p>Relevant citations:</p> <ul> <li>S. Kim, J. Chen, T. Cheng, A. Gindulyte, J. He, S. He, Q. Li, B. A. Shoemaker, P. A. Thiessen, B. Yu, L. Zaslavsky, J. Zhang, E. E. Bolton, PubChem 2023 update.&nbsp;<em>Nucleic Acids Research</em>&nbsp;<strong>51</strong>, D1373&ndash;D1380 (2023).</li> <li>Y. Deng, S. S. Ericksen, A. Gitter, Chemical Language Model Linker: blending text and molecules with modular adapters. <em>Journal of Chemical Information and Modeling</em> (2025).</li> </ul> <p><strong>Models</strong></p> <p>The `.pth` files are saved PyTorch models. The filenames correspond to the ChemLML models in Table 1 of the ChemLML manuscript. These ChemLML models use the following models from Hugging Face:<br>- <a href="https://huggingface.co/laituan245/molt5-base-smiles2caption">MolT5</a><br>- <a href="https://huggingface.co/GT4SD/multitask-text-and-chemistry-t5-base-standard">Text+Chem T5</a><br>- <a href="https://huggingface.co/zjunlp/MolGen-large">MolGen</a><br>- <a href="https://huggingface.co/zjunlp/MolGen-7b">MolGen-7B</a><br>- <a href="https://huggingface.co/zjunlp/llama2-molinst-molecule-7b">Fine-tuned LLaMA2-7B</a><br>- <a href="https://huggingface.co/allenai/scibert_scivocab_uncased">SCIBERT</a><br>- <a href="https://huggingface.co/facebook/galactica-125m">Galactica</a><br>- <a href="https://huggingface.co/zequnl/molxpt">MolXPT</a></p> <p>See the Hugging Face model cards for the original models' licenses, limitations, and citations.</p> <p>The models are available under the Creative Commons Attribution 4.0 International license.</p>

opencc-zeroOct 2024View details →
zenodo44/100

Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis

<p>This repository contain datasets and results for the paper:</p> <p><strong>Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis</strong></p> <p>&nbsp;</p> <p><strong>Github repository for the code:&nbsp;</strong></p> <p><a href="https://github.com/siebeniris/QuantifyingLanguageConfusion/tree/main">Quantifying Language Confusion GitHub repo</a></p> <p>&nbsp;</p> <p><strong>DATA</strong> include the following datasets:</p> <p>i) raw language graphs and</p> <p>ii) the calculated language similarities from the language graphs,</p> <p>iii) <strong>MTEI</strong>: the files from the <a href="https://github.com/siebeniris/vec2text_exp/tree/aaai">experimental results of multilingual inversion attacks</a>, and calculated language confusion entropy from the data;</p> <p>iv) <strong>LCB</strong>: the files from the <a href="https://github.com/for-ai/language-confusion?tab=Apache-2.0-1-ov-file#readme">language confusion benchmark</a> and calculated language confusion entropy from the data&nbsp;</p> <p>&nbsp;</p> <p><strong>Results</strong> include&nbsp;aggregated results for further analysis:</p> <p>i) <strong>inversion_language_confusion</strong>: results from MTEI</p> <p>ii) <strong>prompting_language_confusion</strong>: results from LCB</p> <p>&nbsp;</p> <p>&nbsp;</p>

openapache2.0Oct 2024View details →
zenodo44/100

Embeddings from protein language models predict conservation and variant effects

<p>For this work, we used protein language model representations (embeddings) to predict sequence conservation without multiple sequence alignments (MSAs). Embeddings alone predicted residue conservation almost as accurately from single sequences as ConSeq using MSAs (two-state Matthew Correlation Coefficient &ndash; MCC - for ProtT5 embeddings of 0.596&plusmn;0.006 vs. 0.608&plusmn;0.006 for ConSeq).</p> <p><strong><em>ConSurf10k</em>- Dataset for the development of ProtT5cons:</strong> The method (ProtT5cons) predicting residue conservation used <em>ConSurf-DB </em>(Ben Chorin et al. 2020). This resource provided sequences and conservation for 89,673 proteins. For all, experimental high-resolution three-dimensional (3D) structures were available in the Protein Data Bank (PDB) (Berman et al. 2000). As standard-of-truth for the conservation prediction, we used the values from ConSurf-DB generated using HMMER (Mistry et al. 2013), CD-HIT (Fu et al. 2012), and MAFFT-LINSi (Katoh and Standley 2013) to align proteins in the PDB (Burley et al. 2019). For proteins from families with over 50 proteins in the resulting MSA, an evolutionary rate at each residue position is computed and used along with the MSA to reconstruct a phylogenetic tree. The ConSurf-DB conservation scores ranged from 1 (most variable) to 9 (most conserved). The PISCES server (Wang and Dunbrack 2003) was used to redundancy reduce the data set such that no pair of proteins had more than 25% pairwise sequence identity. We removed proteins with resolutions &gt;2.5&Aring;, those shorter than 40 residues, and those longer than 10,000 residues. The resulting data set (ConSurf10k) with 10,507 proteins (or domains) was randomly partitioned into training (9,392 sequences), cross-training/validation (555) and test (519) sets.</p> <p>Uploaded data:</p> <ul> <li>ConSuf10k_PDBid_seq_cons.fasta: fasta file with PDBid, sequence and conservation annotation</li> <li>consurf10k_test_ids.txt: txt file with id&#39;s of test set</li> <li>consurf10k_train_ids.txt: txt file with id&#39;s of train set</li> <li>consurf10k_val_ids.txt: txt file with id&#39;s of cross-validation set</li> </ul>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Ecore Metamodels and EcoreBERT Pre-trained Language Model

<p>This dataset contains ecore metamodels from the MAR dataset&nbsp;transformed into tree representations.&nbsp;The original dataset can be found here:&nbsp;<a href="http://mar-search.org/experiments/models20/">http://mar-search.org/experiments/models20/</a></p> <p>The data contained in this repository were used to conduct the experiments in the paper: <strong>Recommending Metamodel Concepts during Modeling Activities with Pre-Trained Language Models.&nbsp;</strong>Link to the paper:&nbsp;<a href="https://arxiv.org/abs/2104.01642">https://arxiv.org/abs/2104.01642</a></p> <p>The data are organized as follows:</p> <ul> <li>model : our model trained on the tree representations of metamodels with RoBERTa architecture.</li> <li>tokenizers : the byte-level BPE tokenizer we used to train our model.</li> <li>train : the training data separated into a training and validation set.</li> <li>test : the test data of all experiments conducted in the paper.</li> </ul> <p>This data repository is linked with the following Github repository containing our code:&nbsp;<a href="https://github.com/mweyssow/ecore-bert">https://github.com/martiwey/metamodel-concepts-bert</a></p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Replication package for "Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation"

<p>This repository contains the replication package for the paper "Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation" by Fernando Vallecillos Ruiz, Anastasiia Grishina, Max Hort and Leon Moonen, accepted for publication in ACM Transactions on Software Engineering and Methodology on 2025-10-09.</p> <p>A preprint is deposited on arXiv with DOI: <a href="https://doi.org/10.48550/arXiv.2401.07994">10.48550/arXiv.2401.07994</a>.</p> <p>The replication package is archived on Zenodo with DOI: <a href="https://doi.org/10.5281/zenodo.10500593">10.5281/zenodo.10500593</a>.&nbsp;It is maintained on GitHub at <a href="https://github.com/secureIT-project/RTT_for_APR">https://github.com/secureIT-project/RTT_for_APR</a>.</p> <p>This project builds on code from the <a href="https://github.com/lin-tan/clm/">clm</a> project, which is (c) 2023, The ASSET research group led by Lin Tan,&nbsp;Purdue University, licensed under the BSD 3-Clause License (see jasper/LICENSE.BSD).&nbsp;All modifications and new contributions are (c) 2025 by the authors of this replication package&nbsp;and distributed under the MIT License (see LICENSE.MIT).&nbsp;The data, models and preprint are distributed under the CC BY 4.0 license.</p> <h2>Citation<code> </code></h2> <p>If you build on this data or code, please cite this work by referring to the paper:</p> <div> <pre><code>@article{ruiz2025:rtt, title = {Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation}, author = {Vallecillos Ruiz, Fernando and Anastasiia Grishina and Max Hort and Leon Moonen}, journal = {ACM Transactions on Software Engineering and Methodology (TOSEM)}, year = {2025}, publisher = {{ACM}} }</code></pre> </div> <h2>Organization</h2> <p>The replication package is organized as follows:</p> <ul> <li>clm-apr <ul> <li>plbart: code to generate patches with PLBART models.</li> <li>codet5: code to generate patches with CodeT5 models.</li> <li>transcoder: code to generate patches with the TransCoder model.</li> <li>incoder: code to generate patches with InCoder models.</li> <li>santacoder: code to generate patches with the SantaCoder model.</li> <li>starcoder: code to generate patches with the StarCoderBase model.</li> <li>quixbugs: code to validate patches generated for the QuixBugs benchmark.</li> <li>defects4j: code to validate patches generated for any of the Defects4J benchmarks.</li> <li>humaneval: code to validate patches generated for the HumanEval-Java benchmark.</li> </ul> </li> <li>humaneval-java: the HumanEval-Java benchmark proposed by Jiang et al. 2023</li> <li>jasper: a Java tool to parse Java programs needed to preprocess input.</li> <li>model: folder to download the language models.</li> <li>analysis_wandb: data from WandB and Jupyter notebook to create graphs.</li> <li>tmp_benchmarks: folder for temporary files used in patch validation. The folder may contain pairs of `paralell&rsquo; folders src and src_org for each benchmark, used to replace buggy code with candidate patches.</li> </ul> <h2>Replication</h2> <h3>Prerequisites</h3> <ul> <li>Python version: 3.8&mdash;3.10.</li> <li><a href="https://git-lfs.com/">Git LFS</a> is required for model downloading.</li> </ul> <h4>Weight and Biases (WandB)</h4> <ol> <li>Create an account on <a href="https://wandb.ai/">Weights and Biases</a></li> <li>Install the <a href="https://docs.wandb.ai/ref/python">Weights and Biases</a> library</li> <li>Run <code>wandb login</code> and follow the instructions</li> </ol> <h4>Set up OpenAI access</h4> <p>OpenAI account is needed with access to <code>gpt-3.5-turbo</code> and <code>gpt-4</code> . The <code>OPENAI_API_KEY</code> environment variable should be set to your OpenAI API access token.</p> <h3>Dependencies</h3> <ul> <li><a href="https://github.com/rjust/defects4j">Defects4J</a> - To generate inputs for the Defects4J datasets or to validate them, you need&nbsp;to have installed <a href="https://github.com/rjust/defects4j">their tool</a>.</li> <li>Java 8</li> <li>Apache Maven</li> </ul> <h3>Setup</h3> <p>We recommend the use of the setup script:</p> <pre><code>setup.sh </code></pre> <p>which performs the following:</p> <ol> <li>Creates a virtual environment for Python and activate it.</li> <li>Install the packages in <code>requirements.txt</code>.</li> <li>Compiles Jasper.</li> <li>Downloads parsers.</li> <li>Check if the Defects4J installation is correct.</li> </ol> <h3>Download models</h3> <p>The following bash script contains the code to download all of the models used:</p> <pre><code>models/download_models.sh </code></pre> <p>We recommend downloading only the models you are going to use due to their size</p> <pre><code>cd models chmod +x download_models.sh ./download_models.sh </code></pre> <p>To run one specific model, for example, PLBART (C#), use the following commands:</p> <pre><code>cd models git lfs install git clone https://huggingface.co/uclanlp/plbart-java-cs git clone https://huggingface.co/uclanlp/plbart-cs-java cd ../.. </code></pre> <h3>Step 1: Preprocessing and Prompting:</h3> <p>Each script in each <code>clm-apr/[model]</code> folder connects one or more models with<br>one dataset. These scripts follow the template: [benchmark]_[model]_[technique].py.<br>The scripts first create an <code>[model]_input.json</code> file with the preprocessed<br>input. Then generate outputs based on that file with one or more models.<br>For example:</p> <pre><code>cd clm-apr/plbart python quixbugs_plbart_round.py # Generates input for QuixBugs and generate patches using Java&lt;-&gt;C# RTT. python quixbugs_plbart_round_nl.py # Generates input for QuixBugs and generate patches using Java&lt;-&gt;NL RTT. </code></pre> <p>Optionally, use argument <code>--device_map cpu</code> if you wish to run the script on<br>CPU, for example:</p> <pre><code>python quixbugs_plbart_round.py --device_map cpu </code></pre> <p>Otherwise, the script will be run on all available CUDA GPU&rsquo;s.</p> <p>We have commented the generation of inputs in the scripts. Users are free to<br>uncomment this method and try for themselves. It is easily recognizable by<br>their name template <code>[model]_[benchmark]_input()</code>. In the previous case:</p> <pre><code>quixbugs_plbart_input() </code></pre> <h3>Step 2 and 3: Round Trip Translation and Postprocessing</h3> <p>These steps are also included in the [benchmark]_[model]_[technique].py<br>script mentioned above. They are modularized in the method recognizable by<br>their name template [model]_[benchmark]_output().<br>For example:</p> <pre><code>quixbugs_incoder_output() </code></pre> <p>This method:</p> <ol> <li>Reads the input json file.</li> <li>Generates outputs through the LLM.</li> <li>Postprocess the output (extract the patch, clean up extra token, etc.).</li> <li>Creates [model]_output_[technique]_[extra].json.</li> </ol> <p>The last 3 steps are repeated according to the number of runs set to performed<br>(10 in our experiments). Each run will produce a different file with the seed<br>used in its generation. For example, <code>quixbugs\_plbart\_round.py</code> and<br><code>quixbugs\_plbart\_round_nl.py</code> scripts create:</p> <pre><code>clm-apr/quixbugs/plbart_results/run_0/plbart_java_cs_java_output_round_csharp_batch.json clm-apr/quixbugs/plbart_results/run_0/plbart_java_nl_java_output_round_nl_batch.json </code></pre> <h3>Step 4: Evaluation of RTT Results:</h3> <p>The last step evaluates the generated outputs against the test-suites of each<br>benchmark. This script reads the previous outputs files and generates a new one<br>with the results of the test for one model. Furthermore, it connects with the<br><em>WandB</em> tool to calculate metrics and send them to analyze.</p> <p>Following the previous examples, to validate the results previously obtained,<br>we execute the following:</p> <pre><code>cd clm-apr/quixbugs python validate_quixbugs_parallel.py </code></pre> <p>Given the included JSON, this script would create:</p> <pre><code>clm-apr/quixbugs/plbart_results/run_0/plbart_java_cs_java_validate_round_csharp_batch.json </code></pre> <p>We have disabled <em>WandB</em> in the script to allow users to try the script first.<br>However, it can be easily activated by changing the parameter <code>mode="disabled"</code><br>to <code>mode="online"</code>.<br>We have set the variable <code>total_runs = 1</code>, as well as <code>input_file</code> and <code>output_file</code><br>to the results included. They should be modified accordingly to validate more runs<br>or to validate other files/models.</p> <h3>Included Results</h3> <p>We include two CSV files obtained through WandB.</p> <pre><code>'data_cleaned_grouped.csv': Aggregated metrics of the 25 outputs for all runs. 'full_data_all_runs.csv': All metrics for all outputs on all runs. </code></pre> <h2>Changelog</h2> <ul> <li>v1.0 - updates corresponding to the accepted version of the manuscript in TOSEM</li> <li>v0.1 - initial replication package corresponding to v1 of arXiv deposit: includes raw data, code, and example outputs.</li> </ul> <h2>References</h2> <p>Jiang, N.; Liu, K.; Lutellier, T.; and Tan, L. 2023. Impact of Code Language<br>Models on Automated Program Repair. In 45th International Conference on<br>Software Engineering (ICSE), 1430&ndash;1442. IEEE. ISBN 978-1-66545-701-9.</p> <div>&nbsp;</div>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Delphi Study: Exploring the Implications of Large Language Models on the Science System

<p><strong>Sample description:</strong> Our target audience consisted of researchers working in the fields of science, technology, and society with a specific interest in Large Language Models (LLMs).</p> <p><strong>Collection method: </strong>Participants were recruited through the professional and personal networks of the authors, as well as the Alexander von Humboldt Institute (HIIG), using a combination of generic emails via LimeSurvey and personal contacts.</p> <p><strong>Description. </strong>The aim of this study was to explore the impact of large language models, specifically ChatGPT, on scholarly practice and academic writing, targeting researchers and experts in the fields of artificial intelligence, science, and technology who publish their research and scientific work. The two-stage Delphi survey sought to identify and assess the potential opportunities and challenges associated with the use of ChatGPT in academic work and scientific writing, with a specific focus on research impact rather than university teaching. Phase 1 yielded 72 responses, while Phase 2 had 52 responses.</p> <p>To conduct our analysis, we developed two distinct codebooks (see Files ChatGPT Delphi Codebook Phase 1.csv and ChatGPT Delphi Codebook Phase 2.csv) for the Delphi study. The first codebook was created by examining approximately half of the responses, extracting relevant information, and generating codes through inductive reasoning. We then categorized and developed subcodes based on these initial codes, assigning them to each participant&#39;s answers using deductive reasoning. For example, when addressing the potential applications of ChatGPT and other language models (LLMs), we identified six subcategories with precise definitions and illustrative examples. The analysis in Phase 1 led to the formulation of ranking questions for Phase 2, focusing on determining the most frequently utilized applications of ChatGPT and other LLMs based on the established codes.</p> <p>During Phase 2, we introduced two additional open-ended questions to explore the impact of ChatGPT and LLMs on the scientific system and society, aiming to envision future scenarios. The analysis of these questions in the second codebook followed a similar approach to Phase 1, including inductive reasoning for code generation and deductive reasoning for assigning codes to the answers. We observed overlapping codes with the Phase 1 codebook and assigned them to the second codebook. Additionally, we noted a shift in the connotation of certain answers from neutral in Phase 1 to being perceived as either positive or negative consequences of ChatGPT and other LLMs. This observation prompted the bifurcation of specific codes to capture the nuanced perspectives. For instance, applications such as reducing administrative tasks initially seen as valuable aids for researchers were sometimes viewed as potential causes for job replacement, implying negative outcomes.</p> <p>For detailed information on the analytical approach employed, including references to these methodologies, please refer to the methodology chapter in the official publication.</p> <p><strong>Content</strong></p> <ol> <li> <p>Questionaire-ChatGPT-Delphi-Phase1-Limesurvey-Export.pdf &ndash; This file file is an exported version of the Phase 1 questionnaire from Limesurvey. It includes the description, socio demographic questions, content questions, and a request for participant naming.</p> </li> <li> <p>Questionaire-ChatGPT-Delphi-Phase2-Limesurvey-Export.pdf &ndash; This file file is an exported version of the Phase 2 questionnaire from Limesurvey. It includes the description, socio demographic questions, content questions, and a request for participant naming.</p> </li> <li> <p>ChatGPT Delphi - Results Phase 1.pdf &ndash; This file contains the responses and corresponding questions from Phase 1 of the Delphi study. The responses provided by the participants are in the form of open-ended answers. As part of this publication, we have ensured the anonymity of the participants.</p> </li> <li> <p>ChatGPT Delphi - Results Phase 2. pdf &ndash; This file contains the responses and corresponding questions from Phase 2 of the Delphi study.&nbsp; It encompasses the ranking answers provided by the participants, as well as two open-ended answers. To maintain anonymity consistently, all participants have been anonymized again in this publication of our results.</p> </li> <li> <p>ChatGPT Delphi Codebook Phase 1.pdf &ndash; This file contains the Phase 1 codebook, which presents the primary codes, their respective subcodes, detailed definitions, and noteworthy examples.</p> </li> <li> <p>ChatGPT Delphi Codebook Phase 2.pdf &ndash; This file contains the Phase 1 codebook, which provides a comprehensive overview of the primary codes within the given scenario. It includes their corresponding subcodes, detailed definitions, and notable examples to enhance understanding and interpretation.</p> </li> </ol>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Dataset Worldwide Survey on the Impact of AI Chatbots and Large Language Models in Dental Education: Insights from Dental Educators

<p><strong>This dataset contains responses from participants regarding their awareness, knowledge, and perceptions of AI-powered tools in dental education. The data was collected during May-June 2023 to investigate the potential enhancement that AI can bring to dental education. The dataset includes variables related to participants&#39; demographics, experiences, perceptions, and opinions.</strong></p> <p><strong>Details in the published protocol by Uribe, S. E., &amp; Maldupa, I. (2023, June 2). Chatbots In Dental Education - Research Protocol. https://doi.org/10.17605/OSF.IO/3BSG2</strong></p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

lilGym: Natural Language Visual Reasoning with Reinforcement Learning, model files

<p>Baselines models&nbsp;for the paper <a href="https://lil.nlp.cornell.edu/lilgym"><em>lil</em>Gym: Natural Language Visual Reasoning with Reinforcement Learning</a>.</p>

openmit-licenseJul 2023View details →
zenodo44/100

PuoBERTa + PuoBERTaJW300 Setswana Language Models

<p><strong>PuoBERTa +&nbsp;PuoBERTaJW300: Setswana Language Models</strong></p><p>A Roberta-based language model specially designed for Setswana, using the new PuoData dataset (PuoBERTa) and PuoData + JW300 TSN (PuoBERTaJW300)</p><p><strong>Cite&nbsp;</strong></p><p>@inproceedings{marivate2023puoberta, title &nbsp; = {PuoBERTa: Training and evaluation of a curated language model for Setswana}, author &nbsp;= {Vukosi Marivate and Moseli Mots'Oehli and Valencia Wagner and Richard Lastrucci and Isheanesu Dzingirai}, year &nbsp; &nbsp;= {2023}, booktitle= {Artificial Intelligence Research. SACAIR 2023. Communications in Computer and Information Science}, url= {https://link.springer.com/chapter/10.1007/978-3-031-49002-6_17}, keywords = {NLP}, preprint_url = {https://arxiv.org/abs/2310.09141}, dataset_url = {https://github.com/dsfsi/PuoBERTa}, software_url = {https://huggingface.co/dsfsi/PuoBERTa} }</p><p><strong>Model Details</strong></p><p>Model Description</p><p>This is a masked language model trained on Setswana corpora, making it a valuable tool for a range of downstream applications from translation to content creation. It's powered by the PuoData dataset to ensure accuracy and cultural relevance.</p><ul><li><strong>Developed by:</strong>&nbsp;Vukosi Marivate (<a href="https://huggingface.co/@vukosi">@vukosi</a>), Moseli Mots'Oehli (<a href="https://huggingface.co/@MoseliMotsoehli">@MoseliMotsoehli</a>) , Valencia Wagner, Richard Lastrucci and Isheanesu Dzingirai</li><li><strong>Model type:</strong>&nbsp;RoBERTa Model</li><li><strong>Language(s) (NLP):</strong>&nbsp;Setswana</li><li><strong>License:</strong>&nbsp;CC BY 4.0</li></ul><p>Usage</p><p>Use this model filling in masks or finetune for downstream tasks. Here's a simple example for masked prediction:</p><p>from transformers import RobertaTokenizer, RobertaModel # Load model and tokenizer model = RobertaModel.from_pretrained('dsfsi/PuoBERTa') tokenizer = RobertaTokenizer.from_pretrained('dsfsi/PuoBERTa')</p><p>&nbsp;</p><p>Downstream Use</p><p>Downstream Performance</p><p><i><strong>MasakhaPOS</strong></i></p><p>Performance of models on the MasakhaPOS downstream task.</p><p>Model Test Performance&nbsp;</p><p><strong>Multilingual Models</strong></p><p>AfroLM</p><p>83.8</p><p>AfriBERTa</p><p>82.5</p><p>AfroXLMR-base</p><p>82.7</p><p>AfroXLMR-large</p><p>83.0</p><p><strong>Monolingual Models</strong></p><p>NCHLT TSN RoBERTa</p><p>82.3</p><p>PuoBERTa</p><p><strong>83.4</strong></p><p>PuoBERTa+JW300</p><p><strong>84.1</strong></p><p>&nbsp;</p><p><i><strong>MasakhaNER</strong></i></p><p>Performance of models on the MasakhaNER downstream task.</p><p>Model Test Performance (f1 score)&nbsp;</p><p><strong>Multilingual Models</strong></p><p>AfriBERTa</p><p>83.2</p><p>AfroXLMR-base</p><p>87.7</p><p>AfroXLMR-large</p><p>89.4</p><p><strong>Monolingual Models</strong></p><p>NCHLT TSN RoBERTa</p><p>74.2</p><p>PuoBERTa</p><p><strong>78.2</strong></p><p>PuoBERTa+JW300</p><p><strong>80.2</strong></p><p>&nbsp;</p><p><strong>Dataset</strong></p><p>We used the PuoData dataset, a rich source of Setswana text, ensuring that our model is well-trained and culturally attuned.</p><p>Citation Information</p><p><strong>Bibtex Reference</strong></p><p>@inproceedings{marivate2023puoberta, title &nbsp; = {PuoBERTa: Training and evaluation of a curated language model for Setswana}, author &nbsp;= {Vukosi Marivate and Moseli Mots'Oehli and Valencia Wagner and Richard Lastrucci and Isheanesu Dzingirai}, year &nbsp; &nbsp;= {2023}, booktitle= {Artificial Intelligence Research. SACAIR 2023. Communications in Computer and Information Science}, url= {https://link.springer.com/chapter/10.1007/978-3-031-49002-6_17}, keywords = {NLP}, preprint_url = {https://arxiv.org/abs/2310.09141}, dataset_url = {https://github.com/dsfsi/PuoBERTa}, software_url = {https://huggingface.co/dsfsi/PuoBERTa} }</p><p>Contributing</p><p>Your contributions are welcome! Feel free to improve the model.</p><p>Model Card Authors</p><p>Vukosi Marivate</p><p>Model Card Contact</p><p>For more details, reach out or check our&nbsp;<a href="https://dsfsi.github.io/">website</a>.</p><p>Email:&nbsp;<a href="mailto:vukosi.marivate@cs.up.ac.za">vukosi.marivate@cs.up.ac.za</a></p><p><strong>Enjoy exploring Setswana through AI!</strong></p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Language Model for Mathematics

<p>This file provides an n-gram&nbsp;language model for mathematics. It was created by parsing papers from arXiv. It is in&nbsp;ARPA format.</p>

opencc-by-sa-4.0Jun 2016View details →
zenodo40/100

HoneyBee: Progressive Instruction Finetuning of Large Language Models for Materials Science

<p>We propose an instruction-based process for trustworthy data curation in materials science (MatSci-Instruct), which we then apply to finetune a LLaMa-based language model targeted for materials science (HoneyBee). MatSci-Instruct helps alleviate the scarcity of relevant, high-quality materials science textual data available in the open literature, and HoneyBee is the first billion-parameter language model specialized to materials science. In MatSci-Instruct we improve the trustworthiness of generated data by prompting multiple commercially available large language models for generation with an Instructor module (e.g. Chat-GPT) and verification from an independent Verifier module (e.g. Claude). Using MatSci-Instruct, we construct a dataset of multiple tasks and measure the quality of our dataset along multiple dimensions, including accuracy against known facts, relevance to materials science, as well as completeness and reasonableness of the data. Moreover, we iteratively generate more targeted instructions and instruction-data in a finetuning-evaluation-feedback loop leading to progressively better performance for our finetuned HoneyBee models. Our evaluation on the MatSci-NLP benchmark shows HoneyBee's outperformance of existing language models on materials science tasks and iterative improvement in successive stages of instruction-data refinement. We study the quality of HoneyBee's language modeling through automatic evaluation and analyze case studies to further understand the model's capabilities and limitations. Our code and relevant datasets are publicly available at https://github.com/BangLab-UdeM-Mila/NLP4MatSci-HoneyBee.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code

<p>Artifact repository for the paper&nbsp;<a href="http://arxiv.org/abs/2308.03109" rel="nofollow"><em>Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code</em></a>, accepted at&nbsp;<em>ICSE 2024</em>, Lisbon, Portugal. Authors are&nbsp;<a href="https://rangeetpan.github.io/" rel="nofollow">Rangeet Pan</a>*&nbsp;<a href="https://alirezai.cs.illinois.edu/" rel="nofollow">Ali Reza Ibrahimzada</a>*,&nbsp;<a href="http://rkrsn.us/" rel="nofollow">Rahul Krishna</a>, Divya Sankar, Lambert Pougeum Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and&nbsp;<a href="https://reyhaneh.cs.illinois.edu/index.htm" rel="nofollow">Reyhaneh Jabbarvand</a>.</p> <h3>Install</h3> <p>This repository contains the source code for reproducing the results in our paper. Please start by cloning this repository:</p> <div> <pre><code>git clone https://github.com/Intelligent-CAT-Lab/PLTranslationEmpirical </code></pre> </div> <p>We recommend using a virtual environment for running the scripts. Please download <code>conda 23.11.0</code>&nbsp;from this&nbsp;<a href="https://docs.conda.io/projects/miniconda/en/latest/miniconda-other-installer-links.html" rel="nofollow">link</a>. You can create a virtual environment using the following command:</p> <div> <pre><code>conda create -n plempirical python=3.10.13 </code></pre> </div> <p>After creating the virtual environment, you can activate it using the following command:</p> <div> <pre><code>conda activate plempirical </code></pre> </div> <p>You can run the following command to make sure that you are using the correct version of Python:</p> <div> <pre><code>python3 --version &amp;&amp; pip3 --version </code></pre> </div> <h3>Dependencies</h3> <p>To install all software dependencies, please execute the following command:</p> <div> <pre><code>pip3 install -r requirements.txt </code></pre> </div> <p>As for hardware dependencies, we used 16 NVIDIA A100 GPUs with 80GBs of memory for inferencing models. The models can be inferenced on any combination of GPUs as long as the reader can properly distribute the model weights across the GPUs. We did not perform weight distribution since we had enough memory (80 GB) per GPU.</p> <p>Moreover, for compiling and testing the generated translations, we used Python 3.10, g++ 11, GCC Clang 14.0, Java 11, Go 1.20, Rust 1.73, and .Net 7.0.14 for Python, C++, C, Java, Go, Rust, and C#, respectively. Overall, we recommend using a machine with Linux OS and at least 32GB of RAM for running the scripts.</p> <p>For running scripts of alternative approaches, you need to make sure you have installed&nbsp;<a href="https://github.com/immunant/c2rust">C2Rust</a>,&nbsp;<a href="https://github.com/gotranspile/cxgo">CxGO</a>, and&nbsp;<a href="https://github.com/paulirwin/JavaToCSharp">Java2C#</a>&nbsp;on your machine. Please refer to their repositories for installation instructions. For Java2C#, you need to create a&nbsp;<code>.csproj</code>&nbsp;file like below:</p> <div> <pre><code>&lt;Project Sdk="Microsoft.NET.Sdk"&gt; &lt;PropertyGroup&gt; &lt;OutputType&gt;Exe&lt;/OutputType&gt; &lt;TargetFramework&gt;net7.0&lt;/TargetFramework&gt; &lt;ImplicitUsings&gt;enable&lt;/ImplicitUsings&gt; &lt;Nullable&gt;enable&lt;/Nullable&gt; &lt;/PropertyGroup&gt; &lt;/Project&gt; </code></pre> </div> <h3>Dataset</h3> <p>We uploaded the dataset we used in our empirical study to&nbsp;<a href="../doi/10.5281/zenodo.8190051" rel="nofollow">Zenodo</a>. The dataset is organized as follows:</p> <ol> <li><a href="https://github.com/IBM/Project_CodeNet">CodeNet</a></li> <li><a href="https://github.com/wasiahmad/AVATAR">AVATAR</a></li> <li><a href="https://github.com/evalplus/evalplus">Evalplus</a></li> <li><a href="https://github.com/apache/commons-cli">Apache Commons-CLI</a></li> <li><a href="https://github.com/pallets/click">Click</a></li> </ol> <p>Please download and unzip the&nbsp;<code>dataset.zip</code>&nbsp;file from Zenodo. After unzipping, you should see the following directory structure:</p> <div> <pre><code>PLTranslationEmpirical ├── dataset ├── codenet ├── avatar ├── evalplus ├── real-life-cli ├── ... </code></pre> </div> <p>The structure of each dataset is as follows:</p> <p>1. CodeNet &amp; Avatar: Each directory in these datasets correspond to a source language where each include two directories&nbsp;<code>Code</code>&nbsp;and&nbsp;<code>TestCases</code>&nbsp;for code snippets and test cases, respectively. Each code snippet has an&nbsp;<code>id</code>&nbsp;in the filename, where the&nbsp;<code>id</code>&nbsp;is used as a prefix for test I/O files.</p> <p>2. Evalplus: The source language code snippets follow a similar structure as CodeNet and Avatar. However, as a one time effort, we manually created the test cases in the target Java language inside a maven project,&nbsp;<code>evalplus_java</code>. To evaluate the translations from an LLM, we recommend moving the generated Java code snippets to the&nbsp;<code>src/main/java</code>&nbsp;directory of the maven project and then running the command&nbsp;<code>mvn clean test surefire-report:report -Dmaven.test.failure.ignore=true</code>&nbsp;to compile, test, and generate reports for the translations.</p> <p>3. Real-life Projects: The&nbsp;<code>real-life-cli</code>&nbsp;directory represents two real-life CLI projects from Java and Python. These datasets only contain code snippets as files and no test cases. As mentioned in the paper, the authors manually evaluated the translations for these datasets.</p> <h3>Scripts</h3> <p>We provide bash scripts for reproducing our results in this work. First, we discuss the translation script. For doing translation with a model and dataset, first you need to create a&nbsp;<code>.env</code>&nbsp;file in the repository and add the following:</p> <div> <pre><code>OPENAI_API_KEY=&lt;your openai api key&gt; LLAMA2_AUTH_TOKEN=&lt;your llama2 auth token from huggingface&gt; STARCODER_AUTH_TOKEN=&lt;your starcoder auth token from huggingface&gt; </code></pre> </div> <p>1. Translation with GPT-4: You can run the following command to translate all&nbsp;<code>Python -&gt; Java</code>&nbsp;code snippets in&nbsp;<code>codenet</code>&nbsp;dataset with the&nbsp;<code>GPT-4</code>&nbsp;while top-k sampling is&nbsp;<code>k=50</code>, top-p sampling is&nbsp;<code>p=0.95</code>, and&nbsp;<code>temperature=0.7</code>:</p> <div> <pre><code>bash scripts/translate.sh GPT-4 codenet Python Java 50 0.95 0.7 0 </code></pre> </div> <p>2. Translation with CodeGeeX: Prior to running the script, you need to clone the CodeGeeX repository from&nbsp;<a href="https://github.com/THUDM/CodeGeeX">here</a>&nbsp;and use the instructions from their artifacts to download their model weights. After cloning it inside&nbsp;<code>PLTranslationEmpirical</code>&nbsp;and downloading the model weights, your directory structure should be like the following:</p> <div> <pre><code>PLTranslationEmpirical ├── dataset ├── codenet ├── avatar ├── evalplus ├── real-life-cli ├── CodeGeeX ├── codegeex ├── codegeex_13b.pt # this file is the model weight ├── ... ├── ... </code></pre> </div> <p>You can run the following command to translate all&nbsp;<code>Python -&gt; Java</code>&nbsp;code snippets in&nbsp;<code>codenet</code>&nbsp;dataset with the&nbsp;<code>CodeGeeX</code>&nbsp;while top-k sampling is&nbsp;<code>k=50</code>, top-p sampling is&nbsp;<code>p=0.95</code>, and&nbsp;<code>temperature=0.2</code>&nbsp;on GPU&nbsp;<code>gpu_id=0</code>:</p> <div> <pre><code>bash scripts/translate.sh CodeGeeX codenet Python Java 50 0.95 0.2 0 </code></pre> </div> <p>3. For all other models (StarCoder, CodeGen, LLaMa, TB-Airoboros, TB-Vicuna), you can execute the following command to translate all&nbsp;<code>Python -&gt; Java</code>&nbsp;code snippets in&nbsp;<code>codenet</code>&nbsp;dataset with the&nbsp;<code>StarCoder|CodeGen|LLaMa|TB-Airoboros|TB-Vicuna</code>&nbsp;while top-k sampling is&nbsp;<code>k=50</code>, top-p sampling is&nbsp;<code>p=0.95</code>, and&nbsp;<code>temperature=0.2</code>&nbsp;on GPU&nbsp;<code>gpu_id=0</code>:</p> <div> <pre><code>bash scripts/translate.sh StarCoder codenet Python Java 50 0.95 0.2 0 </code></pre> </div> <p>4. For translating and testing pairs with traditional techniques (i.e., C2Rust, CxGO, Java2C#), you can run the following commands:</p> <div> <pre><code>bash scripts/translate_transpiler.sh codenet C Rust c2rust fix_report bash scripts/translate_transpiler.sh codenet C Go cxgo fix_reports bash scripts/translate_transpiler.sh codenet Java C# java2c# fix_reports bash scripts/translate_transpiler.sh avatar Java C# java2c# fix_reports </code></pre> </div> <p>5. For compile and testing of CodeNet, AVATAR, and Evalplus (Python to Java) translations from GPT-4, and generating fix reports, you can run the following commands:</p> <div> <pre><code>bash scripts/test_avatar.sh Python Java GPT-4 fix_reports 1 bash scripts/test_codenet.sh Python Java GPT-4 fix_reports 1 bash scripts/test_evalplus.sh Python Java GPT-4 fix_reports 1 </code></pre> </div> <p>6. For repairing unsuccessful translations of Java -&gt; Python in CodeNet dataset with GPT-4, you can run the following commands:</p> <div> <pre><code>bash scripts/repair.sh GPT-4 codenet Python Java 50 0.95 0.7 0 1 compile bash scripts/repair.sh GPT-4 codenet Python Java 50 0.95 0.7 0 1 runtime bash scripts/repair.sh GPT-4 codenet Python Java 50 0.95 0.7 0 1 incorrect </code></pre> </div> <p>7. For cleaning translations of open-source LLMs (i.e., StarCoder) in codenet, you can run the following command:</p> <div> <pre><code>bash scripts/clean_generations.sh StarCoder codenet </code></pre> </div> <p>Please note that for the above commands, you can change the dataset and model name to execute the same thing for other datasets and models. Moreover, you can refer to&nbsp;<a href="https://github.com/Intelligent-CAT-Lab/PLTranslationEmpirical/blob/main/prompts/README.md"><code>/prompts</code></a>&nbsp;for different vanilla and repair prompts used in our study.</p> <h3>Artifacts</h3> <p>Please download the&nbsp;<code>artifacts.zip</code>&nbsp;file from our&nbsp;<a href="../doi/10.5281/zenodo.8190051" rel="nofollow">Zenodo</a>&nbsp;repository. We have organized the artifacts as follows:</p> <ol> <li>RQ1 - Translations: This directory contains the translations from all LLMs and for all datasets. We have added an excel file to show a detailed breakdown of the translation results.</li> <li>RQ2 - Manual Labeling: This directory contains an excel file which includes the manual labeling results for all translation bugs.</li> <li>RQ3 - Alternative Approaches: This directory contains the translations from all alternative approaches (i.e., C2Rust, CxGO, Java2C#). We have added an excel file to show a detailed breakdown of the translation results.</li> <li>RQ4 - Mitigating Translation Bugs: This directory contains the fix results of GPT-4, StarCoder, CodeGen, and Llama 2. We have added an excel file to show a detailed breakdown of the fix results.</li> </ol> <h3>Contact</h3> <p>We look forward to hearing your feedback. Please contact&nbsp;<a href="mailto:rangeet.pan@ibm.com">Rangeet Pan</a>&nbsp;or&nbsp;<a href="mailto:alirezai@illinois.edu">Ali Reza Ibrahimzada</a> for any questions or comments 🙏.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record