Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
167
datasets available to search
ShareScore release 0.7.1
Dataset results
167 results for “Large Language Models”
Supplemental material for: Software System Testing assisted by Large Language Models: An Exploratory Study
<p>This is the supplemental material of the paper titled as “Software System Testing Assisted by Large Language Models: An Exploratory Study” presented at the 36th International Conference on Testing Software and Systems.</p> <p>It contains the raw execution data generated by both models, GPT-4o and GPT-4omini, during the exploratory study. The supplementary material includes the following files:</p> <ul> <li><em>GPT-4ominiRQ1-2ExecutionData.zip</em>: contains the JSON outputs from the OpenAI API for the GPT-4o mini model. Each output is labeled according to the research question number and the corresponding timestamp (for RQ1) or the requested test case (for RQ2), all provided in plain text format.</li> <li><em>GPT-4oRQ1-2ExecutionData.zip</em>: contains the JSON outputs from the OpenAI API for the GPT-4o model. Like the previous file, each output is named in plain text format based on the research question number and timestamp (for RQ1) or the requested test case (for RQ2).</li> </ul> <p>To cite this work: </p> <p>C. Augusto, J. Morán, A. Bertolino, C. de la Riva and J. Tuya, “S<em>oftware System Testing assisted by Large Language Models: An Exploratory Study</em>”, in <em>Testing Software and Systems</em> (pp. 239–255). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-80889-0_17</p>
ThoughtSource: A central hub for large language model reasoning data (code snapshot)
<p><strong>ThoughtSource is a meta-dataset and software library for chain-of-thought reasoning in large language models (LLMs). This repository contains a snapshot of the associated GitHub repository.</strong></p>
Results and log of LLM-KG-Bench runs described in article "Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering", Meyer et al. 2023
<p>Results and logs of <a href="https://github.com/AKSW/LLM-KG-Bench">LLM-KG-Bench</a> runs described in article "Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering", Meyer et al., to appear in <a href="https://2023-eu.semantics.cc/page/accepted_posters">SEMANTICS 2023 poster track</a> proceedings.</p>
Data and code on the Moral Machine experiment on large language models (LLMs)
<p>As large language models (LLMs) have become more deeply integrated into various sectors, understanding how they make moral judgments has become crucial, particularly in the realm of autonomous driving. This study utilized the Moral Machine framework to investigate the ethical decision-making tendencies of prominent LLMs, including GPT-3.5, GPT-4, PaLM 2, and Llama 2, to compare their responses to human preferences. While LLMs' and humans' preferences such as prioritizing humans over pets and favoring saving more lives are broadly aligned, PaLM 2 and Llama 2, especially, evidence distinct deviations. Additionally, despite the qualitative similarities between the LLM and human preferences, there are significant quantitative disparities, suggesting that LLMs might lean toward more uncompromising decisions, compared to the milder inclinations of humans. These insights elucidate the ethical frameworks of LLMs and their potential implications for autonomous driving.</p>
Robustness of large language models in moral judgments
Open the record for dataset details and reuse information.
Data and code on the Moral Machine experiment on large language models (LLMs)
Open the record for dataset details and reuse information.
Data and code from: Learning a deep language model for microbiomes: The power of large scale unlabeled microbiome data
Open the record for dataset details and reuse information.
Evaluation of large language model chatbot responses to psychotic prompts: numerical ratings of prompt-response pairs
Open the record for dataset details and reuse information.
Replication Package for "Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis"
<h1>Replication Package for the Paper: “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis”</h1> <p>This replication package includes the raw data, questionnaire answers, and a Python notebook needed for reproducing the results detailed in the paper titled “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis.”</p> <h2><a></a>Repository Structure</h2> <ol> <li><strong>Scenarios:</strong> Contains an Excel file encompassing all 141 scenarios collected (in Italian).</li> <li><strong>Training and Validation Messages:</strong> Includes the jsonl files necessary for fine-tuning the model.</li> <li><strong>Testing Messages and Ground Truth:</strong> Contains the messages utilized for testing the models.</li> <li><strong>Results:</strong> Contains Excel files with the responses from the 2 human experts and the 5 model as well as the review of the 3 human reviewer.</li> <li><strong>Tables:</strong> Contains the full Wilcoxon Test Results for H01 and H02 as well as the raw RQs results.</li> </ol> <h2><a></a>Replication Process</h2> <p>To replicate the results of our study, open the provided Python Notebook in Google Colab and follow the instructions to seamlessly reproduce the results.</p> <h1><a></a>Instructions for Use</h1> <p>To utilize this replicability package, refer to the steps outlined in the notebook file.</p> <h1><a></a>Remarks</h1> <p>If you encounter any issues or have any questions, please reach out to the authors of the paper. We will be glad to assist you!</p>
PRICER: Leveraging Few-Shot Learning with Fine-Tuned Large Language Models for Unstructured Economic Data
<p>Describes the taxonomy used in the paper "PRICER: Leveraging Few-Shot Learning with Fine-Tuned Large Language Models for Unstructured Economic Data", presented at the Second Workshop on Semantic Technologies and Deep Learning Models for Scientific, Technical and Legal Data<em> </em>at the Extended Semantic Web Conference (ESWC) 2024.</p>
Knowledge Discovery from Porous Organic Cages Literature Using a Large Language Model
<p>This article presents a GPT-4-based literature reading method that incorporates multi-label text classification and a follow-up information extraction, in which the potential of GPT-4 can be fully exploited to rapidly extract valid information from the literature. In the process of multi-label text classification, the prompt-engineered GPT-4 demonstrated the ability to label text with proper recall rates according to the type of information contained in text, including authors, affiliations, synthetic procedures, surface area, and the CCDC number of corresponding cages. Additionally, GPT-4 demonstrated proficiency in information extraction, effectively transforming labeled text into concise tabulated data. Furthermore, we built a chatbot based on this database, allowing for quick and comprehensive searching across the entire database and responding for cage-related questions.</p>
Dataset for "Pop Quiz! Can a Large Language Model Help WIth Reverse Engineering?"
<p>Source files and scripts for the test framework used in evaluating OpenAI's code-davinci-001 (also known as davinci-codex) model for reverse engineering tasks. All results presented in the associated manuscript are preserved in this dataset, or you can use it to perform additional experiments or re-generate the results.</p>
Human review for post-training improvement of low-resource language performance in large language models
<p>Large language models (LLMs) have significantly improved natural language processing, holding the potential to support health workers and their clients directly. Unfortunately, there is a substantial and variable drop in performance for low-resource languages. Here we present results from an exploratory case study in Malawi, aiming to enhance the performance of LLMs in Chichewa through innovative prompt engineering techniques. By focusing on practical evaluations over traditional metrics, we assess the subjective utility of LLM outputs, prioritizing end-user satisfaction. Our findings suggest that tailored prompt engineering may improve LLM utility in underserved linguistic contexts, offering a promising avenue to bridge the language inclusivity gap in digital health interventions.</p>
Do Large Language Models Have a Personality? A Psychometric Evaluation with Implications for Clinical Medicine and Mental Health AI Dataset
Open the record for dataset details and reuse information.
Artifacts for paper "Code Refinement with the Assistance of Conversation-Aided Large Language Models"
<p>All the conversation data collected and the source code for our method are inside the uploaded zip file.</p>
Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
<p>Autonomous tuning of particle accelerators is an active and challenging field of research with the goal of enabling novel accelerator technologies cutting-edge high-impact applications, such as physics discovery, cancer research and material sciences. A key challenge with autonomous accelerator tuning remains that the most capable algorithms require an expert in optimisation, machine learning or a similar field to implement the algorithm for every new tuning task. In this work, we propose the use of large language models (LLMs) to tune particle accelerators. We demonstrate on a proof-of-principle example the ability of LLMs to successfully and autonomously tune a particle accelerator subsystem based on nothing more than a natural language prompt from the operator, and compare the performance of our LLM-based solution to state-of-the-art optimisation algorithms, such as Bayesian optimisation (BO) and reinforcement learning-trained optimisation (RLO). In doing so, we also show how LLMs can perform numerical optimisation of a highly non-linear real-world objective function. Ultimately, this work represents yet another complex task that LLMs are capable of solving and promises to help accelerate the deployment of autonomous tuning algorithms to the day-to-day operations of particle accelerators.</p>
ETimeline: An Extensive Timeline Generation Dataset based on Large Language Model
<div> <div>Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when viewing news in isolation, making it easier to track the development of stories and understand the interrelationships between key events. Timelines have appeared in many commercial products, but there is a noticeable lack of research in this field in academia, and existing datasets need improvement in terms of effectiveness and scale. We propose the ETimeline dataset, which contains over 13,000 news articles, covering 600 bilingual timelines across 23 news domains. We collected more than 120,000 news articles as a candidate news pool and used the large language model (LLM) Pipeline to enhance performance, ultimately obtaining the ETimeline, and the news pool data will also be provided. This work contributes to the advancement of timeline generation research and supports a wide range of tasks, including topic generation and event relationships. We believe that this dataset will serve as a catalyst for innovative research and bridge the gap between academia and industry in understanding the practical application of technology services.</div> </div>
(supplementary material) Fine-Tuning and Prompt Engineering for Large Language Models-based Code Review Automation
<div> <div> <div> <div>Supplementary material for paper <strong>"Fine-Tuning and Prompt Engineering for Large Language Models-based Code Review Automation"</strong></div> <div> </div> <div> <div> <div>The script for the paper can be found in this GitHub repository: https://github.com/awsm-research/LLM-for-code-review-automatiton</div> </div> </div> </div> </div> </div>
Incentivizing news consumption on social media platforms using large language models and realistic bot accounts
<p>This project examines how to enhance users' exposure to and engagement with verified and ideologically balanced news in an ecologically valid setting. We rely on a large-scale two-week long field experiment on 28,457 Twitter users. We created 28 bots utilizing GPT-2 that replied to users tweeting about sports, entertainment, or lifestyle with a contextual reply containing two hardcoded elements: a URL to the topic-relevant section of quality news organization and an encouragement to follow its Twitter account. Treated users were randomly assigned to receive responses by bots presented as female or male. We examine whether our intervention enhances the following of news media organization, the sharing/liking of news content and the tweeting/liking of political content. We find that the treated users followed more news accounts and the users in the female bot treatment were more likely to like news content than the control.</p>
Results outputs for "Beyond the Hype: Identifying and Analyzing Math Word Problem-Solving Challenges for Large Language Models"
<p>The provided files contain outputs generated by various Large Language Models (LLMs) for solving problems in the SVAMP dataset. Additionally, they include tagged statements of problems that LLMs incorrectly resolved.</p> <p>This repository includes the following two files:</p> <ul> <li>all_data.json --> Contains the generated samples for the SVAMP dataset.</li> <li>df_combined.pkl --> Contains the tagged SVAMP statements of problems that CodeLlama failed to resolve.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.