Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.9.0
Dataset results
132 results for “CHATGPT”
Data from: ChatGPT performance on radiation technologist and therapist entry to practice exams
<p>This dataset contains the data needed to reproduce all results and figures described in "ChatGPT performance on radiation technologist and therapist entry to practice exams".</p> <p>Details about the data collection can be found in the paper referenced below. Briefly, ChatGPT (GPT-4) was prompted with multiple choice questions from 4 practice exams provided by the Canadian Association of Medical Radiation Technologists (CAMRT). ChatGPT was promted with the questions from each exam 5 times between July 17 and August 13, 2023. Table 1, below, provides details about the dates for data collection.<br><br></p> <p><strong>Variable descriptions</strong></p> <ul> <li><code>question</code>: Question number, provided by CAMRT. Skipped question numbers indicate image-based questions that were excluded from the study.</li> <li><code>discipline</code>: Indicates the CAMRT exam discipline, abbreviated as follows <ul> <li>RAD: radiological technology</li> <li>MRI: magnetic resonance</li> <li>NUC: nuclear medicine</li> <li>RTT: radiation therapy</li> </ul> </li> <li><code>question_type</code>: Indicates the type of competency being assessed by the question (Knowledge, Application, or Critical thinking). Competency categories were assigned by CAMRT.</li> <li><code>corrrect_response</code>: The correct multiple choice response ("A", "B", "C", or "D"), assigned by CAMRT.</li> <li><code>attempt1-5</code>: ChatGPT's response to the multiple choice questions for attempts 1 through 5, indicated using the letters "A", "B", "C", or "D". In a few cases, ChatGPT did not provide a reference to a multiple choice response and "NA" is recorded in the dataset. </li> </ul> <p><em>Note: The long-form questions from CAMRT and answers provided by ChatGPT are not available as a part of this dataset.<br><br></em></p> <p><strong>Table 1</strong>: Dates for data collection</p> <table> <tbody> <tr> <td> </td> <td><strong>Attempt 1</strong></td> <td><strong>Attempt 2</strong></td> <td><strong>Attempt 3</strong></td> <td><strong>Attempt 4</strong></td> <td><strong>Attempt 5</strong></td> </tr> <tr> <td><strong>Radiological technology</strong></td> <td>2 Aug 2023</td> <td>2 Aug 2023</td> <td>8 Aug 2023</td> <td>9 Aug 2023</td> <td>11 Aug 2023</td> </tr> <tr> <td><strong>Magnetic resonance </strong></td> <td>17 Jul 2023</td> <td>18 Jul 2023</td> <td>18 Jul 2023</td> <td>9 Aug 2023</td> <td>12 Aug 2023</td> </tr> <tr> <td><strong>Nuclear medicine</strong></td> <td>8 Aug 2023</td> <td>9 Aug 2023</td> <td>12 Aug 2023</td> <td>12 Aug 2023</td> <td>12 Aug 2023</td> </tr> <tr> <td><strong>Radiation therapy</strong></td> <td>9 Aug 2023</td> <td>12 Aug 2023</td> <td>12 Aug 2023</td> <td>13 Aug 2023</td> <td>13 Aug 2023</td> </tr> </tbody> </table> <p> </p>
Supporting Data for "Exploring ChatGPT-4 for Transforming Taxonomic Data into OWL: Lessons Learned and Implications for Ontology Development"
<p>Data from the trials with ChatGPT to generate OWL files for taxonomic data from the GBIF Backbone Taxonomy.</p> <p>Updates of version 2: additional prompts from the experiments with Gemini and DeepSeek.</p>
Dataset from the study "Analysis of the accuracy of scientific literature references provided by ChatGPT"
<p>This dataset corresponds to the study carried out to analyse 10 bibliographic references of 10 Spanish authors in the field of Information Sciences requested to the ChatGPT chatbot.</p> <p>The file "Bibliographic_references_ analysis" contains the 10 references returned by ChatGPT for each of the 10 authors (a total of 100 references), together with the variables analysed to check their authenticity.</p> <p>The "Keywords_analysis" file contains the normalisation carried out on the words considered to be key words extracted from the titles of the works, according to which a word cloud showing the frequency of occurrence could be drawn up.</p>
TrainTicket microservice testbench extracted information for our work: Evaluating ChatGPT's Proficiency in Understanding and Answering Microservice Architecture Queries Using Source Code Insights
<p>It contains the CSV file output of our tool implemented in the paper: "Evaluating ChatGPT’s Proficiency in Understanding and Answering Microservice Architecture Queries Using Source Code Insights." applied to the TrainTicket microservice testbench. The information in this CSV was used for In-Context-Learning for ChatGPT.</p>
Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper
<p>Corpora used in the publication:</p> <ul> <li>Cristina España-Bonet. 2023. <strong>Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper. </strong>In <em>Findings of the Association for Computational Linguistics: EMNLP 2023</em>, Singapore. Pages 11757–11777. Association for Computational Linguistics.</li> </ul> <p>Three corpora are included:</p> <ol> <li>Newspaper articles extracted from the OSCAR corpus in English, German, Spanish and Catalan automatically annotated for political stance (left vs right) and topic</li> <li>Newspaper-like article generations by different versions of ChatGPT for 101 topics in the 4 languages</li> <li>Newspaper-like article generations by Bard for 101 topics in the 4 languages</li> </ol> <p>See the README file and the original article for further details.</p>
ChatGPT's Aptitude in Utilizing UML Diagrams for Software Engineering Exercise Generation
<p>The integration of Artificial Intelligence (AI) technologies into educational settings has paved the way for innovative teaching and learning approaches. In Software Engineering (SE) education, using Unified Modeling Language (UML) diagrams is a fundamental teaching element for understanding complex software systems. This research addresses the ability of ChatGPT to utilize UML class and sequence diagrams for creating SE modeling exercises. We use ChatGPT to generate exercises based on the information from uploaded UML diagrams by analyzing textual UML representations such as Mermaid and graphical diagrams. The research explores ChatGPT's ability to synthesize UML-specific information from class and sequence diagrams, enabling the generation of various exercises tailored to strengthen conceptual understanding and practical application. Furthermore, we investigate generating graphical UML class and sequence diagrams based on natural language as input. By bridging the gap between AI-driven natural language understanding and the comprehension of UML diagrams, this study highlights the potential of ChatGPT to improve SE education. Our concise findings address educators, practitioners, and other researchers engaged in the field of SE education with a special focus on UML.</p>
Gene embeddings used in GenePT: A Simple But Hard-to-Beat Foundation Model for Genes and Cells Built From ChatGPT
<p>These are the pulled NCBI (and UniProt, when applicable) summaries of genes, as well as the corresponding OpenAI text embeddings (text-embedding-ada-002 and text-embedding-3-large) computed on the summaries. See methods details in Chen and Zou (2024+).</p> <p>The unzipped folder contains four different files: </p> <ol> <li>NCBI_summary_of_genes.json (NCBI gene card summary of human genes)</li> <li>NCBI_UniProt_summary_of_genes.json (NCBI gene card and UniProt protein (when applicable) summary of human genes)</li> <li>GenePT_gene_embedding_ada_text.pickle (a dictionary of numpy array where gene names (upper case) are keys and text-embedding-ada-002 embeddings of the summary in 1. are the values)</li> <li>GenePT_gene_protein_embedding_model_3_text.pickle (a dictionary of numpy array where gene names (upper case) are keys and text-embedding-3-large embeddings of the summary in 1. are the values)</li> </ol> <p>Reference:</p> <p>Chen YT, Zou J. (2024+) GenePT: A Simple But Effective Foundation Model for Genes and Cells Built From ChatGPT. bioRxiv preprint: <a href="https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1">https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1</a>.</p>
Humans vs ChatGPT texts on TOEFL questions and HC3 dataset
<h2>Description</h2> <p>Human-ChatGPT (gpt4-o) comparison corpus. It extends the <a href="https://huggingface.co/datasets/Hello-SimpleAI/HC3"><strong>HC3</strong></a> dataset and <a href="https://github.com/rexshijaku/chatgpt-generated-text-detection-corpus"><strong>ChatGPT Generated Text Detection corpus</strong></a>. The original datasets include the questions and human & ChatGPT3.5 answers. These datasets extend the originals with answers from gpt4-o. Each line of each file is the gpt4-o answer to each of the questions.</p> <ul> <li><a href="https://huggingface.co/datasets/Hello-SimpleAI/HC3"><strong>HC3 dataset</strong></a>: <ul> <li>finance</li> <li>medicine</li> <li>computing</li> <li>open questions</li> </ul> </li> <li><a href="https://github.com/rexshijaku/chatgpt-generated-text-detection-corpus"><strong>ChatGPT Generated Text Detection corpus</strong></a>: <ul> <li>toefl</li> </ul> </li> <li>Program.py: python script to lemmatize, POS, and clean the human/chatgpt texts</li> </ul> <p> </p> <h2>Paper</h2> <ul> <li>Paper: <a href="https://arxiv.org/abs/2308.07462">Playing with Words: Comparing the Vocabulary and Lexical Richness of ChatGPT and Humans</a></li> <li>Cite:</li> </ul> <p><code>@misc{reviriego2023playing,</code><br><code> title={Playing with Words: Comparing the Vocabulary and Lexical Richness of ChatGPT and Humans}, </code><br><code> author={Pedro Reviriego and Javier Conde and Elena Merino-Gómez and Gonzalo Martínez and José Alberto Hernández},</code><br><code> year={2023},</code><br><code> eprint={2308.07462},</code><br><code> archivePrefix={arXiv},</code><br><code> primaryClass={cs.CL}</code><br><code>}</code></p> <p> </p>
Modèle d'une démarche pour la génération de requête adaptée dans ChatGPT afin d'activer son apprentissage et développer un langage performant
<p>En m'inspirant de diverses sources existant dans la littérature actuelle sur la conception de ressources éducatives libres tel que les lignes directrices de l’UNESCO, j'ai procédé à un entrainement de ChatGPT afin d'obtenir des réponses exactes pour m'aider dans le processus de transformation d'une ressource éducative en ressource éducative libre. Les suggestions de l'outil, furent à chaque fois, traitées, refusées ou validées. <br>Ma démarche consista à suggérer, étape après étape, pour chaque section du contenu web fourni, de proposer à l’outil de réaliser une analyse approfondie basée sur la connaissance experte du domaine.<br>À chaque, étape il fallait indiquer précisément à l’outil la tâche qu’il fallait réaliser. Les premières informations fournies par l’IA furent considérées comme un tremplin pour commencer à construire un premier tableau récapitulatif de données qui allait nous servir de guide/ de base pour la construction d’un langage plus performant et généralisant. </p>
Decálogo para un uso responsable de ChatGPT en la escritura académica y en su enseñanza en la universidad
<p><span><span>Se presenta un decálogo para introducir ChatGPT en la escritura académica en la universidad como asistente. Estas prácticas pueden introducirse en el aula para formar a los estudiantes. El decálogo surge de la siguiente publicación en The Conversation: https://theconversation.com/decalogo-para-escribir-mejor-con-ayuda-de-chatgpt-228547</span></span></p> <p><span><span>Presentamos un decálogo para introducir ChatGPT en la escritura académica en la universidad como asistente. Estas prácticas se pueden introducir en el aula para formar a los estudiantes. The Decalogue was created from the following publication in The Conversation: <span><span>https://theconversation.com/decalogo-para-escribir-mejor-con-ayuda-de-chatgpt-228547</span></span></span></span></p>
DATASET Estimating the Use of ChatGPT in Dental Research Publications
<p>These files contain the data and the research script for analyzing the use of ChatGPT in dental research writing from 2018 to 2024. The dataset includes various CSV files with publication details and analysis scripts to investigate trends and patterns in the usage of specific signaling words associated with ChatGPT.</p>
ChatGPT's responses to TUG-K 4.0 survey (October 2023)
<p>This dataset was created in October 2023, after the initial public release of ChatGPT vision. It contains 60 completed Test of Understanding Graphs in Kinematics 4.0 (TUG-K 4.0) surveys. The responses are sorted by item number. An analysis of the responses was published in <a href="https://doi.org/10.1103/PhysRevPhysEducRes.20.010109">https://doi.org/10.1103/PhysRevPhysEducRes.20.010109</a></p> <p>v2 of the dataset takes care of some accidentally duplicated responses present in v1. Because the analysis in the above cited paper used data directly out of the chatbot, the updates in no way impact the analysis or findings in the paper.</p>
Toward multimodal information and AI interaction: a quasi-experiment with ChatGPT
<p>The development of argumentative text and information comprehension (CoI) skills related to the critical reconstruction of meaning (CT) is crucial in undergraduate education. Especially now in the era of social media and AI-mediated information. Generative AI aids in information creation, but its unconscious use can complicate complex information navigation. Argument maps (AM), commonly used for analyzing analog and static texts, can help visualize, understand, and rework multimodal and dynamic arguments and information.</p> <p>Stemming from the Vygotskian idea, our study used a design-based research approach on the use of AMs and ChatGPT as socio-technical artifacts to stimulate and support the understanding of information (CoI) and thus the development of critical thinking (CT). The workshop introduced the multimodal element through a 3-group quasi-experiment. The first group dealt with fully analog texts, the second group used maps with multimodal textual modes, and the third group only interacted with ChatGPT. The research focused on comparing the three groups and focusing on the two experimental groups (experimental macro-focus). </p> <p>The research had three main objectives: 1) to test whether AMs improved students' CoI enhancement and critical processing (CT); 2) to determine whether interaction with ChatGPT supported information reprocessing and critical construction of opinions and assessment tools; and 3) to determine whether interaction with ChatGPT alone, without AMs, still fostered greater integration of information and viewpoints.</p> <p>Our preliminary analysis showed that AMs improved students' CoI and CT, especially when exposed to multimodal information. ChatGPT interaction increased critical reflection and awareness of AI's role in education. Students using only ChatGPT performed well in argumentative reworking, suggesting that interaction with the chatbot can be effective. However, integrating AMs and ChatGPT could provide optimal support for comprehension and critical thinking skills.</p> <p>This Zenodo record follows the full analysis process with R (https://cran.r-project.org/bin/windows/base/ ) and Nvivo (https://lumivero.com/products/nvivo/) composed of the following datasets, script and results:</p> <p>1. Comprehension of Text and AMs Results - Arg_Map.xlsx</p> <p>2. Critical Thinking level - CriThink.xlsx</p> <p>3. Descriptive and Inferential Statistics Comprehension and Critical Thinking - Preliminary Analysis.R</p> <p>4. Elaboration and Integration Opinion - Opi_G1.xlsx; Opi_G2.xlsx & Opi_G3.xlsx</p> <p>5. Descriptive and Inferential Statistics Opinion level - Preliminary Analysis_opi.R</p> <p>6. Sentiment Analysis - Sentiment Analysis.R</p> <p>7. Vocabulary Frequent words - Vocabulary.csv</p> <p>8. Codebook qualitative Analysis with Nvivo (Codebook.xlsx)</p> <p>9. Results Nvivo Analysis G1 & G2 - Codebook-ChatGPT_G1&G2.docx</p> <p> </p> <p>Any comments or improvements are welcome!</p>
Gold dataset for checking transliteration output from ChatGPT 4.0 and local tool
<p>This dataset includes 385 names and short phrases in Ancient and Modern Greek. This dataset has been used as gold dataset to evaluate the transliteration output from ChatGPT 4.0 and a local transliteration tool developed by the 'International Hellenic University, Department of Information and Electronic Engineering' and 'Open Knowledge Greece'. On the 13th-14th August 2024, both ChatGPT and local tool were tested against two well-known standards, i.e, the ALA-LC Romanization for Greek and the ISO 843 standard. </p> <p>The Gold dataset includes the names and short phrases in Ancient and Modern Greek, the expected output, the ChatGPT output, and the local tool output. For evaluating the results the EXACT excel formula was used. The results of the evaluation are also included in the dataset. </p> <p>The Gold dataset complements a manuscript submitted to the Metadata and Semantics Research (MTSR'24) conference. <br> </p>
Dados abertos do artigo 'Inteligência artificial no levantamento bibliográfico em bases de dados científicos: comparando expressões de busca no ChatGPT, Copilot e Gemini'
<p>Resultado por IA de todas os comandos executados na pesquisa. Texto em formado PDF.<br><br><br>Artigo disponível em → https://doi.org/10.20396/rdbci.v23i00.8678378 ou https://periodicos.sbu.unicamp.br/ojs/index.php/rdbci/article/view/8678378</p>
ChatGPT's answers to requests about the librarian profession
<p>Dataset containing questions and answers about the library profession requested to ChatGPT.</p>
ChatGPT Software Testing Study
<p>This repository contains the dataset and replication package for the TestEd'23 paper </p> <blockquote> <p>Sajed Jalil, Suzzana Rafi, Thomas LaToza, Kevin Moran, and Wing Lam, "ChatGPT and Software Testing Education: Promises & Perils", in Proceedings of the 2nd International Workshop on Software Testing Educaiton (co-located with ICST'23), Dublin Ireland</p> </blockquote> <p> </p>
Dataset of the study: "Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard"
<p>This dataset contains the 30 questions that were posed to the chatbots (i) ChatGPT-3.5; (ii) ChatGPT-4; and (iii) Google Bard, in May 2023 for the study “Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard”. These 30 questions describe mathematics and logic problems that have a unique correct answer. The questions are fully described with plain text only, without the need for any images or special formatting. The questions are divided into two sets of 15 questions each (Set A and Set B). The questions of Set A are 15 “Original” problems that cannot be found online, at least in their exact wording, while Set B contains 15 “Published” problems that one can find online by searching on the internet, usually with their solution. Each question is posed three times to each chatbot. This dataset contains the following: (i) The full set of the 30 questions, A01-A15 and B01-B15; (ii) the correct answer for each one of them; (iii) an explanation of the solution, for the problems where such an explanation is needed, (iv) the 30 (questions) × 3 (chatbots) × 3 (answers) = 270 detailed answers of the chatbots. For the published problems of Set B, we also provide a reference to the source where each problem was taken from.</p>
Investigating the Use of AI-Generated Exercises for Beginner and Intermediate Programming Courses: A ChatGPT Case Study
<p>In recent years, artificial intelligence (AI) has been increasingly used in education and supports teachers in creating educational material and students in their learning progress. AI- driven learning support has recently been further strengthened by the release of ChatGPT, in which users can retrieve expla- nations for various concepts in a few minutes through chat. However, to what extent the use of AI models, such as ChatGPT, is suitable for the creation of didactically and content-wise good exercises for programming courses is not yet known. Therefore, in this paper, we investigate the use of AI-generated exercises for beginner and intermediate programming courses in higher education using ChatGPT. We created 12 exercise sheets with ChatGPT for a beginner to intermediate programming course focusing on the objects-first approach. We report our process, prompts, and experience using ChatGPT for this task and outline good practices we identified. The generated exercises are assessed and revised, primarily using ChatGPT, until they met the requirements of the programming course. We assessed the quality of these exercises by using them in our course as external teaching assignment at the University of Education Ludwigsburg and let the students evaluate them. Results indicate the quality of the generated exercises and the time-saving for creating them using ChatGPT. However, our experience showed that while it is fast to generate a good version of an exercise, almost every exercise requires minor manual changes to improve its quality.</p>
Replication Package for "Improving the Readability of Generated Tests Using GPT-4 and ChatGPT Code Interpreter"
<p>While automated test generation can decrease the human burden associated with testing, it does not eliminate this burden. Humans must still work with generated test cases to interpret testing results, debug the code, build and maintain a comprehensive test suite, and many other tasks. Therefore, a major challenge with automated test generation is understandability of generated test test cases. </p> <p>Large language models (LLMs), machine learning models trained on massive corpora of textual data - including both natural language and programming languages - are an emerging technology with great potential for performing language-related predictive tasks such as translation, summarization, and decision support. </p> <p>In this study, we are exploring the capabilities of LLMs with regard to improving test case understandability.</p> <p>This package contains the data produced during this exploration:</p> <ul> <li>The examples directory contains the three case studies we tested our transformation process on: <ul> <li>queue_example: Tests of a basic queue data structure</li> <li>httpie_sessions: Tests of the sessions module from the httpie project. </li> <li>string_utils_validation: Tests of the validation module from the python-string-utils project.</li> <li>Each directory contains the modules-under-test, the original test cases generated by Pynguin, and the transformed test cases. </li> <li>Two trials were performed per case example of the transformation technique to assess the impact of different results from the LLM.</li> </ul> </li> <li>The survey directory contains the survey that was sent to assess the impact of the transformation on test readability. <ul> <li>survey.pdf contains the survey questions.</li> <li>responses.xlsx contains the survey results.</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.