Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
296
datasets available to search
ShareScore release 0.9.0
Dataset results
296 results for “language models”
Understanding the Rare Inflammatory Disease Using Large Language Models and Social Media Data
Open the record for dataset details and reuse information.
Topologies, Checkpoints, and Configurations for the paper "GVI-RL: Graph-Invariant RL for Attack Paths Discovery using Vulnerabilities Embedded with Large Language Models"
<p>This repository consists of the <strong>files</strong> related to the <strong>paper</strong> "GVI-RL: Graph-Invariant RL for Attack Paths Discovery using Vulnerabilities Embedded with Large Language Models". In particular, this repository contains tensorboard logs, topologies, checkpoints, seeds, and results to ensure reproducibility of the results of the paper.</p> <p>The results included are related to the training/validation and hyper-parameters optimization of the outcome multi-label classifier, the GVI-RL agent, and the world model.<br>The data folder contains also the topologies used in the study, the vulnerabilities data used to generate them and the dataset for multi-label classification.</p> <p>The README.md file describes the folders' structure.</p>
The comparison of structured abstracts generated by the ChatGPT language model with the author's original abstracts derived from research papers
<p>The study utilized publications in the field of information science, both in Polish and English, which appeared in the journal <em>Zagadnienia Informacji Naukowej – Studia Informacyjne</em> during the 2022-2023 period. A total of 10 research papers were selected – 5 in Polish (PL1-PL5) and 5 in English (EN1-EN5).</p>
Artifacts for "On Hardware Security Bug Code Fixes By Querying Large Language Models""
<p>This repository contains the benchmarks and results obtained for the work "On Hardware Security Bug Code Fixes<br> By Querying Large Language Models".<br> Follow the README.md file for more information on how to use the tools yourself.</p>
RNA large language models embeddings on benchmark datasets
<p>This repository contains pre-computed embeddings for several RNA sequences, using most recent Large Language Models (LLM) pre-trained on RNA sequences. </p> <p>compressed files for each combination of RNA-LLM models and benchmarking RNA datasets. </p>
Data for: "Can Language Models Recognize Convincing Arguments?"
<p>For a description, see: https://go.epfl.ch/persuasion-llm</p>
Character Level GPT Language Model for Biomedical Tables
<p>This tarball contains a character level GPT language model (LM) trained on over 11 million tables from PMC OAS full papers. The tarball also contains a row merger classifier finetuned from the Table LM for key resource table reconstruction from papers/preprints in PDF format. This tarball is to be used together with the codebases <a href="https://github.com/SciCrunch/table_lm">https://github.com/SciCrunch/table_lm</a> and <a href="https://github.com/SciCrunch/key_resource_table_extractor">https://github.com/SciCrunch/key_resource_table_extractor</a>.</p>
Supplementary material for Evaluating Legal Compliance of Smart Contracts Generated by Large Language Models
<p>This repository contains the supplementary material for the paper titled "Evaluating Legal Compliance of Smart Contracts Generated by Large Language Models". It includes natural-language legal contracts, their smart contract implementations, and Petri net models of said legal contracts contracts.</p>
Duplicate Bug Report Detection using an Attention-basedPre-trained Neural Language Model
<p>A BERT based Approach for Automatic Duplicate Bug Report Detection</p>
Topic model of English-language fiction, 1880-1999, with 200 topics.
<p>A topic model of 29,341 volumes of fiction, written in English and published between 1880 and 1999. The underlying corpus was organized by Ted Underwood for an experiment on period and cohort effects in cultural change. Metadata is in finalcorpus.tsv (which also has rows for 10 volumes not actually included in the model). </p> <p>To identify volumes as fiction, we relied on the NovelTM Dataset of English-Language Fiction (https://culturalanalytics.org/article/13147-noveltm-datasets-for-english-language-fiction-1700-2009). To confirm birth years of authors and publication dates of books, we compared NovelTM metadata both to the Chicago Novel Corpus and to a copy of the US Copyright Registry, digitized by the New York Public Library (https://github.com/NYPL/catalog_of_copyright_entries_project).</p> <p>The corpus itself is in cohort4.txt.gz; each line represents a roughly 10,000-word "chunk" of a document. The first 15% and last 5% of pages in each volume were discarded; the remaining pages were divided into chunks of roughly equal size. Chunk id is the first token on each line; it is formed by taking a HathiTrust volume id and adding an underscore + sequential integer (chunk number). Removing the underscore and integer produces a "document id" that can be paired to the metadata. The words in the line are not presented in original order; they are taken from HathiTrust Extracted Features, which records only page-level word counts.</p> <p>The topic model was produced using MALLET (http://mallet.cs.umass.edu/index.php), and has 200 topics.</p> <p>The top words in each topic are listed in the "keys" file; document-topic proportions are listed in "doctopics."</p> <p>For more information on the construction of the corpus and the experiment it is designed to support, see https://github.com/tedunderwood/period-cohort and/or a permanent Zenodo object created from that repository.</p>
Data for: Can language representation models think in bets?
<p>The dataset contains three files: *Item_Sets.xlsx*, *Value_Questions.xlsx*, and *Bet_Questions.xlsx*. Each of the three files corresponds to each of the three benchmarks in the manuscript currently under submission to Royal Society Open Science and is also available as a preprint: <a href="https://arxiv.org/abs/2210.07519">https://arxiv.org/abs/2210.07519</a>.</p> <p>The items in Item_Sets.xlsx is used to create the other two files. Value_Questions.xlsx is used in RQ1, and Bet_Questions.xlsx is used in both RQ2 and RQ3.</p>
Species-aware DNA language modeling - data
<p>Data accompanying the publication Species-aware DNA language modeling.</p> <p>For code, see: https://github.com/DennisGankin/species-aware-DNA-LM (for the latest version) or the code.zip file.</p> <p>The data directory contains model checkpoints, baselines models, evaluation results and datasets used for training, testing and downstream tasks. It has the following structure:</p> <p>data/ Datasets and subdirectories</p> <p> - data/results/ Results from different test runs</p> <p> - data/models/ Model checkpoints</p> <p> - data/baselines/ Baseline models</p>
RESTful API Testing with the Power of Large Language Models
<p>Sample data for <a href="https://ase2023.hotcrp.com/u/1/paper/519">RESTful API Testing with the Power of Large Language Models</a></p>
Geographical modelling of language decline for Cornish and Welsh
<p>Competition between languages affects the lives of people all over the globe, and a huge number of languages are at risk of extinction. In this work, statistical physics is applied to modelling the decline of one language in competition with another. A model from the literature is used and adapted to model the interactions among speakers in a population distribution over time, and is applied to historical data for Cornish and Welsh. Visual, geographical models show the simulated decline of the real languages studied, and a number of qualitative and quantitative features from the historical data are captured by the model. The applicability of the model to further real situations is discussed, as well as adaptations that would be needed to better take account of migration and population dynamics.</p>
FIGURE 1 in Harnessing the power of AI language models for taxonomy and systematics: a follow-up to "Can ChatGPT be leveraged for taxonomic investigations? Potential and limitations of a new technology" by Davinack (2023)
FIGURE 1. Python code generated by ChatGPT that allows the preparation of a NEX file format. For Python tutorials, source code and installers, see https://www.python.org.
Datasets and Trained Models for "Unblind Your Apps: Predicting Natural-Language Labels for Mobile GUI Components by Deep Learning"
<p>Datasets and Trained models for ICSE 2020 "Unblind Your Apps: Predicting Natural-Language Labels for Mobile GUI Components by Deep Learning"</p>
Replication Package for "An Exploratory Literature Study on Sharing and Energy Use of Language Models for Source Code"
<p>This repository contains the replication package for the paper <em>"</em>An Exploratory Literature Study on Sharing and Energy Use of Language Models for Source Code" by Max Hort, Anastasiia Grishina, and Leon Moonen, accepted for publication in the 17th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM 2023).</p> <p>The paper is deposited on arXiv, will be available later at the publisher's site (<a href="https://ieeexplore.ieee.org/Xplore/home.jsp">IEEE</a>), and a copy is included in this repository.</p> <p>The replication package is archived on Zenodo with DOI: <a href="https://doi.org/10.5281/zenodo.8058667">10.5281/zenodo.8058667</a>. The data is distributed under the CC BY 4.0 license.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you build on this data or code, please cite this work by referring to the paper:</p> <pre><code>@inproceedings{hort2023:sharing, title = {An Exploratory Literature Study on Sharing and Energy Use of Language Models for Source Code}, author = {Max Hort and Anastasiia Grishina and Leon Moonen}, booktitle = {17th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM 2023)}, year = {2023}, publisher = {IEEE} note = {To appear. Pre-print on arXiv.} }</code></pre>
No More In-Context Learning? Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
<p>Data and models part of the replication package of the ICSE 24 submission entitled "<em>No More In-Context Learning? Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models</em>".</p>
Anonymized data for paper "Automatic Bug Fixing in the Era of Large Language Models: Interactive Simulation of Programmer Behavior" submitted to ICSE 2024
<p>The project includes the BFP benchmark used in the submitted ICSE 2024 paper titled "Automatic Bug Fixing in the Era of Large Language Models: Interactive Simulation of Programmer Behavior"</p>
Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.
<p>List of 200 most abundant VOG descriptions.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.