Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
296
datasets available to search
ShareScore release 0.9.0
Dataset results
296 results for “language models”
Investigating Gender Bias in Large Language Models Through Text Generation
<p>Presented at the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024). Published in the ACL Anthology.</p>
Dataset of the Automatic Unit Test Generation for Programming Assignments Using Large Language Models
Open the record for dataset details and reuse information.
Dataset for "The Politics of AI – An Evaluation of Political Preferences in Large Language Models from a European Perspective"
Open the record for dataset details and reuse information.
Advancing Automated Code Review Comment Generation Using Large Language Models
Open the record for dataset details and reuse information.
Automated Program Repair in the Era of Large Pre-trained Language Models
<p>Code used for the paper along with the generated outputs</p>
Enhancing diversity in language based models for single-step retrosynthesis
<p>Dataset for publication: <a href="https://doi.org/10.1039/D2DD00110A">https://doi.org/10.1039/D2DD00110A</a></p>
Data for "No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models"
<p>Datasets to reproduce the experiments associated with the paper: https://doi.org/10.48550/arXiv.2307.06440</p> <p>The readme contains instructions for how to use them: https://github.com/JeanKaddour/NoTrainNoGain/blob/main/bert/README.md</p> <p>c4-subset-random.tar.bz2 is a subset of the C4 dataset (https://arxiv.org/abs/1910.10683), licensed under ODC-BY 1.0.</p>
Deepurify: a multi-modal deep language model to remove contamination from metagenome-assembled genomes
<p>The SIM1 testing set.</p>
The Plastic Surgery Hypothesis in the Era of Large Language Models
<p>Correct Patches generated by FitRepair and Source code.</p>
Evaluating Accuracy Improvements of Large Language Models Using Retrieval Augmented Generation
<p>Original data</p>
Model checkpoints for "XFEVER: Exploring Fact Verification across Languages"
<p>This is the collection of model checkpoints (as well as example outputs and data) for reproducing the experiments in "XFEVER: Exploring Fact Verification across Languages". Our code is available at: <a href="https://github.com/nii-yamagishilab/xfever">https://github.com/nii-yamagishilab/xfever</a>.</p>
Data and code for, "Large language models design sequence-defined macromolecules via evolutionary optimization"
<div> <pre># Codes and data for "Large language models design sequence-defined macromolecules via evolutionary optimization"<br><br>Note this repository contains codes and data files for the manuscript. This is a snapshot of the repository, frozen at the time of submission.<br><br># Codes<br><br>## LLM codes<br>- `run_claude.py` - the routine for performing LLM-based rollouts; intended for command line execution using argparse<br>- `message_utils.py` - utilities for constructing and parsing messages for LLM I/O<br>- `model_utils.py` - lightweight utilities for retrieving formatted predictions from the RNN ensemble<br>- `target_defs.py` - defines the sequence, locations, and natural language descriptions of the target structures<br>- `ask_about_oracle.ipynb` - asks the LLM to speculate about the nature of the optimization task<br><br>## other algorithms<br>- `active_learning.ipynb` - use EI acquisition with RF surrogate to label new sequences; includes an unused tokenization scheme<br>- `evolutionary_algorithm.ipynb` - use DEAP library to perform evolutionary optimization<br>- `random_sampling.ipynb` - sample sequences randomly from all possible sequences<br><br>## postprocessing<br>- `process_aggregated_logs.py` - reads data from the raw log files and prepares them for visualization<br>- `process_sample_rollouts.py` - reads data from the raw log files and prepares individual rollouts<br><br>## visualization<br>- `figure1b.ipynb` - renders panel b of Fig. 1<br>- `figure1efg.ipynb` - renders the last row of Fig. 1 (panels e-g)<br>- `figure2.ipynb` - renders all of Fig. 2<br>- `figure_si.ipynb` - renders Figs. S1 and S2<br>- `figure_md_validation.ipynb` - renders Fig. S3<br><br># Data files<br><br>- `prompts/`<br> - `prompt-scientific-v4.4.yml` - the full text of the scientific prompt, to be read by `run_claude.py`<br> - `prompt-oracle-v4.4.yml` - the full text of the oracle prompt, to be read by `run_claude.py`<br>- `models/` - the TorchScript RNN models used to make predictions<br>- `data/`<br> - `embeddings` - calculated embeddings for a collection of sequences from our prior work<br> - `llm-logs` - the raw logs obtained from the Claude 3.5 Sonnet LLM (other algorithms made to look like the LLM logs after the fact)<br> - `llm-logs-opus` - the raw logs obtained from the Claude 3.0 Opus LLM (used in the first draft of the article, replaced by Claude 3.5 Sonnet) <br> - `all-rollouts-kltd.csv` - postprocessed logs for all the rollouts using the "top $k < d^*$" metric<br> - `all-rollouts-topkd.csv` - postprocessed logs for all the rollouts using the "mean $d$ for top $k$" metric<br> - `sample-rollout-membranes-x-3.csv` - postprocessed logs for a single rollout replica, `x` = each algorithm type<br> - `snapshots` - png snapshots of MD simulation results at different locations in the manifold</pre> </div>
Can Feedback From a Large Language Model Improve Health Care Quality?
ClinicalTrials.gov study NCT06823765. IPD Sharing: YES. Countries: 1. Publications: 0.
Enhancing Interdisciplinary Understanding of Ophthalmology Notes Through a Local Large Language Model
ClinicalTrials.gov study NCT06624605. IPD Sharing: YES. Countries: 1. Publications: 0.
Data from: Markovian language model of the DNA and its information content
Open the record for dataset details and reuse information.
Data from: Natural language processing and recurrent network models for identifying genomic mutation-associated cancer treatment change from patient progress notes
Open the record for dataset details and reuse information.
A Genomic Language Model for Chimera Artifact Detection in Nanopore Direct RNA Sequencing
GEO Series GSE277934. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.
Data for Investigating the Technical Debt in Procedural Model Transformation Languages
<p>The content presents the data for Investigating the Technical Debt in Procedural Model Transformation Languages </p>
Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models
<p>Data and checkpoints for 'Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models'</p>
Application of large language models in radiology: evaluating longitudinal diagnostic yield of imaging for abdominal pain.
<p>Annotated dataset belonging to the paper "Application of large language models in radiology: evaluating longitudinal diagnostic yield of imaging for abdominal pain".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.