Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

296

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

296 results for “language models”

Learn how ShareScore rates datasets ↗
zenodo28/100

Investigating Gender Bias in Large Language Models Through Text Generation

<p>Presented at the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024). Published in the ACL Anthology.</p>

opencc-by-sa-4.0Sep 2024View details →
zenodo28/100

Dataset of the Automatic Unit Test Generation for Programming Assignments Using Large Language Models

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo28/100

Dataset for "The Politics of AI – An Evaluation of Political Preferences in Large Language Models from a European Perspective"

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

Advancing Automated Code Review Comment Generation Using Large Language Models

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Automated Program Repair in the Era of Large Pre-trained Language Models

<p>Code used for the paper along with the generated outputs</p>

opencc-by-4.0May 2023View details →
zenodo28/100

Enhancing diversity in language based models for single-step retrosynthesis

<p>Dataset for publication:&nbsp;<a href="https://doi.org/10.1039/D2DD00110A">https://doi.org/10.1039/D2DD00110A</a></p>

opencc-by-4.0Feb 2023View details →
zenodo28/100

Data for "No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models"

<p>Datasets to reproduce the experiments associated with the paper: https://doi.org/10.48550/arXiv.2307.06440</p> <p>The readme contains instructions for how to use them: https://github.com/JeanKaddour/NoTrainNoGain/blob/main/bert/README.md</p> <p>c4-subset-random.tar.bz2 is a subset of the C4 dataset (https://arxiv.org/abs/1910.10683), licensed under ODC-BY 1.0.</p>

openodc-byJul 2023View details →
zenodo28/100

Deepurify: a multi-modal deep language model to remove contamination from metagenome-assembled genomes

<p>The SIM1 testing set.</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

The Plastic Surgery Hypothesis in the Era of Large Language Models

<p>Correct Patches generated by FitRepair and Source code.</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

Evaluating Accuracy Improvements of Large Language Models Using Retrieval Augmented Generation

<p>Original data</p>

opencc-by-4.0Oct 2023View details →
zenodo28/100

Model checkpoints for "XFEVER: Exploring Fact Verification across Languages"

<p>This is the collection of model checkpoints (as well as example outputs and data) for reproducing the experiments in "XFEVER: Exploring Fact Verification across Languages". Our code is available at: <a href="https://github.com/nii-yamagishilab/xfever">https://github.com/nii-yamagishilab/xfever</a>.</p>

openOct 2023View details →
zenodo28/100

Data and code for, "Large language models design sequence-defined macromolecules via evolutionary optimization"

<div> <pre># Codes and data for "Large language models design sequence-defined macromolecules via evolutionary optimization"<br><br>Note this repository contains codes and data files for the manuscript. This is a snapshot of the repository, frozen at the time of submission.<br><br># Codes<br><br>## LLM codes<br>- `run_claude.py` - the routine for performing LLM-based rollouts; intended for command line execution using argparse<br>- `message_utils.py` - utilities for constructing and parsing messages for LLM I/O<br>- `model_utils.py` - lightweight utilities for retrieving formatted predictions from the RNN ensemble<br>- `target_defs.py` - defines the sequence, locations, and natural language descriptions of the target structures<br>- `ask_about_oracle.ipynb` - asks the LLM to speculate about the nature of the optimization task<br><br>## other algorithms<br>- `active_learning.ipynb` - use EI acquisition with RF surrogate to label new sequences; includes an unused tokenization scheme<br>- `evolutionary_algorithm.ipynb` - use DEAP library to perform evolutionary optimization<br>- `random_sampling.ipynb` - sample sequences randomly from all possible sequences<br><br>## postprocessing<br>- `process_aggregated_logs.py` - reads data from the raw log files and prepares them for visualization<br>- `process_sample_rollouts.py` - reads data from the raw log files and prepares individual rollouts<br><br>## visualization<br>- `figure1b.ipynb` - renders panel b of Fig. 1<br>- `figure1efg.ipynb` - renders the last row of Fig. 1 (panels e-g)<br>- `figure2.ipynb` - renders all of Fig. 2<br>- `figure_si.ipynb` - renders Figs. S1 and S2<br>- `figure_md_validation.ipynb` - renders Fig. S3<br><br># Data files<br><br>- `prompts/`<br> - `prompt-scientific-v4.4.yml` - the full text of the scientific prompt, to be read by `run_claude.py`<br> - `prompt-oracle-v4.4.yml` - the full text of the oracle prompt, to be read by `run_claude.py`<br>- `models/` - the TorchScript RNN models used to make predictions<br>- `data/`<br> - `embeddings` - calculated embeddings for a collection of sequences from our prior work<br> - `llm-logs` - the raw logs obtained from the Claude 3.5 Sonnet LLM (other algorithms made to look like the LLM logs after the fact)<br> - `llm-logs-opus` - the raw logs obtained from the Claude 3.0 Opus LLM (used in the first draft of the article, replaced by Claude 3.5 Sonnet) <br> - `all-rollouts-kltd.csv` - postprocessed logs for all the rollouts using the "top $k &lt; d^*$" metric<br> - `all-rollouts-topkd.csv` - postprocessed logs for all the rollouts using the "mean $d$ for top $k$" metric<br> - `sample-rollout-membranes-x-3.csv` - postprocessed logs for a single rollout replica, `x` = each algorithm type<br> - `snapshots` - png snapshots of MD simulation results at different locations in the manifold</pre> </div>

opencc-by-4.0Aug 2024View details →
ClinicalTrials.gov28/100

Can Feedback From a Large Language Model Improve Health Care Quality?

ClinicalTrials.gov study NCT06823765. IPD Sharing: YES. Countries: 1. Publications: 0.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov28/100

Enhancing Interdisciplinary Understanding of Ophthalmology Notes Through a Local Large Language Model

ClinicalTrials.gov study NCT06624605. IPD Sharing: YES. Countries: 1. Publications: 0.

controlledIPD-YESFeb 2026View details →
dryad28/100

Data from: Markovian language model of the DNA and its information content

Open the record for dataset details and reuse information.

publicNov 2015View details →
dryad28/100

Data from: Natural language processing and recurrent network models for identifying genomic mutation-associated cancer treatment change from patient progress notes

Open the record for dataset details and reuse information.

publicFeb 2019View details →
geo24/100

A Genomic Language Model for Chimera Artifact Detection in Nanopore Direct RNA Sequencing

GEO Series GSE277934. Homo sapiens. 5 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2024View details →
zenodo24/100

Data for Investigating the Technical Debt in Procedural Model Transformation Languages

<p>The content presents the data for Investigating the Technical Debt in Procedural Model Transformation Languages&nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo24/100

Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models

<p>Data and checkpoints for 'Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language Models'</p>

opencc-zeroFeb 2024View details →
zenodo24/100

Application of large language models in radiology: evaluating longitudinal diagnostic yield of imaging for abdominal pain.

<p>Annotated dataset belonging to the paper "Application of large language models in radiology: evaluating longitudinal diagnostic yield of imaging for abdominal pain".</p>

restrictedcc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record