Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

296

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

296 results for “language models”

Learn how ShareScore rates datasets ↗
zenodo32/100

Molecules used to train or generated by chemical language models

<p>This upload contains training datasets or generated molecules&nbsp;from the paper &ldquo;Invalid SMILES are helpful, not harmful, for chemical language models.&rdquo;</p> <p>The contents of the directories are as follows:</p> <ul> <li>training_sets: sets of molecules from ChEMBL or GDB-13 used to train chemical language models,&nbsp;represented either as SMILES or SELFIES</li> <li>sampled-*: unprocessed samples of 10 million molecules from each model&nbsp;trained on ChEMBL or GDB-13</li> <li>prior_inputs: sets of molecules from LOTUS, COCONUT, FooDB and NORMAN, split into ten folds and used to train chemical language models</li> <li>priors-*: samples of 100 million molecules from chemical language models trained on each cross-validation fold, with unique molecules represented as canonical SMILES and sorted in descending order by their sampling frequency</li> </ul>

opencc-by-4.0Sep 2023View details →
zenodo32/100

[Replication Package] Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation

<p>This repository contains scripts, datasets, and results of the work <em>"Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation"</em></p> <p><strong>Scripts contained in this Zenodo repository can also be visualized at the following link: <a href="https://anonymous.4open.science/r/lowbit-quantization-D070/README.md">https://anonymous.4open.science/r/lowbit-quantization-D070/README.md</a><br></strong></p>

opencc-by-4.0Sep 2024View details →
ClinicalTrials.gov32/100

Physician Reasoning on Diagnostic Cases With Large Language Models

ClinicalTrials.gov study NCT06157944. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

The Effects of a Large Language Model on Clinical Questioning Skills

ClinicalTrials.gov study NCT06229379. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Evaluating the Potential of Large Language Models for Respiratory Disease Consultations

ClinicalTrials.gov study NCT06457269. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Duet 2.0 Starting the Conversation: A New Intervention Model to Stimulate Language Growth in Underserved Populations

ClinicalTrials.gov study NCT04692519. IPD Sharing: NO. Countries: 1. Publications: 23.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Effectiveness of Large Language Model for Anaesthesia and Procedural Consent

ClinicalTrials.gov study NCT06949462. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov32/100

Nutritional Language Model

ClinicalTrials.gov study NCT06661590. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

The Application of Large Language Model in Emergency Chest Pain Triage

ClinicalTrials.gov study NCT06493175. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Enhancing Medical Researchers' Self-learning With an Intelligent Language Model

ClinicalTrials.gov study NCT06015178. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Application of Large Language Models in Emergency Neurology

ClinicalTrials.gov study NCT06779292. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Large Language Model-Generated Messages to Improve Guideline-Directed Medical Therapy in Heart Failure

ClinicalTrials.gov study NCT07337577. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Physician Reasoning on Management Cases With Large Language Models

ClinicalTrials.gov study NCT06208423. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Multi-Disciplinary Treatment on the Anthropomorphism of Large Language Models

ClinicalTrials.gov study NCT06627985. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov32/100

Application of Multimodal Large Language Model in HFpEF

ClinicalTrials.gov study NCT06486649. IPD Sharing: NO. Countries: 1. Publications: 9.

closedIPD-NOFeb 2026View details →
dryad32/100

Data for: Can language representation models think in bets?

Open the record for dataset details and reuse information.

publicDec 2022View details →
dryad32/100

Geographical modelling of language decline for Cornish and Welsh

Open the record for dataset details and reuse information.

publicMay 2023View details →
zenodo28/100

The results of model learning on base an annotated text, compiled on the basis of the English-language news feed of the Yuri Gagarin State Technical University of Saratov

<p>The results of model learning on base an annotated text, compiled from&nbsp;the English-language news feed of the Yuri Gagarin State Technical University of Saratov.</p> <p>This file can be used in conjunction with&nbsp;Data for Model Learning on base OPENNLP&nbsp;DOI&nbsp;10.5281/zenodo.3550016</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo28/100

Dataset from the EMNLP 2020 article "Modeling the Music Genre Perception across Language-Bound Cultures"

<p>We release the data required to reproduce the experiments from the article&nbsp;<em>Modeling the Music Genre Perception across Language-Bound Cultures</em>&nbsp;presented at the&nbsp;<a href="https://2020.emnlp.org">EMNLP 2020</a>&nbsp;conference.</p> <p>More information about this&nbsp;data&nbsp;and how it should be used in the experiments can be found&nbsp;in the GitHub repository&nbsp;<a href="https://github.com/deezer/CrossCulturalMusicGenrePerception">deezer/CrossCulturalMusicGenrePerception</a>.</p> <p>Please cite our paper if you use the code or data in your work.</p>

opencc-by-4.0Nov 2020View details →
zenodo28/100

A Korean raw text collection for creating a language model

<p><strong>A very large Korean raw text collection for creating a language model&nbsp;</strong></p> <p>&nbsp;</p> <p>We collected&nbsp;a very large monolingual dataset for Korean, which contains <strong>over 9.6M sentences and 130.6M eojeols</strong>, to create a language model: Korean Wikipedida (https://dumps.wikimedia.org/kowiki/20201101/, 5.3M sentences and 71.8M eojeols, respectively), the Sejong morphologically analyzed corpus (3.0M and 40.0M), and articles from <em>The Hankyoreh</em>&nbsp;daily newspaper during 2016 (1.2M and 18.6M).&nbsp;</p> <p>&nbsp;</p> <p>We preprocessed&nbsp;raw text into morpheme-segmented text &nbsp;using the POS tagging system (<a href="https://www.aclweb.org/anthology/W19-4022/">park-tyers:2019:LAW</a>).&nbsp;We also attached the POS label to the morpheme-segmented lexicon, and explicitly include a + symbol for consecutive morphemes.&nbsp;</p> <blockquote> <p>시인/NNG 윤동주/NNP +,/SP 이준익/NNP 감독/NNG 영화/NNG +로/JKB 부활/NNG</p> </blockquote> <p>&nbsp;</p> <p>See&nbsp;https://github.com/jungyeul/sjmorph for the&nbsp;POS tagging system described in&nbsp;<a href="https://www.aclweb.org/anthology/W19-4022/">park-tyers:2019:LAW</a>.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record