Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
296
datasets available to search
ShareScore release 0.9.0
Dataset results
296 results for “language models”
Molecules used to train or generated by chemical language models
<p>This upload contains training datasets or generated molecules from the paper “Invalid SMILES are helpful, not harmful, for chemical language models.”</p> <p>The contents of the directories are as follows:</p> <ul> <li>training_sets: sets of molecules from ChEMBL or GDB-13 used to train chemical language models, represented either as SMILES or SELFIES</li> <li>sampled-*: unprocessed samples of 10 million molecules from each model trained on ChEMBL or GDB-13</li> <li>prior_inputs: sets of molecules from LOTUS, COCONUT, FooDB and NORMAN, split into ten folds and used to train chemical language models</li> <li>priors-*: samples of 100 million molecules from chemical language models trained on each cross-validation fold, with unique molecules represented as canonical SMILES and sorted in descending order by their sampling frequency</li> </ul>
[Replication Package] Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
<p>This repository contains scripts, datasets, and results of the work <em>"Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation"</em></p> <p><strong>Scripts contained in this Zenodo repository can also be visualized at the following link: <a href="https://anonymous.4open.science/r/lowbit-quantization-D070/README.md">https://anonymous.4open.science/r/lowbit-quantization-D070/README.md</a><br></strong></p>
Physician Reasoning on Diagnostic Cases With Large Language Models
ClinicalTrials.gov study NCT06157944. IPD Sharing: NO. Countries: 1. Publications: 1.
The Effects of a Large Language Model on Clinical Questioning Skills
ClinicalTrials.gov study NCT06229379. IPD Sharing: NO. Countries: 1. Publications: 1.
Evaluating the Potential of Large Language Models for Respiratory Disease Consultations
ClinicalTrials.gov study NCT06457269. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Duet 2.0 Starting the Conversation: A New Intervention Model to Stimulate Language Growth in Underserved Populations
ClinicalTrials.gov study NCT04692519. IPD Sharing: NO. Countries: 1. Publications: 23.
Effectiveness of Large Language Model for Anaesthesia and Procedural Consent
ClinicalTrials.gov study NCT06949462. IPD Sharing: YES. Countries: 1. Publications: 1.
Nutritional Language Model
ClinicalTrials.gov study NCT06661590. IPD Sharing: NO. Countries: 1. Publications: 1.
The Application of Large Language Model in Emergency Chest Pain Triage
ClinicalTrials.gov study NCT06493175. IPD Sharing: NO. Countries: 1. Publications: 0.
Enhancing Medical Researchers' Self-learning With an Intelligent Language Model
ClinicalTrials.gov study NCT06015178. IPD Sharing: NO. Countries: 1. Publications: 1.
Application of Large Language Models in Emergency Neurology
ClinicalTrials.gov study NCT06779292. IPD Sharing: NO. Countries: 1. Publications: 0.
Large Language Model-Generated Messages to Improve Guideline-Directed Medical Therapy in Heart Failure
ClinicalTrials.gov study NCT07337577. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.
Physician Reasoning on Management Cases With Large Language Models
ClinicalTrials.gov study NCT06208423. IPD Sharing: NO. Countries: 1. Publications: 2.
Multi-Disciplinary Treatment on the Anthropomorphism of Large Language Models
ClinicalTrials.gov study NCT06627985. IPD Sharing: YES. Countries: 1. Publications: 1.
Application of Multimodal Large Language Model in HFpEF
ClinicalTrials.gov study NCT06486649. IPD Sharing: NO. Countries: 1. Publications: 9.
Data for: Can language representation models think in bets?
Open the record for dataset details and reuse information.
Geographical modelling of language decline for Cornish and Welsh
Open the record for dataset details and reuse information.
The results of model learning on base an annotated text, compiled on the basis of the English-language news feed of the Yuri Gagarin State Technical University of Saratov
<p>The results of model learning on base an annotated text, compiled from the English-language news feed of the Yuri Gagarin State Technical University of Saratov.</p> <p>This file can be used in conjunction with Data for Model Learning on base OPENNLP DOI 10.5281/zenodo.3550016</p> <p> </p> <p> </p>
Dataset from the EMNLP 2020 article "Modeling the Music Genre Perception across Language-Bound Cultures"
<p>We release the data required to reproduce the experiments from the article <em>Modeling the Music Genre Perception across Language-Bound Cultures</em> presented at the <a href="https://2020.emnlp.org">EMNLP 2020</a> conference.</p> <p>More information about this data and how it should be used in the experiments can be found in the GitHub repository <a href="https://github.com/deezer/CrossCulturalMusicGenrePerception">deezer/CrossCulturalMusicGenrePerception</a>.</p> <p>Please cite our paper if you use the code or data in your work.</p>
A Korean raw text collection for creating a language model
<p><strong>A very large Korean raw text collection for creating a language model </strong></p> <p> </p> <p>We collected a very large monolingual dataset for Korean, which contains <strong>over 9.6M sentences and 130.6M eojeols</strong>, to create a language model: Korean Wikipedida (https://dumps.wikimedia.org/kowiki/20201101/, 5.3M sentences and 71.8M eojeols, respectively), the Sejong morphologically analyzed corpus (3.0M and 40.0M), and articles from <em>The Hankyoreh</em> daily newspaper during 2016 (1.2M and 18.6M). </p> <p> </p> <p>We preprocessed raw text into morpheme-segmented text using the POS tagging system (<a href="https://www.aclweb.org/anthology/W19-4022/">park-tyers:2019:LAW</a>). We also attached the POS label to the morpheme-segmented lexicon, and explicitly include a + symbol for consecutive morphemes. </p> <blockquote> <p>시인/NNG 윤동주/NNP +,/SP 이준익/NNP 감독/NNG 영화/NNG +로/JKB 부활/NNG</p> </blockquote> <p> </p> <p>See https://github.com/jungyeul/sjmorph for the POS tagging system described in <a href="https://www.aclweb.org/anthology/W19-4022/">park-tyers:2019:LAW</a>. </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.