Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “Pre-trained language model”
Ecore Metamodels and EcoreBERT Pre-trained Language Model
<p>This dataset contains ecore metamodels from the MAR dataset transformed into tree representations. The original dataset can be found here: <a href="http://mar-search.org/experiments/models20/">http://mar-search.org/experiments/models20/</a></p> <p>The data contained in this repository were used to conduct the experiments in the paper: <strong>Recommending Metamodel Concepts during Modeling Activities with Pre-Trained Language Models. </strong>Link to the paper: <a href="https://arxiv.org/abs/2104.01642">https://arxiv.org/abs/2104.01642</a></p> <p>The data are organized as follows:</p> <ul> <li>model : our model trained on the tree representations of metamodels with RoBERTa architecture.</li> <li>tokenizers : the byte-level BPE tokenizer we used to train our model.</li> <li>train : the training data separated into a training and validation set.</li> <li>test : the test data of all experiments conducted in the paper.</li> </ul> <p>This data repository is linked with the following Github repository containing our code: <a href="https://github.com/mweyssow/ecore-bert">https://github.com/martiwey/metamodel-concepts-bert</a></p>
BioVAE: a pre-trained latent variable language model for biomedical text mining
<p>We release BioVAE, the first large-scale pre-trained latent variable language model for the biomedical domain, which uses the OPTIMUS framework to train on large volumes of biomedical text.</p> <p>This version contains the pre-trained models for text mining tasks such as named entity recognition or relation extraction, and text generation task.</p> <p>Explanation of each file: (lt32: latent_size = 32, beta05: beta=0.5)</p> <ul> <li>pm-full-lt32-beta00</li> <li>pm-full-lt32-beta05</li> <li>pm-full-lt768-beta00</li> <li>pm-full-lt768-beta05</li> <li>pm-full-generation</li> </ul>
Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair
<p>This page serves as supplementary material for the article: <strong>Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair</strong>. Here, we provide the ARJACLM code utilized in the study, enabling other researchers to replicate the experiments and further develop the tool. </p>
Automated Program Repair in the Era of Large Pre-trained Language Models
<p>Code used for the paper along with the generated outputs</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.