Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
296
datasets available to search
ShareScore release 0.9.0
Dataset results
296 results for “language models”
Data and code on the Moral Machine experiment on large language models (LLMs)
Open the record for dataset details and reuse information.
Data from: Protein Set Transformer: A protein-based genome language model to power high diversity viromics
Open the record for dataset details and reuse information.
RadCases evaluation results: Evaluating acute image ordering for real-world patient cases via language model alignment with radiological guidelines
Open the record for dataset details and reuse information.
Data and code from: Learning a deep language model for microbiomes: The power of large scale unlabeled microbiome data
Open the record for dataset details and reuse information.
Evaluation of large language model chatbot responses to psychotic prompts: numerical ratings of prompt-response pairs
Open the record for dataset details and reuse information.
Wbbyyr: FastText language models for Mandarin Chinese, trained on 14m Sina Weibo posts for each year in 2012-2018 (Fold 1 of 10)
<p>Wbbyyr: FastText language models for Mandarin Chinese, trained on 14,440,000 Sina Weibo posts for each year in 2012-2018.</p> <p>The 14,440,000 posts from each year are split into 10 folds. Due to Zenodo size limit, this dataset contains only the first fold from each year.</p> <p>Each model is trained for 20 iterations. Each vector is 300 dimensions long.</p>
SemEval-2020 Task 5: Modelling Causal Reasoning in Language: Detecting Counterfactuals
<p><strong>SemEval-2020 Task 5</strong></p> <p> </p> <p><strong>Subtask-1:</strong> Recognizing Counterfactual Statements (RCS) -- Determine whether a given sentence is counterfactual or not.</p> <p><strong>Subtask-2: </strong>Detecting Antecedent and Consequent (DAC) -- Extract the antecedent and consequent part in a given counterfactual sentence.</p> <p> </p> <p>The released dataset consists of train/test data of both subtask-1 and subtask-2. In our competition, participants could only use the corresponding dataset in each subtask.</p> <p> </p> <p><strong>Task 5 Codalab Website:</strong> <a href="https://competitions.codalab.org/competitions/21691">https://competitions.codalab.org/competitions/21691</a></p>
Language models
<p>English and Spanish language models for the NLP Assistant</p>
Representations of language in a model of visually grounded speech signal: Data
<p>The set of datafiles to reproduce results from:</p> <ul> <li>Chrupała, G., Gelderloos, L., & Alishahi, A. (2017). Representations of language in a model of visually grounded speech signal. ACL. arXiv preprint: https://arxiv.org/abs/1702.01991</li> </ul>
pLMMoRF: A web server that accurately predicts membrane-interacting molecular recognition features by employing a protein language model
<p>pLMMMoRF predictor scrips and MemMoRF prediction of the human proteome.</p>
Gene-language models are whole genome representation learners
<p>The language of genetic code embodies a complex grammar and rich syntax of interacting molecular elements. Recent advances in self-supervision and feature learning suggest that statistical learning techniques can identify high-quality quantitative representations from inherent semantic structure. We present a gene-based language model that generates whole-genome vector representations from a population of 16 disease-causing bacterial species by leveraging natural contrastive characteristics between individuals. To achieve this, we developed a set-based learning objective, AB learning, that compares the annotated gene content of two population subsets for use in optimization. Using this foundational objective, we trained a Transformer model to backpropagate information into dense genome vector representations. The resulting bacterial representations, or embeddings, captured important population structure characteristics, like delineations across serotypes and host specificity preferences. Their vector quantities encoded the relevant functional information necessary to achieve state-of-the-art genomic supervised prediction accuracy in 11 out of 12 antibiotic resistance phenotypes.</p>
Data from: Genome-scale annotation of protein binding sites via language model and geometric deep learning
<p>The dataset contains the training and test sets of protein binding sites with DNA, RNA, peptide, protein, ATP, HEM, Zn2+, Ca2+, Mg2+ and Mn2+. Each protein is associated with 3 lines indicating the protein name (PDB accession code and chain), sequence and residue labels (0 for non-binding and 1 for binding), respectively. The ESMFold-predicted structures are also provided.</p>
Replication Package for "Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis"
<h1>Replication Package for the Paper: “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis”</h1> <p>This replication package includes the raw data, questionnaire answers, and a Python notebook needed for reproducing the results detailed in the paper titled “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis.”</p> <h2><a></a>Repository Structure</h2> <ol> <li><strong>Scenarios:</strong> Contains an Excel file encompassing all 141 scenarios collected (in Italian).</li> <li><strong>Training and Validation Messages:</strong> Includes the jsonl files necessary for fine-tuning the model.</li> <li><strong>Testing Messages and Ground Truth:</strong> Contains the messages utilized for testing the models.</li> <li><strong>Results:</strong> Contains Excel files with the responses from the 2 human experts and the 5 model as well as the review of the 3 human reviewer.</li> <li><strong>Tables:</strong> Contains the full Wilcoxon Test Results for H01 and H02 as well as the raw RQs results.</li> </ol> <h2><a></a>Replication Process</h2> <p>To replicate the results of our study, open the provided Python Notebook in Google Colab and follow the instructions to seamlessly reproduce the results.</p> <h1><a></a>Instructions for Use</h1> <p>To utilize this replicability package, refer to the steps outlined in the notebook file.</p> <h1><a></a>Remarks</h1> <p>If you encounter any issues or have any questions, please reach out to the authors of the paper. We will be glad to assist you!</p>
PRICER: Leveraging Few-Shot Learning with Fine-Tuned Large Language Models for Unstructured Economic Data
<p>Describes the taxonomy used in the paper "PRICER: Leveraging Few-Shot Learning with Fine-Tuned Large Language Models for Unstructured Economic Data", presented at the Second Workshop on Semantic Technologies and Deep Learning Models for Scientific, Technical and Legal Data<em> </em>at the Extended Semantic Web Conference (ESWC) 2024.</p>
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
<p>We introduce a new image-text dataset, providing high-quality natural language descriptions for global-scale satellite data. Specifically, we utilize Sentinel-2 data for its global coverage as the foundational image source, employing semantic segmentation labels from the European Space Agency's WorldCover project to enrich the descriptions of land covers. By conducting in-depth semantic analysis, we formulate detailed prompts to elicit rich descriptions from ChatGPT. We then include a manual verification process to enhance the dataset's quality further. This step involves manual inspection and correction to refine the dataset. Finally, we offer the community ChatEarthNet, a large-scale image-text dataset characterized by global coverage, high quality, wide-ranging diversity, and detailed descriptions. ChatEarthNet consists of 163,488 image-text pairs with captions generated by ChatGPT-3.5 and an additional 10,000 image-text pairs with captions generated by ChatGPT-4V(ision). This dataset has significant potential for both training and evaluating vision-language geo-foundation models for remote sensing.</p>
Knowledge Discovery from Porous Organic Cages Literature Using a Large Language Model
<p>This article presents a GPT-4-based literature reading method that incorporates multi-label text classification and a follow-up information extraction, in which the potential of GPT-4 can be fully exploited to rapidly extract valid information from the literature. In the process of multi-label text classification, the prompt-engineered GPT-4 demonstrated the ability to label text with proper recall rates according to the type of information contained in text, including authors, affiliations, synthetic procedures, surface area, and the CCDC number of corresponding cages. Additionally, GPT-4 demonstrated proficiency in information extraction, effectively transforming labeled text into concise tabulated data. Furthermore, we built a chatbot based on this database, allowing for quick and comprehensive searching across the entire database and responding for cage-related questions.</p>
Dataset for "Pop Quiz! Can a Large Language Model Help WIth Reverse Engineering?"
<p>Source files and scripts for the test framework used in evaluating OpenAI's code-davinci-001 (also known as davinci-codex) model for reverse engineering tasks. All results presented in the associated manuscript are preserved in this dataset, or you can use it to perform additional experiments or re-generate the results.</p>
Flakify: A Black-Box, Language Model-based Predictor for Flaky Tests – Replication Package
<p>This is the replication package associated with the paper: <em>Flakify: A Black-Box, Language Model-based Predictor for Flaky Tests.</em> We explain how to use it to reproduce the results reported in the paper. A maintainable version of this replication package is available on GitHub (<a href="https://github.com/uOttawa-Nanda-Lab/Flakify">https://github.com/uOttawa-Nanda-Lab/Flakify</a>).</p> <p><strong>Flakify Test Smell Detector</strong></p> <p>This is a step-by-step guideline to detect test smells in the source code of test cases and retain statements that match them.</p> <p><em><strong>Requirements:</strong></em></p> <ul> <li>Eclipse IDE (the version we used was 2021-12)</li> <li>The libraries (the <strong><em>.jar</em></strong> files in the <strong><code>lib\</code></strong> directory)</li> </ul> <p><em><strong>Input Files:</strong></em></p> <p>This is a list of input files that are required to accomplish this step:</p> <ul> <li> <p><em>dataset/FlakeFlagger/FlakeFlagger_filtered_dataset.csv</em></p> </li> <li> <p><em>dataset/FlakeFlagger/FlakeFlagger_class_files/</em></p> </li> <li> <p><em>dataset/IDoFT/IDoFT_filtered_dataset.csv</em></p> </li> <li> <p><em>dataset/IDoFT/IDoFT_class_files/</em></p> </li> </ul> <p>The <strong><code>dataset/FlakeFlagger/FlakeFlagger_filtered_dataset.csv</code></strong> and <strong><code>dataset/IDoFT/IDoFT_filtered_dataset.csv</code></strong> are used to obtain the label (<em>flaky</em>=1 or <em>non-flaky</em>=0) and project name for each test case parsed from <strong><code>dataset/FlakeFlagger/FlakeFlagger_class_files/</code></strong> and <strong><code>dataset/IDoFT/IDoFT_class_files/</code></strong>, respectively.</p> <p><strong><em>Output Files:</em></strong></p> <ul> <li> <p><em>dataset/FlakeFlagger/FlakeFlagger_dataset.csv</em></p> </li> <li> <p><em>dataset/FlakeFlagger/FlakeFlagger_test_cases_full_code/</em></p> </li> <li> <p><em>dataset/FlakeFlagger/FlakeFlagger_test_cases_preprocessed_code/</em></p> </li> <li> <p><em>dataset/IDoFT/IDoFT_dataset.csv</em></p> </li> <li> <p><em>dataset/IDoFT/IDoFT_test_cases_full_code/</em></p> </li> <li> <p><em>dataset/IDoFT/IDoFT_test_cases_preprocessed_code/</em></p> </li> </ul> <p> </p> <p><strong>Replicating the experiment</strong></p> <p>To detect test smells and retain only code statements related to them, the <strong><code>src/FlakifySmellsDetector.java</code></strong> file should be compiled and run using the Eclipse IDE by having all the <em>.jar</em> files in the classpath.</p> <p>The pre-generated executable Jar file <strong><code>src/FlakifySmellsDetector.jar</code></strong> can be executed using the shell script <strong><code>src/FlakifySmellsDetector.sh</code></strong> after changing paths for each dataset as needed, using the following commands:</p> <pre><code class="language-bash">bash FlakifySmellsDetector.sh FlakeFlagger bash FlakifySmellsDetector.sh IDoFT</code></pre> <p>It will generate the dataset required to run Flakify's flaky test prediction model for the datasets given as input. The class file containing each of the test cases is then parsed to produce the corresponding full code and pre-processed code of the test case. The full and pre-processed source code of all test cases are also combined and saved in a CSV file, along with test smells found, project names, and labels.</p> <p> </p> <p><strong>Flakify Replication</strong></p> <p>This is the guideline for replicating the experiments we used to evaluate Flakify for classifying test cases as <em>flaky</em> and <em>non-flaky</em> using both cross-validation and per-project validation.</p> <p><em><strong>Requirements:</strong></em></p> <p>This is a list of all required python packages:</p> <ul> <li><em>python =3.8.5</em></li> <li><em>imbalanced_learn= 0.8.1</em></li> <li><em>numpy= 1.19.5</em></li> <li><em>pandas= 1.3.3</em></li> <li><em>transformer= 4.10.2</em></li> <li><em>torch=1.5.0</em></li> <li><em>scikit_learn= 0.22.1</em></li> </ul> <p><em><strong>Input Files:</strong></em></p> <p>This is a list of input files that are required to accomplish this step:</p> <ul> <li><em>dataset/FlakeFlagger/Flakify_FlakeFlagger_dataset.csv</em></li> <li><em>dataset/IDoFT/Flakify_IDoFT_dataset.csv</em></li> </ul> <p>This file contains the full code and pre-processed code of the test cases in both FlakeFlagger and IDOFT datasets, along with their ground truth labels (<em>flaky</em> and <em>non-flaky</em>).</p> <p><em><strong>Output File:</strong></em></p> <ul> <li> <p><em>results/Flakify_cross_validation_results_on_FlakeFlagger_dataset.csv</em></p> </li> <li> <p><em>results/Flakify_per_project_results_on_FlakeFlagger_dataset.csv</em></p> </li> <li> <p><em>results/Flakify_model_trained_on_FlakeFlagger_dataset.pt</em></p> </li> <li> <p><em>results/Flakify_cross_validation_results_on_IDoFT_dataset.csv</em></p> </li> <li> <p><em>results/Flakify_per_project_results_on_IDoFT_dataset.csv</em></p> </li> <li> <p><em>results/Flakify_model_trained_on_IDoFT_dataset.pt</em></p> </li> </ul> <p> </p> <p><strong>Replicating Flakify experiments</strong></p> <p><strong>Cross-Validation</strong></p> <p>To run the Flakify experiment using cross-validation on the two datasets, navigate to <code>src\</code> folder and run the following commands:</p> <pre><code class="language-bash">bash Flakify_predictor_cross_validation.sh FlakeFlagger bash Flakify_predictor_cross_validation.sh IDoFT</code></pre> <p>This will generate the classification results into <strong><code>results/Flakify_cross_validation_results_on_FlakeFlagger_dataset.csv</code></strong> and <strong><code>results/Flakify_cross_validation_results_on_IDoFT_dataset.csv</code></strong> for the cross-validation experiments on both datasets. It will also save the weights of the two models trained on the FlakeFlagger and IDoFT datasets into <strong><code>results/Flakify_model_trained_on_FlakeFlagger_dataset.pt</code></strong> and <code><strong>results/Flakify_model_trained_on_IDoFT_dataset.pt</strong></code>, respectively.</p> <p> </p> <p><strong>Per-project Validation</strong></p> <p>To run the Flakify experiment using per-project validation on the two datasets, navigate to <code>src\</code> folder and run the following commands:</p> <pre><code class="language-bash">bash Flakify_predictor_per_project.sh FlakeFlagger bash Flakify_predictor_per_project.sh IDoFT</code></pre> <p>This will generate the classification results into <strong><code>results/Flakify_per_project_results_on_FlakeFlagger_dataset.csv</code></strong> and <strong><code>results/Flakify_per_project_results_on_IDoFT_dataset.csv</code></strong> for the whole per-project validation experiments on both datasets.</p> <p> </p> <p><strong>FlakeFlagger Replication</strong></p> <p>This is the guideline for replicating the experiments we used to evaluate the two versions of FlakeFlagger, white-box and black-box, for classifying test cases as <em>flaky</em> and <em>non-flaky</em> using cross-validation on the FlakeFlagger dataset.</p> <p><em><strong>Requirements:</strong></em></p> <p>This is a list of all required python packages:</p> <ul> <li><em>python =3.8.5</em></li> <li><em>imbalanced_learn= 0.8.1</em></li> <li><em>pandas= 1.3.3</em></li> <li><em>scikit_learn= 0.22.1</em></li> </ul> <p><em><strong>Input File:</strong></em></p> <p>This is a list of input files that are required to accomplish this step:</p> <ul> <li><em>dataset/FlakeFlagger/FlakeFlagger_filtered_dataset.csv</em></li> <li><em>dataset/FlakeFlagger/FlakeFlaggerFeaturesTypes.csv</em></li> <li><em>dataset/FlakeFlagger/Information_gain_per_feature.csv</em></li> </ul> <p><em><strong>Output File:</strong></em></p> <ul> <li><em>results/FlakeFlagger_black-box_results.csv</em></li> <li><em>results/FlakeFlagger_white-box_results.csv</em></li> </ul> <p> </p> <p><strong>Replicating FlakeFlagger experiments</strong></p> <p>To run the FlakeFlagger experiments, navigate to <code>src\</code> folder and run the following command:</p> <pre><code class="language-bash">bash FlakeFlagger_predictor.sh white-box bash FlakeFlagger_predictor.sh black-box</code></pre> <p>This will generate the classification results into <strong><code>results/FlakeFlagger_white-box_results.csv</code> </strong>and <strong><code>results/FlakeFlagger_black-box_results.csv</code> </strong>for both white-box and black-box experiments, respectively.</p>
Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"
<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States. </p> <p>The set of 47 news media outlets analysed (listed in Figure 1 of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets’ ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript. </p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets’ online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles’ HTML raw data using outlet-specific XPath expressions. </p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000’s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000. Figure S 1 in the SI reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small. Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth. </p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely: </p> <p>Siebert/sentiment-roberta-large-english from https://huggingface.co/siebert/sentiment-roberta-large-english. This model is a fine-tuned checkpoint of <a href="https://huggingface.co/roberta-large">RoBERTa-large</a> (<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with sentiment-roberta-large-english </p> <p>DistilRoberta j-hartmann/emotion-english-distilroberta-base from https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of <a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with Ekman's 6 basic emotions, plus a neutral class. The model was trained on 6 diverse datasets. Please refer to the original author at https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the siebert/sentiment-roberta-large-english Transformer model. https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset. https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the j-hartmann/emotion-english-distilroberta-base Transformer model. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>
Data and model weights for a series of antibody language models
<p>This repo contains the sequence dataset for finetuning and model weights including pre-trained meta model and a series of finetuned antibody language models.</p> <p>The antibody language models were used in the paper "<em>Physics-driven structural docking and protein language models accelerate antibody screening and design for broad-spectrum antiviral therapy"</em> for antibody sequence embedding.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.