Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

65

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

65 results for “question answering”

Learn how ShareScore rates datasets ↗
zenodo52/100

BioASQ-QA: A manually curated corpus for Biomedical Question Answering

<p>The BioASQ question answering (QA) benchmark dataset contains questions in English, along with golden standard (reference) answers and related material. The dataset has been designed to reflect real information needs of biomedical experts and is therefore more realistic and challenging than most existing datasets. Furthermore, unlike most previous QA benchmarks that contain only exact answers, the BioASQ-QA dataset also includes ideal answers (in effect summaries), which are particularly useful for research on multi-document summarization. The dataset combines structured and unstructured data. The material linked with each question comprise documents and snippets, which are useful for Information Retrieval and Passage Retrieval experiments, as well as concepts that are useful in concept-to-text Natural Language Generation. Researchers working on paraphrasing and textual entailment can also measure the degree to which their methods improve the performance of biomedical QA systems. Last but not least, the dataset is continuously extended, as the BioASQ challenge is running and new data are generated.</p>

opencc-by-2.5Dec 2022View details →
zenodo48/100

Event-QA: A Dataset for Event-Centric Question Answering over Knowledge Graphs

<p>Event-QA dataset contains 1000 semantic queries and the corresponding verbalisations&nbsp;for EventKG -&nbsp;a recently proposed event-centric knowledge graph containing over 970 thousand events.</p>

opencc-by-4.0Apr 2019View details →
zenodo48/100

Toloka Visual Question Answering Dataset

<p>Our dataset consists of the images associated with textual questions. One entry (instance) in our dataset is a question-image pair labeled with the ground truth coordinates of a bounding box containing the visual answer to the given question. The images were obtained from a CC BY-licensed subset of the Microsoft Common Objects in Context dataset,&nbsp;<a href="https://cocodataset.org/">MS COCO</a>. All data labeling was performed on the Toloka crowdsourcing platform,&nbsp;<a href="https://toloka.ai/">https://toloka.ai/</a>.</p> <p>Our dataset has 45,199 instances split among three subsets:&nbsp;<strong>train</strong>&nbsp;(38,990 instances),&nbsp;<strong>public test</strong>&nbsp;(1,705 instances), and&nbsp;<strong>private test</strong>&nbsp;(4,504 instances). The entire train dataset was available for everyone since the start of the challenge. The public test dataset was available since the evaluation phase of the competition, but without any ground truth labels. After the end of the competition, public and private sets were released.</p> <p>The datasets will be provided as files in the comma-separated values (CSV) format containing the following columns.</p> <table> <tbody> </tbody> </table> <table> <tbody> <tr> <td><strong>Column</strong></td> <td><strong>Type</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>image</td> <td>string</td> <td>URL of an image on a public content delivery network</td> </tr> <tr> <td>width</td> <td>integer</td> <td>image width</td> </tr> <tr> <td>height</td> <td>integer</td> <td>image height</td> </tr> <tr> <td>left</td> <td>integer</td> <td>bounding box coordinate: left</td> </tr> <tr> <td>top</td> <td>integer</td> <td>bounding box coordinate: top</td> </tr> <tr> <td>right</td> <td>integer</td> <td>bounding box coordinate: right</td> </tr> <tr> <td>bottom</td> <td>integer</td> <td>bounding box coordinate: bottom</td> </tr> <tr> <td>question</td> <td>string</td> <td>question in English</td> </tr> </tbody> </table> <p>This upload also contains a ZIP file with the images from MS COCO.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Perceptions on the utility of community question and answer websites like Stack Overflow to software developers (Replication package)

<p>Interview Questions on the perception of the utility of CQAs like Stack Overflow to software developers. In this study, we focused on the questions highlighted in yellow.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

E-Commerce Question Answering Dataset

<p>The <strong>E</strong>-<strong>C</strong>ommerce <strong>Q</strong>uestion <strong>A</strong>nswering <strong>D</strong>ataset (ECQuAD) is a reading comprehension dataset for question answering in brazilian e-commerce platforms. It consists&nbsp;of questions annotated&nbsp;by crowdworkers on a set of products&#39; descriptions. It follows the SQuAD-v2 format, so questions might be unanswerable.</p> <p>This is a development set, for public usage, powered by <a href="https://gobots.ai/en/">GoBots</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

TempTabQA: Temporal Question Answering for Semi-Structured Tables

<p>This repository contains resources, namely <strong>TempTabQA</strong>, developed for the paper: Gupta, V., Kandoi, P., Vora, M., Zhang, S., He, Y., Reinanda R., Srikumar V., <i>TempTabQA: Temporal Question Answering for Semi-Structured Tables</i>. In: Proceeding of the The 2023 Conference on Empirical Methods in Natural Language Processing, Dec 2023.</p><p><strong>TempTabQA</strong> is a dataset which comprises 11,454 question-answer pairs extracted from Wikipedia Infobox tables. These question-answer pairs are annotated by human annotators. We provide two test sets instead of one: the <strong>Head</strong> set with popular frequent domains, and the <strong>Tail</strong> set with rarer domains.&nbsp;</p><p>Files to access the annotation follow the below structure:</p><p>Maindata</p><ul><li>qapairs: split into train, dev,&nbsp; head, and tail sets, in both csv and json formats</li><li>Tables: Wikipedia category and tables metadata in csv, json and html formats</li></ul><p>Carefully read the ```LICENCE``` for non-academic usage.</p><p><i>Note : Wherever required consider the year of 2022 as the build date for the dataset.</i></p><p>&nbsp;</p><p>&nbsp;</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Questions for Galops 1-7 horse riding theoretical exams and their textual answers with enumerations

<p><strong>Description in English :</strong></p><p>(Description en Français plus bas)</p><p><strong>Questions for Galops 1-7 horse riding theoretical exams and their textual answers with enumerations</strong></p><p>This French dataset was created to develop a system of automatic correction of textual answers with enumeration for horse riding theoretical exams, as part of Marine Potier's Master's thesis, a student in NLP (Natural Language Processing) at the Université de Franche-Comté.</p><p>This dataset contains 122 questions with their respective possible answers.</p><p>These questions and their answers were manually extracted from the official books published by the French Equestrian Federation.</p><p>We used ChatGPT (v.3.5) in order to get variations of the extracted answers, and every new answer was manually verified to make sure it was correct as well.</p><p>This dataset is available in CSV format and structured according to the following fields :</p><p>- <strong>question_id</strong> : The question ID, representing the level (1 to 7) followed by the question number.</p><p>- <strong>type_question</strong> : Either "sans_ordre" or "avec_ordre". Some answers need to be written in a precise order to be considered correct.</p><p>- <strong>question</strong> : The actual question, which was extracted from the official books.</p><p>- <strong>reponse_correcte</strong> : The correct answer, which was extracted from the official books.</p><p>- <strong>variation_1 to variation_5</strong> : The variations of the correct answer which was reformulated with ChatGPT. All were manually verified afterwards.</p><p>&nbsp;</p><p><strong>Description en Français :</strong></p><p><strong>Questions des examens théoriques Galops 1 à 7 d'équitation et leurs réponses textuelles contenant des énumérations</strong></p><p>Ce dataset en langue française a été créé pour développer un système de correction automatique de réponses textuelles contenant une énumération pour les examens théorique d'équitation (les Galops 1 à 7). Ce projet s'inscrit dans le cadre de mémoire de recherche de Marine Potier, étudiante en Master en Traitement Automatique des Langues à l'Université de Franche-Comté.</p><p>Ce dataset est constitué de 122 questions avec leurs réponses correctes respectives.</p><p>Ces questions et leurs réponses ont été extraites manuellement des guides fédéraux officiels publiés par la Fédération Française d'Équitation.</p><p>Nous avons utilisé ChatGPT (v.3.5) afin de recueillir des variations des réponses extraites. Chaque nouvelle réponse a été vérifiée manuellement pour s'assurer de leur exactitude.</p><p>Ce dataset est disponible en format CSV, et structuré comme expliqué ci-dessous :</p><p>- <strong>question_id</strong> : L'identifiant de la question, représentant le niveau (1 à 7) suivi du numéro de question.</p><p>- <strong>type_question</strong> : Soit "sans_ordre" ou "avec_ordre". Certaines questions nécessitent d'avoir leurs réponses rédigées suivant un ordre précis pour être considérées comme correctes.</p><p>- <strong>question</strong> : La question, extraite des guides fédéraux.</p><p>- <strong>reponse_correcte</strong> : La réponse correcte, extraite des guides fédéraux.</p><p>- <strong>variation_1 à variation_5</strong> : Les variations de la réponse correcte qui a été reformulée à l'aide de ChatGPT. Toutes ont été vérifiées manuellement par la suite.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

On the Helpfulness of Answering Developer Questions on Discord with Similar Conversations and Posts from the Past

<p>Replication Package for &quot;On the Helpfulness of Answering Developer Questions on Discord with Similar Conversations and Posts from the Past&quot;.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

PrimeKGQA, the dataset from paper: Bridging the Gap: Generating a Comprehensive Biomedical Knowledge Graph Question Answering Dataset

<p>Despite the plethora of resources such as large-scale&nbsp;corpora and manually curated Knowledge Graphs (KGs), the ability to perform reasoning with natural language inputs over biomedical graphs remains challenging due to insufficient training data. We&nbsp;propose a novel method for automatically constructing a Biomedical&nbsp;Knowledge Graph Question Answering (BioKGQA) dataset sourced&nbsp;from PrimeKG, the largest precision medicine-oriented KG. In total,<br>we create 83999 question-answer pairs along with their respective&nbsp;SPARQL queries. Our approach generates a diverse array of contextually relevant questions covering a wide spectrum of biomedical&nbsp;concepts and levels of complexity. We evaluate our method based on&nbsp;automatic metrics alongside manual annotations. We establish novel&nbsp;standards tailored for KGQA systems to highlight the linguistic correctness and semantical faithfulness of the generated questions based&nbsp;on extracted KG facts. The compiled dataset &ndash; PrimeKGQA &ndash; serves&nbsp;as a valuable benchmarking resource for advancing knowledge-driven biomedical research and evaluating KGQA system.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Analysis of the main answers to the survey questions [English version] / Análise das principais respostas às questões do survey [Versão em Português]

<p>Resultados da aplica&ccedil;&atilde;o do Survey&nbsp;(https://doi.org/10.5281/zenodo.6850089) com funcion&aacute;rios de uma empresa de TI do Sert&atilde;o Setentrional de Pernambuco em 2022, com nome preservado.</p> <p>&nbsp;</p> <p>Results of the application of the Survey (https://doi.org/10.5281/zenodo.6850089) with employees of an IT company in the Northern Sert&atilde;o of Pernambuco in 2022, with preserved name.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Surveying the Open Science Knowledge in a Southern Brazilian University - Survey Questions and Answers

<p>Surveying the Open Science Knowledge in a Southern Brazilian University - Survey Questions and Answers</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Theological Librarianship (Volume 16, Number 1): More questions than answers

In an effort to see how quickly I can do analysis, against the entire issue of a given journal, I did some reading against Theological Librarianship (Volume 16, Number 1), and in the end, I am going away with more questions than answers, but I did learn about "ethiopic".

opencc-zeroApr 2023View details →
zenodo40/100

Amharic visual question answering on Ethiopian tourism

<p>Visual Question Answering (VQA) is a Vision-to-Text (V2T) task that integrates visual&nbsp;<br>features of images with natural language questions to generate meaningful responses.&nbsp;<br>Most existing research has focused on English, leaving a significant gap for other&nbsp;<br>languages, including Amharic. Tourism, a major global industry, relies heavily on&nbsp;<br>interactions where visitors seek information about natural, historical, cultural, and&nbsp;<br>religious sites. Ethiopia is a remarkable tourist destination, home to unique sites such as&nbsp;<br>the Rock-hewn churches of Lalibela and the Castles of Gondar, as well as natural&nbsp;<br>phenomena like Simien National Park and Lake Tana. Most visitors are local, creating an&nbsp;<br>urgent need for a VQA model that can deliver accurate, culturally relevant information in&nbsp;<br>Amharic. Unfortunately, no such model currently exists to assist tourists at these heritage&nbsp;<br>sites. This research addresses this gap by developing an Amharic Visual Question&nbsp;<br>Answering model specifically tailored for Ethiopian tourism. A new Amharic VQA&nbsp;<br>dataset was created using 2,200 diverse images from Ethiopian tourist sites paired with&nbsp;<br>6,600 questions in Amharic, covering natural landmarks, historical sites, and religious&nbsp;<br>celebrations. Our dataset is collected from various sources, including the UNICCO&nbsp;<br>website, the Amhara Tourism office, and online platforms such as Facebook, Free pixel,&nbsp;<br>and Instagram. Each image is complemented by three corresponding questions&nbsp;<br>formulated by three individual experts and answered by ten candidates. The questions,&nbsp;<br>answers, and images are linked through annotations and fed into the model. We used&nbsp;<br>ResNet-50 for feature extraction and Bidirectional Gated Recurrent Unit (BiGRU) with&nbsp;<br>attention mechanisms, achieving a testing accuracy of 54.98%, demonstrating the model's&nbsp;<br>effectiveness in answering questions about Ethiopian heritage. We will expand this&nbsp;<br>research using external knowledge to gat answer and description beyond image and&nbsp;<br>custom object detection</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Analyzing the Impact of Copying-and-Pasting Vulnerable Solidity Code Snippets from Question-and-Answer Websites

<p>This data comprises all input and output, including intermediate results for the evaluation of the tool cpg-contract-checker(CCC) and the crawled data and results for the study published under the name "Analyzing the Impact of Copying-and-Pasting Vulnerable Solidity Code Snippets from Question-and-Answer Websites".<br>We conducted a study on the impact of vulnerable code reuse from Q&amp;A websites during the development of smart contracts and provided tools uniquely fit to detect vulnerable code patterns in complete and incomplete Smart Contract code. The paper proposes a pattern-based vulnerability detection tool that is able to analyze code snippets (i.e., incomplete code) as well as full smart contracts based on the concept of code property graphs. We also propose a methodology that leverages fuzzy hashing to quickly detect code clones of vulnerable snippets among deployed smart contracts. Our results show that our vulnerability search, as well as our code clone detection, are comparable to state-of-the-art while being applicable to code snippets. The tools are used to realize a study pipeline for which the dataset and (intermediate) results are contained in this archive.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

PhAQ: Intuitive Physics Question Answering forMulti-Modal Neural Network training

<p>Dataset for the paper:</p> <p>PhAQ: Intuitive Physics Question Answering forMulti-Modal Neural Network training</p> <p>The two proposed splits are given in zip files.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

MTL-QA : A dataset and multi-task learning approach for knowledge graph and natural language question answering

<p>The dataset used for this project is created by enhancing the publicly available MetaQA (Movie Text Audio QA), which is primarily a KGQA dataset pertaining to movies, an extension of WikiMovies. This involves questions requiring 1, 2, and 3 hops which can be answered by using a MetaQA Knowledge Graph. The questions are available in text and audio format. The text has vanilla (original) and its paraphrased version, and is called ntm.&nbsp;</p> <p>In order to develop a dataset to support NLQA, a series of dataset augmentation steps has been performed.</p> <p>The dataset consists of natural language questions and a tagged topic entity as ground truth. This topic entity is used to retrieve textual information related to the question from Wikipedia. The introduction section of the entity&#39;s page is used as the context that is required for NLQA. Hence, this dataset has information related to both KGQA and NLQA. Certain preliminary checks and validations are done to only retain those data samples whose context can be used to answer a given question.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Webis Causal Question Answering 2022

<p>The Webis Causal Question Answering 2022 (Webis-CausalQA-22) corpus comprises 1.1M causal question-answer pairs collected from the public QA datasets. This dataset was developed to support the development of tailored approaches that can answer causal questions.</p> <p>Overview:</p> <p>The directory &quot;input&quot; contains the train and validation splits (used for evaluation), the directory &quot;output&quot; contains the evaluation results, and the directory &quot;models&quot; includes the fine-tuned checkpoints.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Multi-Perspective Question Answering (MPQA)

<p>Multi-Perspective Question Answering is an imbalanced dataset of opinion polarity detection task. This dataset contains news documents from many sources and it was classified into positive and negative classes.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_&lt;k&gt;.pkl:&nbsp;&nbsp;pandas DataFrame with k-cross validation partition.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Question Answering auf dem Lehrbuch "Health Information Systems" mit Hilfe von unüberwachtem Training eines Pretrained Transformers

Masterthesis, verwendete Datensätze, Skripte und Ergebnisse des praktischen Teils

opencc-zeroSep 2023View details →
zenodo36/100

Bilingual Dataset for Information Retrieval and Question Answering over the Spanish Workers Statute

<p>A bilingual dataset of questions and answers over a key document in Spanish labor law legislation is presented. The document contains 150 questions and their respective answers in the form of one part number from the 130 parts in which the Workers Statute is divided (articles and other provisions), and with the most relevant excerpt of information for the answer.</p>

opencc-by-4.0Nov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record