Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3 results for “Open Information Extraction”

Learn how ShareScore rates datasets ↗
zenodo40/100

OIE4PA: Open Information Extraction for the Public Administration

<p>Tenders are powerful means of investment of public funds and represent a strategic development resource.<br> Despite the efforts made so far by governments at national and international levels to digitalise documents related to the Public Administration sector, most of the information is still available in an unstructured format only.&nbsp;<br> With the aim of bridging this gap, we present OIE4PA, our latest study on extracting and classifying relations from tenders of the Public Administration.<br> Our work focuses on the Italian language, where the availability of linguistic resources to perform Natural Language Processing tasks is considerably limited.&nbsp;<br> For evaluation purposes, we built a dataset composed of 2,000 triples extracted from Italian tenders, which have been manually annotated by two human experts.</p> <p>The dataset, compressed in a single zip file,&nbsp;is composed of:</p> <ul> <li>The corpus of 6,262 texts extracted from Italian public tenders (corpus_tenders)</li> <li>The training set of 1,600&nbsp;annotated triples (training_set)</li> <li>The test set of 400&nbsp;annotated triples (test_set)</li> <li>The set U&nbsp;of 14,096&nbsp;triples used for the self-training (u_triples_dd)</li> <li>a compressed archive that contains both the extracted triples and the index for each supervised approach (extraction)&nbsp;<br> &nbsp;</li> </ul>

opencc-by-4.0Sep 2023View details →
zenodo32/100

CompactIE: Compact Facts in Open Information Extraction [Dataset]

<p>Dataset for the paper <a href="https://arxiv.org/abs/2205.02880">&quot;CompactIE: Compact Facts in Open Information Extraction&quot;</a>.&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo12/100

MOIED: Magi Open Information Extraction Dataset

<p><strong>Description</strong></p> <p>Magi Open Information Extraction Dataset (MOIED) is a Chinese Open IE dataset containing 7,618,181 records extracted from plain text across 3,319,763 webpages in various domains. Each record in the dataset consists of the (subject, predicate, object) tuple, the associated confidence score, and the context information. The dataset comprises 1,427,742 distinct facts of 272,522 entities and 117,731 predicates.</p> <p>A notable property of MOIED is that each distinct fact has multiple records with URLs referring to mentions in diverse contexts, which enables multiple-instance learning (MIL) and other correlative approaches.</p> <p>As a paragraph level Open IE dataset, at least 45.1% of the records in MOIED can only be extracted through synthesizing information from multiple sentences.</p> <p>Magi is an extraction engine that continuously learns from the Internet, which combines cross-referencing, timeline analysis, and other heuristics to mitigate the inevitable false positives in the extractions. All records in MOIED were randomly sampled from a database dump of <a href="https://magi.com/">magi.com</a> in January 2020. To provide more reliable evaluation results, human annotators examined the dataset and selected 19,161 verified records for the dev and test sets.</p> <p>&nbsp;</p> <p><strong>Disclaimers</strong></p> <ol> <li>The dataset is expected to be used in weakly supervised scenarios since the records in the training set are not human-annotated and could be imprecise or erroneous.</li> <li>Records are not guaranteed to be universally correct. The correctness of extractions should be evaluated based on contexts (specified by the URLs).</li> <li>The extraction was made at a certain time Magi visits the URL, thus it is not guaranteed that the URL is still accessible, or the content is unmodified since the extraction was conducted.</li> <li>Due to legal and regulatory issues, the webpage URLs are mostly ones accessible from Mainland China, yet, the content of certain webpages, as well as the extraction results, could be in violation of law and regulation of certain countries or regions in certain ways.</li> </ol>

restrictedFeb 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record