Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48
datasets available to search
ShareScore release 0.9.0
Dataset results
48 results for “Information Extraction”
Classical art semantics information extraction
<p>Abstract. The paper discusses the application of Natural Language Processing (NLP) techniques in the context of semantic annotation of classical art text via rule-based Information Extraction (IE) techniques combined with ontological and domain vocabulary input. The CASIE (Classical Art Semantics Information Extraction) was a pilot collaborative project between the Hypermedia Research<br> Unit (University of South Wales) and the Beazley Archive (Oxford University), which aims to automatically extract information about cultural objects from classical art scholarly texts and represent this information in terms of the ISO metadata standard for cultural heritage, the International Council of Museum’s CIDOC Conceptual Reference Model (CRM). In total 12 documents (fascicules<br> – high quality catalogues) were processed, originating from the Corpus Vasorum Antiquorum (CVA) collection containing over 350 high quality catalogues of mostly ancient Greek painted pottery, illustrating more than 100,000 vases. The extracted information was expressed in interoperable RDF graphs consistent with the CLAROS project format. The role of CIDOC-CRM is central for enabling semantic interoperability across the range of datasets that contribute to CLAROS. The CASIE pilot enabled a complementary exploitation of terminological and ontological resources via rule-based information extraction techniques, delivering semantic annotation with respect to the CRM in the broader field of digital humanities.</p>
Data from: Robust extraction of quantitative structural information from high-variance histological images of livers from necropsied Soay sheep
Open the record for dataset details and reuse information.
End Sequence Analysis ToolKit (ESAT) expands the extractable information from single cell RNA-Seq data
GEO Series GSE79651. Mus musculus; Rattus norvegicus. 14 samples. Type: Expression profiling by high throughput sequencing.
Platform for Medical Information Extraction From Incomplete Data
ClinicalTrials.gov study NCT01813942. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Information Extraction and Database Construction for Emergent Patients Based on Voice and Image Recognition Technology
ClinicalTrials.gov study NCT04918979. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Anxiety and Information Methods in Molar Extraction
ClinicalTrials.gov study NCT07019688. IPD Sharing: NO. Countries: 1. Publications: 0.
eQTL analysis - Extracting genotype information of recombinant inbred lines from transcript profiles established with high-density oligonucleotide arrays
GEO Series GSE99150. Arabidopsis thaliana. 116 samples. Type: Expression profiling by array.
MOIED: Magi Open Information Extraction Dataset
<p><strong>Description</strong></p> <p>Magi Open Information Extraction Dataset (MOIED) is a Chinese Open IE dataset containing 7,618,181 records extracted from plain text across 3,319,763 webpages in various domains. Each record in the dataset consists of the (subject, predicate, object) tuple, the associated confidence score, and the context information. The dataset comprises 1,427,742 distinct facts of 272,522 entities and 117,731 predicates.</p> <p>A notable property of MOIED is that each distinct fact has multiple records with URLs referring to mentions in diverse contexts, which enables multiple-instance learning (MIL) and other correlative approaches.</p> <p>As a paragraph level Open IE dataset, at least 45.1% of the records in MOIED can only be extracted through synthesizing information from multiple sentences.</p> <p>Magi is an extraction engine that continuously learns from the Internet, which combines cross-referencing, timeline analysis, and other heuristics to mitigate the inevitable false positives in the extractions. All records in MOIED were randomly sampled from a database dump of <a href="https://magi.com/">magi.com</a> in January 2020. To provide more reliable evaluation results, human annotators examined the dataset and selected 19,161 verified records for the dev and test sets.</p> <p> </p> <p><strong>Disclaimers</strong></p> <ol> <li>The dataset is expected to be used in weakly supervised scenarios since the records in the training set are not human-annotated and could be imprecise or erroneous.</li> <li>Records are not guaranteed to be universally correct. The correctness of extractions should be evaluated based on contexts (specified by the URLs).</li> <li>The extraction was made at a certain time Magi visits the URL, thus it is not guaranteed that the URL is still accessible, or the content is unmodified since the extraction was conducted.</li> <li>Due to legal and regulatory issues, the webpage URLs are mostly ones accessible from Mainland China, yet, the content of certain webpages, as well as the extraction results, could be in violation of law and regulation of certain countries or regions in certain ways.</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.