Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

20

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

20 results for “knowledge extraction”

Learn how ShareScore rates datasets ↗
zenodo40/100

Zero-Shot Information Extraction to Enhance a Knowledge Graph Describing Silk Textiles - English and Spanish neighborhood sub-graphs

<p>Two language-specific sub-graphs (English and Spanish) based on the ConceptNet Knowledge Graph. These two files are required to run the code for reproducing the results reported in the paper <a href="https://aclanthology.org/2021.latechclfl-1.16/">&quot;Zero-Shot Information Extraction to Enhancea Knowledge Graph Describing Silk Textiles&quot;</a> at the <a href="https://sighum.wordpress.com/events/latech-clfl-2021/">LaTeCH-CLfL 2021</a> workshop co-located with <a href="https://2021.emnlp.org/">EMNLP 2021</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Dataset for "Unleashing the Power of Knowledge Extraction from Scientific Literature in Catalysis"

<p>JCIM paper link: <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c00359">https://pubs.acs.org/doi/10.1021/acs.jcim.2c00359</a><br><br>Github repo: <a href="https://github.com/nsndimt/CatalysisIE">https://github.com/nsndimt/CatalysisIE</a><br><br><br>Dataset Content:</p> <ul> <li>Pretrained BERT:&nbsp;<code>scibert_domain_adaption.tar.gz</code>&nbsp;extract it to&nbsp;<em>pretrained</em> directory</li> <li>Cross-Validation Checkpoint:&nbsp;<code>cross_validation_checkpoint.tar.gz</code>&nbsp;extract it to&nbsp;<em>checkpoint</em> directory</li> <li>Annotated Data: <code>data.jsonl</code> and <code>split.jsonl</code>&nbsp;put it under&nbsp;<em>data</em> directory</li> </ul>

opencc-by-4.0May 2022View details →
zenodo36/100

Enhanced Kinase Dictionaries associated with KinDER: A Biocuration Tool for Extracting Kinase Knowledge from Biomedical Literature

<p>This zip file contains Kinase dictionaries used for annotating documents with KinDER described&nbsp;in the following paper:&nbsp;</p> <p>Dopp, Daniel, Adam Morrone, and Indika Kahanda. &quot;KinDER: A Biocuration Tool for Extracting Kinase Knowledge from Biomedical Literature.&quot;&nbsp;<em>Proceedings of the BioCreative VI Workshop</em>. 2017.</p>

opencc-by-4.0Dec 2017View details →
zenodo36/100

NED data for the paper Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts

<p>This data repository contains NED data from the paper,&nbsp;<em><a href="https://arxiv.org/abs/2309.01812">Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts.</a></em></p> <p>Additional data for the NER classification task can be found here: <a href="../records/10050681">zenodo</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems

<p>FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems. Dataset for the task force "Vulcan" from ISWS 2024 led by Aldo Gangemi and Andrea Poltronieri</p>

opencc-by-4.0Jun 2024View details →
ClinicalTrials.gov36/100

Multi-Modal Education to Improve Compliance, Knowledge Retention & Anxiety After Dental Extractions

ClinicalTrials.gov study NCT07191132. IPD Sharing: YES. Countries: 1. Publications: 13.

controlledIPD-YESFeb 2026View details →
zenodo32/100

Data extraction sheet: Knowledge, Attitudes, and Practices of Women and Men Towards Infertility: A Scoping Review

<p>This is the data extraction tool that includes all the studies reviewed and inclided in: Knowledge, Attitudes, and Practices of Women and<br>Men Towards Infertility: A Scoping Review</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Screencast of the Lokahi2 Prototype: Search Engine with Interactive Knowledge Network Browser Extracted from Text

<p>This video shows a recording of the prototype system Lokahi2. It supports concept surfing for interacting with&nbsp;the search engine, automated document tagging, exploring tags, generating&nbsp;tags for text, and changing&nbsp;depth&nbsp;and dimension of the graph.</p>

opencc-by-4.0Jul 2018View details →
zenodo32/100

Target Speech Extraction Dataset for Knowledge Boosting (Part 1)

<p><strong>Part 1 of the Target Speech Extraction Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJun 2024View details →
zenodo32/100

Target Speech Extraction Dataset for Knowledge Boosting (Part 2)

<p><strong>Part 2 of the Target Speech Extraction Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJul 2024View details →
zenodo32/100

Cytoscape session for the potato knowledge graph extracted with IBM Watson's supervised NLP model

<p>2 cytoscape session files (.cys) representing the genotypic-phenotypic knowledge networks retrieved from scientific literature using IBM Watson. Input to these&nbsp;files were the following:&nbsp;</p> <ul> <li>cytoscapeSession_trainingSet: A training set of 34 full-test&nbsp;articles about potato flesh color</li> <li>cytoscapesession_testSet: A testing set of a&nbsp;4023 abstracts from PubMed.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2019View details →
zenodo32/100

Dataset and code for 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'

<p>Zip file containing the dataset and code for the paper&nbsp;&#39;AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends&#39;.&nbsp;The dataset is composed of:</p> <ul> <li>A train_data.csv file, containing all annotated keywords used for classifier training.</li> <li>A filt_ls.pkl file, containing the sentences used to train the embeddings model.</li> <li>A train.py file, to train the composite keyword annotation model.</li> </ul> <p>The authors acknowledge the&nbsp;supported by the European Union&rsquo;s Horizon 2020 research and innovation program under the project GIOTTO: &ldquo;Giotto: Active ageing and osteoporosis: The next challenge for smart nanobiomaterials and&nbsp;3D technologies,&rdquo; grant agreement no. 814410.</p>

opencc-by-4.0Dec 2022View details →
zenodo28/100

Test set - 4023 PubMed abstracts (for manuscript: Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait )

<p>A .zip archive containing the set of abstracts used in the test set (4023 abstracts from PubMed) in .txt format.</p> <p>This archive contains supplementary files for the manuscript Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait.</p>

opencc-by-4.0Dec 2019View details →
zenodo28/100

Supplemental Material for Paper 'Taxonomy Extraction Using Knowledge Graph Embeddings and Hierarchical Clustering'

<p>Contains input data and gold standard for the non-expressive extraction task, as well as examples of extracted taxonomies for both the non-expressive and expressive cases. Extracted taxonomies can also be found at&nbsp;<a href="http://labowest.ca/sdb2020/">labowest.ca</a>.</p>

opencc-by-4.0Dec 2019View details →
dryad28/100

Data from: Toward synthesizing our knowledge of morphology: using ontologies and machine reasoning to extract presence/absence evolutionary phenotypes across studies

The reality of larger and larger molecular databases and the need to integrate data scalably have presented a major challenge for the use of phenotypic data. Morphology is currently primarily described in discrete publications, entrenched in noncomputer readable text, and requires enormous investments of time and resources to integrate across large numbers of taxa and studies. Here we present a new methodology, using ontology-based reasoning systems working with the Phenoscape Knowledgebase (KB; kb.phenoscape.org), to automatically integrate large amounts of evolutionary character state descriptions into a synthetic character matrix of neomorphic (presence/absence) data. Using the KB, which includes more than 55 studies of sarcopterygian taxa, we generated a synthetic supermatrix of 639 variable characters scored for 1051 taxa, resulting in over 145,000 populated cells. Of these characters, over 76% were made variable through the addition of inferred presence/absence states derived by machine reasoning over the formal semantics of the source ontologies. Inferred data reduced the missing data in the variable character-subset from 98.5% to 78.2%. Machine reasoning also enables the isolation of conflicts in the data, that is, cells where both presence and absence are indicated; reports regarding conflicting data provenance can be generated automatically. Further, reasoning enables quantification and new visualizations of the data, here for example, allowing identification of character space that has been undersampled across the fin-to-limb transition. The approach and methods demonstrated here to compute synthetic presence/absence supermatrices are applicable to any taxonomic and phenotypic slice across the tree of life, providing the data are semantically annotated. Because such data can also be linked to model organism genetics through computational scoring of phenotypic similarity, they open a rich set of future research questions into phenotype-to-genome relationships.

opencc-zeroDec 2014View details →
zenodo28/100

OPC UA Knowledge Extraction (OKE) Dataset

Open the record for dataset details and reuse information.

openmit-licenseDec 2023View details →
zenodo28/100

The Brill Knowledge Graph: A Database of Bibliographic References and Index Terms extracted from Books in Humanities and Social Sciences

<p>We present a complete dataset of linked bibliography and index data, partially disambiguated and augmented with references to external resources, extracted from the Brill&rsquo;s archive in the field of Classics. Processed book identifiers are listed in a separate&nbsp;text file. Text fragments extracted from different books via this process are then parsed and compared using a string-based similarity metric to form clusters of bibliographic references to the same published work or (variants of) the same subjects discussed in these books. The entire set of references was then disambiguated using Google Books and Crossref APIs.</p> <p><a href="https://jdmdh.episciences.org/11062">Paper about extraction pipeline</a></p> <p><a href="https://www.nkokash.com/documents/KIEM-RDJ.pdf">Paper about extracted KG</a></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
dryad28/100

Data from: Toward synthesizing our knowledge of morphology: using ontologies and machine reasoning to extract presence/absence evolutionary phenotypes across studies

Open the record for dataset details and reuse information.

publicJun 2015View details →
ClinicalTrials.gov24/100

Effectiveness of Active and Passive Distraction Techniques on Reducing Fear and Anxiety and Improving Oral Health Knowledge of Children Undergoing Extraction in the Dental Operatory- A Randomized Cont

ClinicalTrials.gov study NCT03247959. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo20/100

Dataset: Relationship extraction for knowledge graph creation from biomedical literature (Gene-Disease relationships)

<p>This is the dataset used for classifying Gene-Disease relationship types from sentences. The dataset consists of 3 files:</p> <ul> <li>manually_annotated_set.xlsx - set of 2000 manualy annotated sentences with entities</li> <li>Unbalanced_dataset.xlsx - set of 12000 sentences, out of which 2000 are from the first set, manually annotated, and the rest have been added using rule based method by adding sentences where extraction had confidence 1.</li> <li>Balanced_dataset_SUB_PRED.xlsx - balanced dataset generated by taking 2000 manually annotated sentences, but then adding sentences from the rule-based method with confidence 1 in such a way that each relationship class had at least 1400 sentences (for biomarkers, we could obtain 1243 sentences with confidence 1 from a processed portion of the data we had at the time of building the dataset).</li> </ul> <p>&nbsp;</p>

restrictedApr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record