Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20
datasets available to search
ShareScore release 0.9.0
Dataset results
20 results for “knowledge extraction”
Zero-Shot Information Extraction to Enhance a Knowledge Graph Describing Silk Textiles - English and Spanish neighborhood sub-graphs
<p>Two language-specific sub-graphs (English and Spanish) based on the ConceptNet Knowledge Graph. These two files are required to run the code for reproducing the results reported in the paper <a href="https://aclanthology.org/2021.latechclfl-1.16/">"Zero-Shot Information Extraction to Enhancea Knowledge Graph Describing Silk Textiles"</a> at the <a href="https://sighum.wordpress.com/events/latech-clfl-2021/">LaTeCH-CLfL 2021</a> workshop co-located with <a href="https://2021.emnlp.org/">EMNLP 2021</a>.</p>
Dataset for "Unleashing the Power of Knowledge Extraction from Scientific Literature in Catalysis"
<p>JCIM paper link: <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c00359">https://pubs.acs.org/doi/10.1021/acs.jcim.2c00359</a><br><br>Github repo: <a href="https://github.com/nsndimt/CatalysisIE">https://github.com/nsndimt/CatalysisIE</a><br><br><br>Dataset Content:</p> <ul> <li>Pretrained BERT: <code>scibert_domain_adaption.tar.gz</code> extract it to <em>pretrained</em> directory</li> <li>Cross-Validation Checkpoint: <code>cross_validation_checkpoint.tar.gz</code> extract it to <em>checkpoint</em> directory</li> <li>Annotated Data: <code>data.jsonl</code> and <code>split.jsonl</code> put it under <em>data</em> directory</li> </ul>
Enhanced Kinase Dictionaries associated with KinDER: A Biocuration Tool for Extracting Kinase Knowledge from Biomedical Literature
<p>This zip file contains Kinase dictionaries used for annotating documents with KinDER described in the following paper: </p> <p>Dopp, Daniel, Adam Morrone, and Indika Kahanda. "KinDER: A Biocuration Tool for Extracting Kinase Knowledge from Biomedical Literature." <em>Proceedings of the BioCreative VI Workshop</em>. 2017.</p>
NED data for the paper Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts
<p>This data repository contains NED data from the paper, <em><a href="https://arxiv.org/abs/2309.01812">Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts.</a></em></p> <p>Additional data for the NER classification task can be found here: <a href="../records/10050681">zenodo</a></p> <p> </p> <p> </p>
FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems
<p>FLIP-KG: Enriching Automated Knowledge Graph Extraction with Tacit Knowledge. The Case of Lyrical Implicatures within Poems. Dataset for the task force "Vulcan" from ISWS 2024 led by Aldo Gangemi and Andrea Poltronieri</p>
Multi-Modal Education to Improve Compliance, Knowledge Retention & Anxiety After Dental Extractions
ClinicalTrials.gov study NCT07191132. IPD Sharing: YES. Countries: 1. Publications: 13.
Data extraction sheet: Knowledge, Attitudes, and Practices of Women and Men Towards Infertility: A Scoping Review
<p>This is the data extraction tool that includes all the studies reviewed and inclided in: Knowledge, Attitudes, and Practices of Women and<br>Men Towards Infertility: A Scoping Review</p>
Screencast of the Lokahi2 Prototype: Search Engine with Interactive Knowledge Network Browser Extracted from Text
<p>This video shows a recording of the prototype system Lokahi2. It supports concept surfing for interacting with the search engine, automated document tagging, exploring tags, generating tags for text, and changing depth and dimension of the graph.</p>
Target Speech Extraction Dataset for Knowledge Boosting (Part 1)
<p><strong>Part 1 of the Target Speech Extraction Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
Target Speech Extraction Dataset for Knowledge Boosting (Part 2)
<p><strong>Part 2 of the Target Speech Extraction Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
Cytoscape session for the potato knowledge graph extracted with IBM Watson's supervised NLP model
<p>2 cytoscape session files (.cys) representing the genotypic-phenotypic knowledge networks retrieved from scientific literature using IBM Watson. Input to these files were the following: </p> <ul> <li>cytoscapeSession_trainingSet: A training set of 34 full-test articles about potato flesh color</li> <li>cytoscapesession_testSet: A testing set of a 4023 abstracts from PubMed.</li> </ul> <p> </p>
Dataset and code for 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'
<p>Zip file containing the dataset and code for the paper 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'. The dataset is composed of:</p> <ul> <li>A train_data.csv file, containing all annotated keywords used for classifier training.</li> <li>A filt_ls.pkl file, containing the sentences used to train the embeddings model.</li> <li>A train.py file, to train the composite keyword annotation model.</li> </ul> <p>The authors acknowledge the supported by the European Union’s Horizon 2020 research and innovation program under the project GIOTTO: “Giotto: Active ageing and osteoporosis: The next challenge for smart nanobiomaterials and 3D technologies,” grant agreement no. 814410.</p>
Test set - 4023 PubMed abstracts (for manuscript: Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait )
<p>A .zip archive containing the set of abstracts used in the test set (4023 abstracts from PubMed) in .txt format.</p> <p>This archive contains supplementary files for the manuscript Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait.</p>
Supplemental Material for Paper 'Taxonomy Extraction Using Knowledge Graph Embeddings and Hierarchical Clustering'
<p>Contains input data and gold standard for the non-expressive extraction task, as well as examples of extracted taxonomies for both the non-expressive and expressive cases. Extracted taxonomies can also be found at <a href="http://labowest.ca/sdb2020/">labowest.ca</a>.</p>
Data from: Toward synthesizing our knowledge of morphology: using ontologies and machine reasoning to extract presence/absence evolutionary phenotypes across studies
The reality of larger and larger molecular databases and the need to integrate data scalably have presented a major challenge for the use of phenotypic data. Morphology is currently primarily described in discrete publications, entrenched in noncomputer readable text, and requires enormous investments of time and resources to integrate across large numbers of taxa and studies. Here we present a new methodology, using ontology-based reasoning systems working with the Phenoscape Knowledgebase (KB; kb.phenoscape.org), to automatically integrate large amounts of evolutionary character state descriptions into a synthetic character matrix of neomorphic (presence/absence) data. Using the KB, which includes more than 55 studies of sarcopterygian taxa, we generated a synthetic supermatrix of 639 variable characters scored for 1051 taxa, resulting in over 145,000 populated cells. Of these characters, over 76% were made variable through the addition of inferred presence/absence states derived by machine reasoning over the formal semantics of the source ontologies. Inferred data reduced the missing data in the variable character-subset from 98.5% to 78.2%. Machine reasoning also enables the isolation of conflicts in the data, that is, cells where both presence and absence are indicated; reports regarding conflicting data provenance can be generated automatically. Further, reasoning enables quantification and new visualizations of the data, here for example, allowing identification of character space that has been undersampled across the fin-to-limb transition. The approach and methods demonstrated here to compute synthetic presence/absence supermatrices are applicable to any taxonomic and phenotypic slice across the tree of life, providing the data are semantically annotated. Because such data can also be linked to model organism genetics through computational scoring of phenotypic similarity, they open a rich set of future research questions into phenotype-to-genome relationships.
OPC UA Knowledge Extraction (OKE) Dataset
Open the record for dataset details and reuse information.
The Brill Knowledge Graph: A Database of Bibliographic References and Index Terms extracted from Books in Humanities and Social Sciences
<p>We present a complete dataset of linked bibliography and index data, partially disambiguated and augmented with references to external resources, extracted from the Brill’s archive in the field of Classics. Processed book identifiers are listed in a separate text file. Text fragments extracted from different books via this process are then parsed and compared using a string-based similarity metric to form clusters of bibliographic references to the same published work or (variants of) the same subjects discussed in these books. The entire set of references was then disambiguated using Google Books and Crossref APIs.</p> <p><a href="https://jdmdh.episciences.org/11062">Paper about extraction pipeline</a></p> <p><a href="https://www.nkokash.com/documents/KIEM-RDJ.pdf">Paper about extracted KG</a></p> <p> </p>
Data from: Toward synthesizing our knowledge of morphology: using ontologies and machine reasoning to extract presence/absence evolutionary phenotypes across studies
Open the record for dataset details and reuse information.
Effectiveness of Active and Passive Distraction Techniques on Reducing Fear and Anxiety and Improving Oral Health Knowledge of Children Undergoing Extraction in the Dental Operatory- A Randomized Cont
ClinicalTrials.gov study NCT03247959. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Dataset: Relationship extraction for knowledge graph creation from biomedical literature (Gene-Disease relationships)
<p>This is the dataset used for classifying Gene-Disease relationship types from sentences. The dataset consists of 3 files:</p> <ul> <li>manually_annotated_set.xlsx - set of 2000 manualy annotated sentences with entities</li> <li>Unbalanced_dataset.xlsx - set of 12000 sentences, out of which 2000 are from the first set, manually annotated, and the rest have been added using rule based method by adding sentences where extraction had confidence 1.</li> <li>Balanced_dataset_SUB_PRED.xlsx - balanced dataset generated by taking 2000 manually annotated sentences, but then adding sentences from the rule-based method with confidence 1 in such a way that each relationship class had at least 1400 sentences (for biomarkers, we could obtain 1243 sentences with confidence 1 from a processed portion of the data we had at the time of building the dataset).</li> </ul> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.