Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
51
datasets available to search
ShareScore release 0.9.0
Dataset results
51 results for “Relation Extraction”
Extracted microbial terms from Wikipedia - Marine Microbiology related pages (refined)
<p>From the dataset titled "Extracted microbial terms from Wikipedia - Marine Microbiology related pages (unfiltered)" available on Zenodo https://doi.org/10.5281/zenodo.12648788, we reviewed the full dataset and narrowed it down to a subset of 100 terms directly related to microbiology. We excluded terms such as "marine", "organisms", "ocean", "life" as they are not strictly related to microbiology. Additionally, terms like "algae" and "plants" were removed since they are not microbial entities. We also avoided including verbs such as "found", "classify", and "form" because they may be related to sentences or methods in microbiology, but not strictly to microbial entities.</p>
Annotated Dataset for Named Entity Recognition and Relation Extraction in French Building Technical Specifications (BTS)
<p>This dataset contains 233 raw requirements extracted from French Building Technical Specifications (BTS), referred to as "<a href="https://www.aglo.ai/cctp/#:~:text=Le%20CCTP%20(Cahier%20des%20Clauses%20Techniques%20Particuli%C3%A8res)%20est%20un%20document,code%20du%20march%C3%A9%20public%20donc."><em>Cahier des Clauses Techniques Particulières (CCTP)</em></a>", specifically focused on carpentry ("<em>lot menuiserie</em>") in public French construction projects. The requirements have been collected from 72 CCTP documents, resulting in a total of 19,725 sentences and 651,948 words.</p> <p>The dataset has been annotated using <a title="Open-source text annotation tool" href="https://github.com/doccano/doccano">Doccano </a>for Named Entity Recognition (NER) and Relation Extraction (RE). The annotations involve identifying entities and the relationships between them within the domain of building requirements. This dataset is intended for research on Natural Language Processing (NLP) models for Requirements Engineering (RE) in the Architecture, Engineering, and Construction (AEC) sector. Potential applications include requirements extraction, compliance analysis, and knowledge management in construction.</p> <p>The dataset includes the following components:</p> <ol> <li><strong>CCTP Documents</strong>: The original CCTP files from which the raw requirements were extracted.</li> <li><strong>Annotated Dataset</strong>: A JSONLines file containing the annotated dataset, including labels for Named Entity Recognition (NER) and Relation Extraction (RE).</li> </ol> <p>Key features of the dataset:</p> <ul> <li>Language: French</li> <li>Number of requirements: 233</li> <li>Number of sentences: 19,725</li> <li>Number of words: 651,948</li> <li>Annotation tasks: Named Entity Recognition (NER) and Relation Extraction (RE)</li> </ul> <p>This dataset is relevant for NLP research focused on structured information extraction from domain-specific texts in the construction industry.</p>
Dataset related to the article "Extraction-Free Absolute Quantification of Circulating miRNAs by Chip-Based Digital PCR"
<p>This record contains raw data related to the article "Extraction-Free Absolute Quantification of Circulating miRNAs by Chip-Based Digital PCR"</p> <p>Circulating microRNAs (miRNA) have been proposed as specific biomarkers for several diseases. Quantitative Real-Time PCR (RT-qPCR) is the gold standard technique currently used to evaluate miRNAs expression from different sources. In the last few years, digital PCR (dPCR) emerged as a complementary and accurate detection method. When dealing with gene expression, the first and most delicate step is nucleic-acid isolation. However, all currently available protocols for RNA extraction suffer from the variable loss of RNA species due to the chemicals and number of steps involved, from sample lysis to nucleic acid elution. Here, we evaluated a new process for the detection of circulating miRNAs, consisting of sample lysis followed by direct evaluation by dPCR in plasma from healthy donors and in the cardiovascular setting. Our results showed that dPCR is able to detect, with high accuracy, low-copy-number as well as highly expressed miRNAs in human plasma samples without the need for RNA extraction. Moreover, we assessed a known myocardial infarction-related miR-133a in acute myocardial infarct patients vs. healthy subjects. In conclusion, our results show the suitability of the extraction-free quantification of circulating miRNAs as disease markers by direct dPCR.</p>
Synthetic datasets for end-to-end Relation Extraction of relationships between Organisms and Natural-Products
<p>Synthetic datasets (training/validation) for end-to-end Relation Extraction of relationships between Organisms and Natural-Products. The datasets are provided for reproducibility purposes, but, can also be used to train new models.</p><p>As in the corresponding article, 3 subtypes of synthetic datasets are provided:</p><ul><li><i>Diversity-synt</i>:<strong> </strong>The seed literature references used in the generation process correspond to the top-500 extracted items per biological kingdoms using the <a href="https://github.com/idiap/gme-sampler">GME-sampler</a>.</li><li><i>Random-synt</i>:<strong> </strong>5 datasets of equivalent sizes as <i>Diversity-synt</i>, but using randomly sampled seed literature references.</li><li><i>Extended-synt</i>: A merge of <i>Diversity-synt and the 5 Random-synt datasets.</i><br> </li></ul><p><strong>All </strong>datasets were produced with <a href="https://huggingface.co/lmsys/vicuna-13b-v1.3">Vicuna-13b-v1.3</a>. Like the model, the produced synthetic data are also submitted to the License of the model used for generation, see the original <a href="https://github.com/facebookresearch/llama/blob/llama_v1/MODEL_CARD.md">LLaMA model card</a>.</p><p>LLaMA is licensed under the <a href="https://github.com/facebookresearch/llama/blob/llama_v1/MODEL_CARD.md">LLaMA License</a>, Copyright (c) Meta Platforms, Inc. All Rights Reserved. </p>
A Randomised, Cross-over, Relative Bioavailability Study of Nicotine Delivery and Nicotine Extraction From Oral Products
ClinicalTrials.gov study NCT04891406. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Dataset of miRNA-Disease Relations Extracted from Textual Data using Transformer-based Neural Networks
<p>Supplementary Data.</p>
Evaluation of SPIRES on Chemical-Disease-Relation extraction task 2023-01
<p>SPIRES CDR evaluation results.</p> <p>See https://github.com/monarch-initiative/ontogpt</p>
Pre-trained Model and Relation Extraction Dataset
<p>NIPA</p>
Effect of Ishige Okamurae Extract on Musculoskeletal Biomarkers in Adults With Relative Sarcopenia
ClinicalTrials.gov study NCT04617951. IPD Sharing: NO. Countries: 1. Publications: 1.
The Effect of Extract of Green Tea on Obese Women and Obese Related Hormone Peptides
ClinicalTrials.gov study NCT02147041. IPD Sharing: Not stated. Countries: 1. Publications: 2.
Effect of Schisandra Chinensis Extract on Musculoskeletal Biomakers in Relatively Sarcopenic Adults: a RCT
ClinicalTrials.gov study NCT03402308. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Traditional vs Orthodontic Extraction of Impacted Teeth Related to the Inferior Alveolar Nerve
ClinicalTrials.gov study NCT06270784. IPD Sharing: NO. Countries: 1. Publications: 1.
Effect of Fermented Oyster Extract on Musculoskeletal Biomarkers in Relative Sarcopenia Adults
ClinicalTrials.gov study NCT04109911. IPD Sharing: NO. Countries: 1. Publications: 1.
Bangla-REX: A Distinct Dataset for Bangla Relation Extraction
<p>The dataset is grounded in theoretical and methodological frameworks that emphasize the importance of structured knowledge bases and annotated corpora for effective relation extraction. To generate this dataset, we compiled a comprehensive Bangla Knowledge Base (KB) consisting of 63,256 entries, which serves as a foundation for automating the labeling process with relation tags. The corpus itself is extensive, comprising 90,441 text entries that have been meticulously processed to include Named Entity Recognition (NER) and Part-of-Speech (POS) tagging, ensuring that it is ready for immediate use in relation extraction tasks.<br>Additionally, we developed mnemonics for 440 distinct locations in Bangla, specifically tailored to enhance performance in location-based relation extraction. These mnemonics are particularly beneficial in the context of distant supervision-based relation extraction, where they help in establishing clear associations between locations and their corresponding entities or contexts.</p>
Data from: Immunotherapy-related adverse events (irAEs): extraction from FDA drug labels and comparative analysis
Objectives: Immune checkpoint inhibitors (ICIs) have dramatically improved outcomes in cancer patients. However, ICIs are associated with significant immune-related adverse events (irAEs) and the underlying biological mechanisms are not well-understood. To ensure safe cancer treatment, research efforts are needed to comprehensively detect and understand irAEs. Materials and Methods: We manually extracted and standardized irAEs from FDA drug labels for six FDA-approved ICIs. We compared irAE profile similarities among ICIs and 1,507 FDA-approved non-ICI drugs. We investigated how irAEs have differential effects on human organs by classifying irAEs based on their targeted organ systems. Finally, we identified broad-spectrum (non-target specific) and narrow-spectrum (target-specific) irAEs. Results: A total of 893 irAEs were extracted. 31.4% irAEs were shared among ICIs as compared to the 8.0% between ICIs and non-ICI drugs. irAEs were resulted from both on- and off-target effects: irAE profiles were more similar for ICIs with same target than different targets, demonstrating the on-target effects; irAE profile similarity for ICIs with the same target is less than 50%, demonstrating unknown off-target effects. ICIs significantly target many organ systems, including endocrine system (3.4-fold enrichment), metabolism (3.7-fold enrichment), immune system (3.6-fold enrichment) and autoimmune system (4.8-fold enrichment). We identified 21 broad-spectrum irAEs shared among all ICIs, 20 irAEs specific for PD-L1/PD-1 inhibition, and 28 irAEs specific for CTLA-4 inhibition. Discussion and Conclusion: Our study presents the first effort toward building a standardized database of irAEs. The extracted irAEs can serve as a goldstandard for automatic irAE extractions from other data resources and set a foundation to understand biological mechanisms of irAEs.
Supplementary data to the paper: Automatic Extraction of Anthropometric Features for the Individualization of the Pinna-Related Transfer Function in the Median Plane
<p>Supplementary research data to the paper (rejected):</p> <blockquote> <p>Davide Fantini, Federico Avanzini, Stavros Ntalampiras and Giorgio Presti (2023) "Automatic Extraction of Anthropometric Features for the Individualization of the Pinna-Related Transfer Function in the Median Plane"</p> </blockquote> <p>The repository includes the research data generated and analyzed in the abovementioned paper describing a method for PRTF individualization. In particular, the following data are included:</p> <ul> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/README.md">README.md</a>: instructions for the data</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/pinna_range_img.mat">pinna_range_img.mat</a>: pinna range images extracted from the 3D head meshes of the <a href="https://depositonce.tu-berlin.de/items/dc2a3076-a291-417e-97f0-7697e332c960">HUTUBS dataset</a></li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/landmarks.mat">landmarks.mat</a>: landmarks coordinates both manually annotated and automatically placed with ASM</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/anthropometry.mat">anthropometry.mat</a>: anthropometric parameters automatically extracted from both manually annotated and ASM-fitted landmarks</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/img_features.mat">img_features.mat</a>: image features pinna cavities extracted from both manually annotated and ASM-fitted landmarks</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/grnn_models.mat">grnn_models.mat</a>: Generalized Regression Neural Network (GRNN) models trained from both HUTUBS anthropometry and the proposed pinna features</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/predicted_dtf.mat">predicted_dtf.mat</a>: Directional Transfer Function (DTF) sets predicted from both HUTUBS anthropometry and the proposed pinna features</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/anthropometry_documentation.pdf">anthropometry_documentation.pdf</a>: documentation of the pinna anthropometric parameters</li> <li><a href="../api/files/bf20ab4d-c913-4cd0-9357-3b4039f1affb/auditory_model_complete_elevation_range.pdf">auditory_model_complete_elevation_range.pdf</a>: auditory model evaluation in the complete elevation range</li> </ul> <p>The data are provided in the Matlab file format MAT. Nevertheless, the MAT files can be read with other programming languages, such as Python (<a href="https://docs.scipy.org/doc/scipy/reference/generated/scipy.io.loadmat.html">scipy.io.loadmat</a>).</p> <p>A GitHub repository to automatically extract the pinna landmarks and features as described in the paper is available <a href="https://github.com/DavideFantini/pinna-landmarks-fitting-and-anthropometry-extraction" target="_blank" rel="noopener">here</a>.</p>
Data from: Immunotherapy-related adverse events (irAEs): extraction from FDA drug labels and comparative analysis
Open the record for dataset details and reuse information.
The Effects of a Food Product Containing Mushroom Extracts (AndoSanTM) in Subjects with Colorectal Cancer-related Fatigue
ClinicalTrials.gov study NCT06599710. IPD Sharing: NO. Countries: 1. Publications: 0.
Relate Tooth Alveolar Extraction Socket Anatomy to Alveolar Remodeling Rate
ClinicalTrials.gov study NCT01963884. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Age-related Changes in Retinal Oxygen Extraction in Healthy Subjects
ClinicalTrials.gov study NCT06643403. IPD Sharing: Not stated. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.