Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
55
datasets available to search
ShareScore release 0.7.1
Dataset results
55 results for “Natural language processing”
Data from: Natural language processing and recurrent network models for identifying genomic mutation-associated cancer treatment change from patient progress notes
Open the record for dataset details and reuse information.
Figure 13 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 13 Gold standard versus NER output.
Figure 12 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 12 An example of a specimen label.
Figure 10 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 10 Results per field from Google Cloud Vision.
Harnessing Natural Language Processing to Deciphers Ancient Temple Texts
<p>This paper explores the use of Natural Language Processing (NLP) and Optical Character Recognition (OCR) to digitize and translate ancient Indian temple inscriptions. By applying machine learning to ancient languages like Sanskrit and Tamil, the research offers methods to decode and reconstruct eroded texts. Semantic analysis tools are also used to uncover the cultural and historical significance of the inscriptions. This interdisciplinary approach combines data science with cultural preservation, providing new ways to understand ancient texts.</p>
An Open Internet-based Survey and Natural Language Processing Project Analysing Written Monologues by Headache Patients
ClinicalTrials.gov study NCT05153876. IPD Sharing: NO. Countries: 1. Publications: 0.
Detecting Delayed Discharge in Acute Geriatric Unit Using Natural Language Processing
ClinicalTrials.gov study NCT04965480. IPD Sharing: NO. Countries: 1. Publications: 0.
Natural Language Processing (NLP) Analysis of Free Text Notes to Investigate Coronavirus (COVID-19)
ClinicalTrials.gov study NCT04432961. IPD Sharing: Not stated. Countries: 1. Publications: 0.
NLP Headache Speech (NLPH-SPEECH): a Cross-sectional Study on Natural Language Processing Analysing Spoken Monologues by Headache Patients
ClinicalTrials.gov study NCT05204316. IPD Sharing: NO. Countries: 1. Publications: 0.
Development of a Natural Language Processing Tool to Enable Clinical Research in Emergency Medicine
ClinicalTrials.gov study NCT06240572. IPD Sharing: YES. Countries: 1. Publications: 0.
Sentiment Analysis using Natural Language Processing implemented in Large Language Model
Open the record for dataset details and reuse information.
Revealing Gender Biases in (TJSP) Court Decisions with Natural Language Processing
<p>Data derived from the realm of the social sciences is often produced in digital text form, which motivates its use as a source for natural language processing methods. Researchers and practitioners have developed and relied on artificial intelligence techniques to collect, process, and analyze documents in the legal field, especially for tasks such as text summarization and classification. In this scenario, we identify an underexplored potential of natural language processing used to delve into human rights issues in the context of artificial intelligence for social good. Qualitative and quantitative social science methods have been used to study matters such as institutional gender biasing in legal settings; however, natural language processing-based approaches can help analyze the issue on a larger scale. The work Revealing Gender Biases in Court Decisions with Natural Language Processing presents a protocol to address the automatic detection of institutional gender biasing in Brazilian courts, which comprises: (a) a pipeline of collection, annotation, and preparation of text extracted from court decisions issued by the São Paulo state Court of Justice in cases of domestic violence and parental alienation, which resulted in two datasets; (b) an experimental protocol of supervised binary classification over the decisions, performed with BERTimbau-based models; (c) methods for evaluating and validating such protocol.</p> <p>Here, we present the two datasets associated with this work: Dataset 1, made of 1,604 decisions issued by the Court between 2012 and 2019 in domestic violence-related criminal cases (DVC), and Dataset 2, made of 49 decisions issued by the Court in the same timeframe in civil and criminal parental alienation-related cases (PAC). Details on the content of each dataset, as well as their pipelines of extraction, annotation, preparation, and use, can be found in the original work, published as a Master's dissertation.</p> <p>The structure of the datasets is presented as follows:</p> <p>├── Dataset 1 (domestic violence cases, DVC): lesao.zip<br>│ ├── files<br>│ └── content<br>├── Dataset 2 (parental alienation cases, PAC): ap.zip<br>│ ├── files<br>│ └── content</p> <ul> <li><strong>files</strong> folder: contains input and output files associated with the pipelines of data extraction, annotation, and preparation as documented in the original work;</li> <li><strong>content</strong> folder: contains TXT and PDF files for each decision.</li> </ul> <p>Please note that, to access and use the datasets, one must abide to a deed of undertaking, whose violation entails legal liability of the breacher. Details on guidelines of legal and ethical compliance regarding this data can be found in the associated publications.</p>
Dataset and Code for the Research Paper: "A Smile is All You Need: Predicting Limiting Activity Coefficients from SMILES with Natural Language Processing."
<p><strong>Dataset and Code for the Research Paper: "A Smile is All You Need: Predicting Limiting Activity Coefficients from SMILES with Natural Language Processing."</strong></p> <p>For detailed instructions on how to utilize the resources, please refer to the <code>README.md</code> and see the original publication: https://doi.org/10.1039/D2DD00058J</p>
Dataset related to article "Connecting the use of innovative treatments and glucocorticoids with the multidisciplinary evaluation through rule-based natural-language processing: a real-world study on patients with rheumatoid arthritis, psoriatic arthritis, and psoriasis"
<p>This record contains raw data related to article "Connecting the use of innovative treatments and glucocorticoids with the multidisciplinary evaluation through rule-based natural-language processing: a real-world study on patients with rheumatoid arthritis, psoriatic arthritis, and psoriasis"</p><p>Abstract</p><p>Background: The impact of a multidisciplinary management of rheumatoid arthritis (RA), psoriatic arthritis (PsA), and psoriasis on systemic glucocorticoids or innovative treatments remains unknown. Rule-based natural language processing and text extraction help to manage large datasets of unstructured information and provide insights into the profile of treatment choices.</p><p>Methods: We obtained structured information from text data of outpatient visits between 2017 and 2022 using regular expressions (RegEx) to define elastic search patterns and to consider only affirmative citation of diseases or prescribed therapy by detecting negations. Care processes were described by binary flags which express the presence of RA, PsA and psoriasis and the prescription of glucocorticoids and biologics or small molecules in each cases. Logistic regression analyses were used to train the classifier to predict outcomes using the number of visits and the other specialist visits as the main variables.</p><p>Results: We identified 1743 patients with RA, 1359 with PsA and 2,287 with psoriasis, accounting for 5,677, 4,468 and 7,770 outpatient visits, respectively. Among these, 25% of RA, 32% of PsA and 25% of psoriasis cases received biologics or small molecules, while 49% of RA, 28% of PsA, and 40% of psoriasis cases received glucocorticoids. Patients evaluated also by other specialists were treated more frequently with glucocorticoids (70% vs. 49% for RA, 60% vs. 28% for PsA, 51% vs. 40% for psoriasis; p < 0.001) as well as with biologics/small molecules (49% vs. 25% for RA, 64% vs. 32% in PsA; 51% vs. 25% for psoriasis; p < 0.001) compared to cases seen only by the main specialist.</p><p>Conclusion: Patients with RA, PsA, or psoriasis undergoing multiple evaluations are more likely to receive innovative treatments or glucocorticoids, possibly reflecting more complex cases.</p><p> </p>
Unveiling the Sentiments and Opinions of Football Fans towards Video Assistant Referee (VAR) Technology: A Natural Language Processing Analysis
<p>The dataset has been used for the titled "Unveiling the Sentiments and Opinions of Football Fans towards Video Assistant Referee (VAR) Technology: A Natural Language Processing Analysis" provides valuable insights into the attitudes and opinions of football fans towards VAR technology. The dataset is the result of a natural language processing analysis of social media posts and online discussions related to VAR technology. It includes a comprehensive collection of sentiments, opinions, and attitudes expressed by football fans towards VAR technology, and provides researchers and analysts with a wealth of data to help them understand the public perception of this emerging technology in football. It can be used for further analysis in text analytics. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.