Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

54

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

54 results for “Text Mining”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Looking at cerebellar malformations through text-mined interactomes of mice and humans

We have generated and made publicly available two very large networks of molecular interactions: 49,493 mouse-specific and 52,518 human-specific interactions. These networks were generated through automated analysis of 368,331 full-text research articles and 8,039,972 article abstracts from the PubMed database, using the GeneWays system. Our networks cover a wide spectrum of molecular interactions, such as bind, phosphorylate, glycosylate, and activate; 207 of these interaction types occur more than 1,000 times in our unfiltered, multi-species data set. Because mouse and human genes are linked through an orthological relationship, human and mouse networks are amenable to straightforward, joint computational analysis. Using our newly generated networks and known associations between mouse genes and cerebellar malformation phenotypes, we predicted a number of new associations between genes and five cerebellar phenotypes (small cerebellum, absent cerebellum, cerebellar degeneration, abnormal foliation, and abnormal vermis). Using a battery of statistical tests, we showed that genes that are associated with cerebellar phenotypes tend to form compact network clusters. Further, we observed that cerebellar malformation phenotypes tend to be associated with highly connected genes. This tendency was stronger for developmental phenotypes and weaker for cerebellar degeneration.

opencc-zeroDec 2011View details →
zenodo28/100

Text Mining Dataset

<p>Text Mining Dataset</p>

opencc-by-4.0Jan 2022View details →
zenodo28/100

MEASURING FLEXIBILITY: A TEXT-MINING APPROACH

<p>In creativity research, ideation flexibility, the ability to generate ideas by utilizing differing semantic categories, has long been the focus of investigation. Psychometric work to develop measurement procedures for flexibility has generally lagged behind fluency and originality. Here, we build from extant research on the measurement of originality using text-mining models to theoretically posit and then empirically validate a text-mining based method for measuring flexibility in verbal divergent thinking (DT) responses. The empirical validation of this method is accomplished in two studies. In the first study, we used the verbal form of the Torrance Test of Creative Thinking (TTCT) to demonstrate that our novel flexibility scoring method strongly and positively correlates with traditionally used TTCT flexibility scores. In the second study, we show that our text-mining based flexibility scoring method can achieve a high level of reliability and internal validity with Alternate Uses Task. In addition, flexibility scores correlated with theoretically relevant outside attributes including artistic expertise and Big 5 personality facets in the second study indicating criterion validity. To assist researchers in applying this flexibility scoring method and provide specific details about how to undertake an analysis of flexibility using text-mining models.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

Text-fig. 2. Čestmír Bůžek and Zlatko Kvaček during the field work in the Kristina Mine in 1964 (photo by František Holý). in A Review Of The Early Miocene Mastixioid Flora Of The Kristina Mine At Hrádek Nad Nisou In North Bohemia (The Czech Republic)

Text-fig. 2. Čestmír Bůžek and Zlatko Kvaček during the field work in the Kristina Mine in 1964 (photo by František Holý).

opencc-by-4.0Dec 2012View details →
zenodo28/100

Text-fig. 1. Geographical position of the studied locality of Hrádek/N. (Kristina Mine) and other floras compared in detail. Symbols: 1. Hrádek/N. (Kristina Mine), 2. Bogatynia (Turów Mine), 3. Hartau, 4. Berzdorf, 5. Wiesa, 6. Libkovice Mb. of the Most Fm. (Most Basin), 7. Cypris Shale (Sokolov Basin), 8. Cypris Shale (Cheb Basin), 9. České Budějovice and Třeboň basins, 10. Wackersdorf, 11. Köflach (Oberdorf Mine). in A Review Of The Early Miocene Mastixioid Flora Of The Kristina Mine At Hrádek Nad Nisou In North Bohemia (The Czech Republic)

Text-fig. 1. Geographical position of the studied locality of Hrádek/N. (Kristina Mine) and other floras compared in detail. Symbols: 1. Hrádek/N. (Kristina Mine), 2. Bogatynia (Turów Mine), 3. Hartau, 4. Berzdorf, 5. Wiesa, 6. Libkovice Mb. of the Most Fm. (Most Basin), 7. Cypris Shale (Sokolov Basin), 8. Cypris Shale (Cheb Basin), 9. České Budějovice and Třeboň basins, 10. Wackersdorf, 11. Köflach (Oberdorf Mine).

opencc-by-4.0Dec 2012View details →
dryad28/100

Data from: Looking at cerebellar malformations through text-mined interactomes of mice and humans

Open the record for dataset details and reuse information.

publicAug 2012View details →
zenodo24/100

Text-fig. 1. Redrawn mining-map of the "Einigkeits-Zeche" of Seifhennersdorf from 1850. in Siliceous Microfossils From The Oligocene Tripoli-Deposit Of Seifhennersdorf

Text-fig. 1. Redrawn mining-map of the "Einigkeits-Zeche" of Seifhennersdorf from 1850.

opencc-by-4.0Dec 2007View details →
nasa24/100

SIAM 2007 Text Mining Competition dataset

**Subject Area:** Text Mining **Description:** This is the dataset used for the SIAM 2007 Text Mining competition. This competition focused on developing text mining algorithms for document classification. The documents in question were aviation safety reports that documented one or more problems that occurred during certain flights. The goal was to label the documents with respect to the types of problems that were described. This is a subset of the Aviation Safety Reporting System (ASRS) dataset, which is publicly available. **How Data Was Acquired:** The data for this competition came from human generated reports on incidents that occurred during a flight. **Sample Rates, Parameter Description, and Format:** There is one document per incident. The datasets are in raw text format. All documents for each set will be contained in a single file. Each row in this file corresponds to a single document. The first characters on each line of the file are the document number and a tilde separats the document number from the text itself. **Anomalies/Faults:** This is a document category classification problem.

restrictednotspecifiedApr 2025View details →
zenodo20/100

The Trend of DH Papers and Patent Applications for text or data mining

<p>it is a figure of paper &quot;A Probe into Patentometrics in Digital Humanities&quot;.</p>

opencc-by-4.0Feb 2020View details →
zenodo20/100

Fig. 1 in The phylogenomic revolution and its conceptual innovations: a text mining approach

Fig. 1 Selected trending concepts in phylogenomic research. Frequency represents the number of word instances divided by the total number of words published in a given year. Error bars represent ± 1 standard error,

opennotspecifiedMar 2019View details →
zenodo20/100

Methodology Flowchart for Thematic Trends Analysis Using Text Mining

<p><strong>Figure: Methodology Flowchart for Thematic Trends Analysis Using Text Mining</strong><br>This flowchart outlines the methodology used to analyze thematic trends in academic research related to trade exhibitions through text mining techniques. The process begins with <strong>Data Collection</strong>, where academic articles are sourced from leading journals. <strong>Data Preprocessing</strong> involves cleaning and preparing the text data for further analysis. The methodology then branches into <strong>Keyword Extraction</strong> and <strong>LDA Topic Modeling</strong>. Keyword extraction identifies key themes and terms within the corpus (Manning, Raghavan, &amp; Sch&uuml;tze, 2008), while LDA (Latent Dirichlet Allocation) is employed to discover latent thematic structures (Blei, Ng, &amp; Jordan, 2003). These steps lead to <strong>Cluster Analysis</strong> to group similar articles based on the extracted keywords (Everitt et al., 2011) and <strong>Sentiment Analysis</strong> to determine the emotional tone associated with the identified themes (Pang &amp; Lee, 2008). The final step is <strong>Data Visualization &amp; Interpretation</strong>, where the results are visualized and insights are derived.</p>

restrictedmit-licenseAug 2024View details →
nasa20/100

Anomaly Detection with Text Mining

Many existing complex space systems have a significant amount of historical maintenance and problem data bases that are stored in unstructured text forms. The problem that we address in this paper is the discovery of recurring anomalies and relationships between problem reports that may indicate larger systemic problems. We will illustrate our techniques on data from discrepancy reports regarding software anomalies in the Space Shuttle. These free text reports are written by a number of different people, thus the emphasis and wording vary considerably. With Mehran Sahami from Stanford University, I'm putting together a book on text mining called "Text Mining: Theory and Applications" to be published by Taylor and Francis.

restrictednotspecifiedMar 2025View details →
geo16/100

Application of text mining to build AOP-based mucus hypersecretion genesets and validation with in vitro and clinical samples

GEO Series GSE142100. Homo sapiens. 72 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2020View details →
geo16/100

A Dual-Filter Strategy Integrating CRISPR-based Target Screening and Text Mining for Hand-Foot Syndrome

GEO Series GSE297714. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record