Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
54
datasets available to search
ShareScore release 0.7.1
Dataset results
54 results for “Text Mining”
Data from: Looking at cerebellar malformations through text-mined interactomes of mice and humans
We have generated and made publicly available two very large networks of molecular interactions: 49,493 mouse-specific and 52,518 human-specific interactions. These networks were generated through automated analysis of 368,331 full-text research articles and 8,039,972 article abstracts from the PubMed database, using the GeneWays system. Our networks cover a wide spectrum of molecular interactions, such as bind, phosphorylate, glycosylate, and activate; 207 of these interaction types occur more than 1,000 times in our unfiltered, multi-species data set. Because mouse and human genes are linked through an orthological relationship, human and mouse networks are amenable to straightforward, joint computational analysis. Using our newly generated networks and known associations between mouse genes and cerebellar malformation phenotypes, we predicted a number of new associations between genes and five cerebellar phenotypes (small cerebellum, absent cerebellum, cerebellar degeneration, abnormal foliation, and abnormal vermis). Using a battery of statistical tests, we showed that genes that are associated with cerebellar phenotypes tend to form compact network clusters. Further, we observed that cerebellar malformation phenotypes tend to be associated with highly connected genes. This tendency was stronger for developmental phenotypes and weaker for cerebellar degeneration.
Text Mining Dataset
<p>Text Mining Dataset</p>
MEASURING FLEXIBILITY: A TEXT-MINING APPROACH
<p>In creativity research, ideation flexibility, the ability to generate ideas by utilizing differing semantic categories, has long been the focus of investigation. Psychometric work to develop measurement procedures for flexibility has generally lagged behind fluency and originality. Here, we build from extant research on the measurement of originality using text-mining models to theoretically posit and then empirically validate a text-mining based method for measuring flexibility in verbal divergent thinking (DT) responses. The empirical validation of this method is accomplished in two studies. In the first study, we used the verbal form of the Torrance Test of Creative Thinking (TTCT) to demonstrate that our novel flexibility scoring method strongly and positively correlates with traditionally used TTCT flexibility scores. In the second study, we show that our text-mining based flexibility scoring method can achieve a high level of reliability and internal validity with Alternate Uses Task. In addition, flexibility scores correlated with theoretically relevant outside attributes including artistic expertise and Big 5 personality facets in the second study indicating criterion validity. To assist researchers in applying this flexibility scoring method and provide specific details about how to undertake an analysis of flexibility using text-mining models. </p>
Text-fig. 2. Čestmír Bůžek and Zlatko Kvaček during the field work in the Kristina Mine in 1964 (photo by František Holý). in A Review Of The Early Miocene Mastixioid Flora Of The Kristina Mine At Hrádek Nad Nisou In North Bohemia (The Czech Republic)
Text-fig. 2. Čestmír Bůžek and Zlatko Kvaček during the field work in the Kristina Mine in 1964 (photo by František Holý).
Text-fig. 1. Geographical position of the studied locality of Hrádek/N. (Kristina Mine) and other floras compared in detail. Symbols: 1. Hrádek/N. (Kristina Mine), 2. Bogatynia (Turów Mine), 3. Hartau, 4. Berzdorf, 5. Wiesa, 6. Libkovice Mb. of the Most Fm. (Most Basin), 7. Cypris Shale (Sokolov Basin), 8. Cypris Shale (Cheb Basin), 9. České Budějovice and Třeboň basins, 10. Wackersdorf, 11. Köflach (Oberdorf Mine). in A Review Of The Early Miocene Mastixioid Flora Of The Kristina Mine At Hrádek Nad Nisou In North Bohemia (The Czech Republic)
Text-fig. 1. Geographical position of the studied locality of Hrádek/N. (Kristina Mine) and other floras compared in detail. Symbols: 1. Hrádek/N. (Kristina Mine), 2. Bogatynia (Turów Mine), 3. Hartau, 4. Berzdorf, 5. Wiesa, 6. Libkovice Mb. of the Most Fm. (Most Basin), 7. Cypris Shale (Sokolov Basin), 8. Cypris Shale (Cheb Basin), 9. České Budějovice and Třeboň basins, 10. Wackersdorf, 11. Köflach (Oberdorf Mine).
Data from: Looking at cerebellar malformations through text-mined interactomes of mice and humans
Open the record for dataset details and reuse information.
Text-fig. 1. Redrawn mining-map of the "Einigkeits-Zeche" of Seifhennersdorf from 1850. in Siliceous Microfossils From The Oligocene Tripoli-Deposit Of Seifhennersdorf
Text-fig. 1. Redrawn mining-map of the "Einigkeits-Zeche" of Seifhennersdorf from 1850.
SIAM 2007 Text Mining Competition dataset
**Subject Area:** Text Mining **Description:** This is the dataset used for the SIAM 2007 Text Mining competition. This competition focused on developing text mining algorithms for document classification. The documents in question were aviation safety reports that documented one or more problems that occurred during certain flights. The goal was to label the documents with respect to the types of problems that were described. This is a subset of the Aviation Safety Reporting System (ASRS) dataset, which is publicly available. **How Data Was Acquired:** The data for this competition came from human generated reports on incidents that occurred during a flight. **Sample Rates, Parameter Description, and Format:** There is one document per incident. The datasets are in raw text format. All documents for each set will be contained in a single file. Each row in this file corresponds to a single document. The first characters on each line of the file are the document number and a tilde separats the document number from the text itself. **Anomalies/Faults:** This is a document category classification problem.
The Trend of DH Papers and Patent Applications for text or data mining
<p>it is a figure of paper "A Probe into Patentometrics in Digital Humanities".</p>
Fig. 1 in The phylogenomic revolution and its conceptual innovations: a text mining approach
Fig. 1 Selected trending concepts in phylogenomic research. Frequency represents the number of word instances divided by the total number of words published in a given year. Error bars represent ± 1 standard error,
Methodology Flowchart for Thematic Trends Analysis Using Text Mining
<p><strong>Figure: Methodology Flowchart for Thematic Trends Analysis Using Text Mining</strong><br>This flowchart outlines the methodology used to analyze thematic trends in academic research related to trade exhibitions through text mining techniques. The process begins with <strong>Data Collection</strong>, where academic articles are sourced from leading journals. <strong>Data Preprocessing</strong> involves cleaning and preparing the text data for further analysis. The methodology then branches into <strong>Keyword Extraction</strong> and <strong>LDA Topic Modeling</strong>. Keyword extraction identifies key themes and terms within the corpus (Manning, Raghavan, & Schütze, 2008), while LDA (Latent Dirichlet Allocation) is employed to discover latent thematic structures (Blei, Ng, & Jordan, 2003). These steps lead to <strong>Cluster Analysis</strong> to group similar articles based on the extracted keywords (Everitt et al., 2011) and <strong>Sentiment Analysis</strong> to determine the emotional tone associated with the identified themes (Pang & Lee, 2008). The final step is <strong>Data Visualization & Interpretation</strong>, where the results are visualized and insights are derived.</p>
Anomaly Detection with Text Mining
Many existing complex space systems have a significant amount of historical maintenance and problem data bases that are stored in unstructured text forms. The problem that we address in this paper is the discovery of recurring anomalies and relationships between problem reports that may indicate larger systemic problems. We will illustrate our techniques on data from discrepancy reports regarding software anomalies in the Space Shuttle. These free text reports are written by a number of different people, thus the emphasis and wording vary considerably. With Mehran Sahami from Stanford University, I'm putting together a book on text mining called "Text Mining: Theory and Applications" to be published by Taylor and Francis.
Application of text mining to build AOP-based mucus hypersecretion genesets and validation with in vitro and clinical samples
GEO Series GSE142100. Homo sapiens. 72 samples. Type: Expression profiling by high throughput sequencing.
A Dual-Filter Strategy Integrating CRISPR-based Target Screening and Text Mining for Hand-Foot Syndrome
GEO Series GSE297714. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.