Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
54
datasets available to search
ShareScore release 0.7.1
Dataset results
54 results for “Text Mining”
Text-fig. 2. (A) Orientation of Schizocrania filosa (HALL, 1847) on articulated shells of benthic brachiopod Rafinesquina sp. (A1–A3 – on dorsal valve of articulated shells, A4 – on ventral valve of articulated shell; forward growth direction is unclear in three specimens) from Upper Ordovician, Corryville Formation, Lawrenceburg, Indiana (after www.drydredgers.org/scizo.htm). (B) Orientation of Schizocrania multistriata (REED, 1905) shells on outer face of conulariid Metaconularia imperialis test (Dobrotivá Formation, Kařízek mine, Barrandian area; after Havlíček and Vaněk 1996); preserved conulariid shell in white, suggested outline of incomplete conulariid test in grey. Arrows indicate direction of forward growth of Schizocrania specimens. in Schizocrania (Brachiopoda, Discinoidea): Taxonomy, Occurrence, Ecology And History Of The Earliest Epizoan Lingulate Brachiopod
Text-fig. 2. (A) Orientation of Schizocrania filosa (HALL, 1847) on articulated shells of benthic brachiopod Rafinesquina sp. (A1–A3 – on dorsal valve of articulated shells, A4 – on ventral valve of articulated shell; forward growth direction is unclear in three specimens) from Upper Ordovician, Corryville Formation, Lawrenceburg, Indiana (after www.drydredgers.org/scizo.htm). (B) Orientation of Schizocrania multistriata (REED, 1905) shells on outer face of conulariid Metaconularia imperialis test (Dobrotivá Formation, Kařízek mine, Barrandian area; after Havlíček and Vaněk 1996); preserved conulariid shell in white, suggested outline of incomplete conulariid test in grey. Arrows indicate direction of forward growth of Schizocrania specimens.
Financial News dataset for text mining
<p>please cite this dataset by :</p> <p>Nicolas Turenne, Ziwei Chen, Guitao Fan, Jianlong Li, Yiwen Li, Siyuan Wang, Jiaqi Zhou (2021) Mining an English-Chinese parallel Corpus of Financial News, BNU HKBU UIC, technical report</p> <p> </p> <p>The dataset comes from Financial Times news website (https://www.ft.com/)</p> <p>news are written in both languages Chinese and English.</p> <p><a href="https://zenodo.org/api/files/74136928-3c77-4388-aa47-b1079efa1650/FTIE.zip?versionId=440c065e-1528-4f01-800d-541a0611ea7f">FTIE.zip</a> contains all documents in a file individually</p> <p><a href="https://zenodo.org/api/files/74136928-3c77-4388-aa47-b1079efa1650/FT-en-zh.rar">FT-en-zh.rar</a> contains all documents in one file</p> <p>Below is a sample document in the dataset defined by these fields and syntax : </p> <p>id;time;english_title;chinese_title;integer;english_body;chinese_body</p> <p> </p> <p>1021892;2008-09-10T00:00:00Z;FLAW IN TWIN TOWERS REVEALED;科学家发现纽约双子塔倒塌的根本原因;1;Scientists have discovered the fundamental reason the Twin Towers collapsed on September 11 2001. The steel used in the buildings softened fatally at 500?C – far below its melting point – as a result of a magnetic change in the metal. @ The finding, announced at the BA Festival of Science in Liverpool yesterday, should lead to a new generation of steels capable of retaining strength at much higher temperatures.;科学家发现了纽约世贸双子大厦(Twin Towers)在2001年9月11日倒塌的根本原因。由于磁性变化,大厦使用的钢在500摄氏度——远远低于其熔点——时变软,从而产生致命后果。 @ 这一发现在昨日利物浦举行的BA科学节(BA Festival of Science)上公布。这应会推动能够在更高温度下保持强度的新一代钢铁的问世。<br> </p> <p>The dataset contains 60,473 bilingual documents.</p> <p>Time range is from 2007 and 2020. </p> <p>This dataset has been used for parallel bilingual news mining in Finance domain.</p>
'Text mining' in the International ADHO Digital Humanities Conferences (2006-2020)
<p>Curated dataset with metadata for all titles that include the English term 'text mining' in CSV format. Full Zotero collection available at: <a href="https://www.zotero.org/silviaegt/collections/SP2JK58V">https://www.zotero.org/silviaegt/collections/SP2JK58V</a></p> <p>Useful to this search was: Eichmann-Kalwara, N., Weingart, S.B., Lincoln, M., et al. <em>The Index of Digital Humanities Conferences</em>. Carnegie Mellon University, 2020. <a href="https://dh-abstracts.library.cmu.edu/">https://dh-abstracts.library.cmu.edu</a>. <a href="https://doi.org/10.34666/k1de-j489">https://doi.org/10.34666/k1de-j489</a></p>
Data and Python script for article "A text mining analysis of the climate change literature in industrial ecology'
<p>The data and Python script are part of the forum article "A text mining analysis of the climate change literature in industrial ecology" authored by Dayeen, F.R., Sharma, A.S., and Derrible, S., and published in the <em>Journal of Industrial Ecology</em> in 2020.</p> <p>The Python script and instructions are included in the LiTCoF_v1.00-py.zip file. The original data is available in two formats: .csv and .pkl.</p> <p>Updates of the script will be posted at https://github.com/csunlab/LiTCoF and at https://csun.uic.edu/codes/LiTCoF.html. The data is also available at https://csun.uic.edu/datasets.html#AbstractsIE.</p> <p>Feel free to contact any of the authors for information and questions about the data and code.</p>
Text Mining of Archaeological Reports for Urban farming (data and code)
<p>This release is created to create a DOI in Zenodo for the data related to Fischer, AD, van Londen, H, Blonk-van den Bercken, AL, Visser, RM and Renes, J. 2021. Urban farming and ruralisation in the Netherlands (1250 up tot the nineteenth century), unravelling farming practice and the use of (open) space by synthesising archaeological reports using text mining. Nederlandse Archeologische Rapporten 68. Amersfoort: Rijksdienst voor het Cultureel Erfgoed. <a href="https://www.cultureelerfgoed.nl/publicaties/publicaties/2021/01/01/urban-farming-and-ruralisation-in-the-netherlands">https://www.cultureelerfgoed.nl/publicaties/publicaties/2021/01/01/urban-farming-and-ruralisation-in-the-netherlands</a></p>
Decoding diabetes biomarkers and related molecular mechanisms using machine learning, text mining, and gene expression analysis
<p>The molecular basis of diabetes mellitus is yet to be fully elucidated. We aimed to identify the most frequently reported and differential expressed genes (DEGs) in diabetes using bioinformatics approaches. Text mining was used to screen 40,225 article abstracts from diabetes literature. These studies highlighted 5939 diabetes-related genes spread across 22 human chromosomes, with 112 genes mentioned in more than 50 studies. Among these genes, HNF4A, PPARA, VEGFA, TCF7L2, HLA- DRB1, PPARG, NOS3, KCNJ11, PRKAA2, and HNF1A were mentioned in more than 200 articles. These genes are correlated with the regulation of glycogen and polysaccharide, adipogenesis, AGE/RAGE, and macrophage differentiation. Three datasets (44 patients and 57 controls) were subjected to gene expression analysis. The analysis revealed 135 significant DEGs, of which CEACAM6, ENPP4, HDAC5, HPCAL1, PARVG, STYXL1, VPS28, ZBTB33, ZFP37 and CCDC58 were the top ten DEGs. These genes were enriched in aerobic respiration, T-Cell antigen receptor pathway, Tricarboxylic acid metabolic process, vitamin D receptor pathway, Toll-like receptor signaling, and endoplasmic reticulum (ER) unfolded protein response. The results of text mining and gene expression analyses used as attribute values for ML analysis . The "Decision tree", "Extra-tree regressor" and "Random forest" algorithms were used in ML analysis to identify unique markers that could be used as diabetes diagnosis tools. These algorithms produced prediction models with accuracy ranges from 0.6364 to 0.88 and overall confidence interval (CI) of 95%. There were 39 biomarkers that could distinguish diabetic and non-diabetic patients, 12 of which were repeated multiple times. The majority of these genes are associated with stress response, signalling regulation, locomotion, cell motility, growth, and muscle adaptation. ML algorithms highlighted the use of the HLA-DQB1 gene as a biomarker for diabetes early detection. Our data mining and gene expression analysis have provided useful information about potential biomarkers in diabetes.</p>
Text Mining as a Support Tool for Research on Climate Change: Theoretical and Technical Considerations
<p>839 companies listed on the Australian Stock Exchange were surveyed for climate risk disclosures, with 201 such disclosures identified.</p>
Appendix-C: Text Data and Mining Licensing Conditions
<p>Appendix C is associated with <em><strong>Chapter 11: Text Data and Mining Ethics</strong></em> of the book -- Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>
Appendix B: Language Corpora Available for Text Mining
<p>Appendix B is associated with <em><strong>Chapter 3: Text Pre-Processing</strong></em> of the book -- Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>
Appendix-A: Online Repositories Available for Text Mining
<p>Appendix A is associated with Chapter 2: Text data and where to find them of the book: Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>
Research exceptions in copyright laws around the world in Legal reform to enhance global text and data mining research.
Research exceptions in copyright laws around the world
Performance of GPT-4o mini and GPT-4o for medical text mining tasks at different temperature settings
Open the record for dataset details and reuse information.
Antibody Watch: Text Mining Antibody Specificity from the Literature
<p><strong>Abstract </strong></p> <p><strong>Motivation</strong>: Antibodies are widely used reagents to test for expression of proteins. However, they might not always reliably produce results when they do not specifically bind to the target proteins that their providers designed them for, leading to unreliable research results.</p> <p><strong>Results</strong>: We developed a deep neural network system and tested its performance with a corpus of more than two thousand articles that reported uses of antibodies. We divided the problem into two tasks. Given an input article, the first task is to identify snippets about antibody specificity and classify if the snippets report any antibody that is nonspecific, and thus problematic. The second task is to link each of these snippets to one or more antibodies that the snippet referred to. We leveraged the Research Resource Identifiers (RRID) to precisely identify antibodies linked to the extracted specificity snippets. The result shows that it is feasible to construct a reliable knowledge base about problematic antibodies by text mining.</p>
Biomedical ELECTRA based deep language representation models for biomedical text mining.
<p>The gzipped tar file contains two biomedical language representation models based on ELECTRA (Clark et al., 2020) deep transformers architecture to be used for down-stream biomedical text mining tasks. </p> <p>Bio-ELECTRA is pre-trained from scratch on PubMed abstracts for 1.8 million steps. Bio-ELECTRA++ is the further pre-trained version of Bio-ELECTRA trained on a corpus of open access full papers from PubMed.</p>
Data from: Using text-mined trait data to test for cooperate-and-radiate co-evolution between ants and plants
Mutualisms may be "key innovations" that spur lineage diversification by augmenting niche breadth, geographic range, or population size, thereby increasing speciation rates or decreasing extinction rates. Whether mutualism accelerates diversification in both interacting lineages is an open question. Research suggests that plants that attract ant mutualists have higher diversification rates than non-ant associated lineages. We ask whether the reciprocal is true: does the interaction between ants and plants also accelerate diversification in ants, i.e. do ants and plants cooperate-and-radiate? We used a novel text-mining approach to determine which ant species associate with plants in defensive or seed dispersal mutualisms. We investigated patterns of lineage diversification across a recent ant phylogeny using BiSSE, BAMM, and HiSSE models. Ants that associate mutualistically with plants had elevated diversification rates compared to non-mutualistic ants in the BiSSE model, with a similar trend in BAMM, suggesting ants and plants cooperate-and-radiate. However, the best-fitting model was a HiSSE model with a hidden state, meaning that diversification models that do no account for unmeasured traits are inappropriate to assess the relationship between mutualism and ant diversification. Against a backdrop of diversification rate heterogeneity, the best-fitting HiSSE model found that mutualism actually decreases diversification: mutualism evolved much more frequently in rapidly diversifying ant lineages, but then subsequently slowed diversification. Thus, it appears that ant lineages first radiated, then cooperated with plants.
RegEl Database: text-mined regulatory elements from the literature and their associations to genes and disease
<pre>@article{garda2022regel, title={RegEl corpus: identifying DNA regulatory elements in the scientific literature}, author={Garda, Samuele and Lenihan-Geels, Freyda and Proft, Sebastian and Hochmuth, Stefanie and Sch{\"u}lke, Markus and Seelow, Dominik and Leser, Ulf}, journal={Database}, volume={2022}, year={2022}, publisher={Oxford Academic} } </pre> <p># RegEl PubMed Database</p> <p>This database contains the annotations generated by running [HunFlair](https://github.com/flairNLP/flair/blob/master/resources/docs/HUNFLAIR.md) models trained on the [RegEl corpus](https://zenodo.org/record/5776679) over >20M PubMed abstracts.</p> <p>By pairing these annotations with the one provided by PubTator this generates a large text mining database of regulatory elements associated with genes (normalized to NCBI Gene ids) and disease (normalized to either MeSH or OMIM).</p> <p>The tables composing the database are:</p> <p>* abstracts.db:<br> - pmid = PubMed ID of the given abstracts<br> - sid = sentence ID of the given abstracts (from 0 to # of sentences)<br> - text = text of the given sentence</p> <p>* gene.db and disease.db:<br> - pmid = PubMed ID of the given abstracts<br> - sid = sentence ID of the given abstracts (from 0 to # of sentences)<br> - etype = entity type (enhancer, promoter, TFBS)<br> - ann_text = mention of the regulatory element as found in the abstract<br> - start = position (# character) in which the mention begins<br> - end = position (# characters) in which the mention ends<br> - score = model's confidence<br> - cui = gene or disease identifier<br> - cui_symbol = official symbol of cui (if available)</p>
Assessing and predicting the quality of peer reviews: a text mining approach
<p>Dataset</p>
Big Data and Text-mining Technologies Applied for Breast Cancer Medical Data Analysis
ClinicalTrials.gov study NCT02810093. IPD Sharing: NO. Countries: 1. Publications: 1.
Data from: Using text-mined trait data to test for cooperate-and-radiate co-evolution between ants and plants
Open the record for dataset details and reuse information.
Semantic text mining in early drug discovery for type 2 diabetes
<p>BACKGROUND: Surveying the scientific literature is an important part of early drug discovery; and with the ever-increasing amount of biomedical publications it is imperative to focus on the most interesting articles. Here we present a project that highlights new understanding (e.g.\ recently discovered modes of action) and identifies potential novel drug target, via a novel, data-driven text mining approach to score type 2 diabetes (T2D) relevance. We focused on monitoring trends and jumps in T2D relevance to help us be timely informed of important breakthroughs.<br> <br> METHODS: We extracted over 7 million <em>n</em>-grams from PubMed and then clustered around 240,000 linked to T2D into almost 50,000 T2D relevant `semantic concepts'. To score papers, these concepts were weighted depending on co-mentioning with core T2D proteins. A protein's current T2D relevance was determined by combining the scores of the papers mentioning it in the preceeding five years. The significance of a jump in a protein's rank was assessed by comparing it to previously observed jumps.<br> <br> RESULTS: We show that T2D relevant papers, also those not mentioning T2D explicitly, got assigned high scores by mentioning semantic concepts often used in connection with T2D, as shown by the enrichment of well known T2D proteins among the top scoring proteins. Our `high jumpers' identified important past developments in the apprehension of how certain key proteins relate to T2D, indicating that our method will make us aware of future breakthroughs. In summary, this project facilitated keeping up with current T2D research by repeatedly providing short lists of potential novel targets into our early drug discovery pipeline.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.