Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
68
datasets available to search
ShareScore release 0.9.0
Dataset results
68 results for “Topic Modeling”
Topic Model October 2023 (200 texts, 40 topics)
<p>Topic Model of the MiMoText roman18 corpus (Oct 2023)</p>
A Social and News Media Benchmark Dataset for Topic Modeling
<p>A novel approach to topic modeling using PSO-based clustering is compared to traditional techniques using the 20 Newsgroups dataset and a collection of posts from the Reddit health forum r/Cancer.</p>
Applying short text topic models to instant messaging communication of software developers
<p>Content related to paper "Applying short text topic models to instant messaging communication of software developers" published in the Journal of Systems and Software.</p> <p><strong>Data available:</strong></p> <ul> <li>Data sets used: JSONs with messages from Gitter chat rooms (downloaded with previous Gitter API - <a href="https://developer.gitter.im/docs/welcome" rel="nofollow">https://developer.gitter.im/docs/welcome</a>): <ul> <li>"Android.json"</li> <li>"ConsenSys.json"</li> <li>"WebpackDocs.json"</li> <li>"Jenkinsci.json"</li> <li>"Locomotive.json"</li> <li>"SpringSecurity.json"</li> <li>"Flutter.rar" - json file was compressed due to its size</li> <li>"GitterHQ.rar" - json file was compressed due to its size</li> <li>"Laravel.rar" - json file was compressed due to its size</li> </ul> </li> <li>"stopwords_list": Customized list of stop words</li> <li>"topics_sttm_results.csv": Topics obtained with each combination of model and corpus (both lemmatized and stemmed corpora)</li> <li>"intrusion_tasks.csv": Results of the survey for the Intrusion Tasks and its participants' background</li> <li>"topicnaming_tasks.csv": Results of the survey for the Topic Naming Tasks and its participants' background</li> <li>"intrinsic_metrics.csv": Scores of topic coherence metrics at topic level ('average' represents the score at model level)</li> <li>"topics_themes_chatrooms.csv": Results of the exercise described in Section 5.2 with the topics and themes identified in each of the 87 Gitter chat rooms.</li> <li>"sensitivity_analysis": Results of a smaller-scale sensitivity analysis to check the impact of the number of topics on the main findings of the paper.</li> </ul>
Mars topical ice model software [dataset]
<p>Data set used to model Mars topical ice spatial temporal distribution. NASA Grant 80NSSC19K1223.</p>
Clusters of topic modelling and naive Bayes classifier - FR CH newspapers
<p>Clusters of articles based on annotations produced by topic modelling and naive Bayes classifier applied to French language newspapers of Switzerland, published between 1900 and 1944 and containing the characters "europ", extracted from the impresso app. </p>
Topic modeling in software engineering research
<p>Raw data collected from 111 papers applying topic modeling techniques in software engineering studies.</p>
Figures from the paper "The diversity of canonical and ubiquitous progress in computer vision: A dynamic topic modeling approach"(v2))
<p>Figures from the paper "The diversity of canonical and ubiquitous progress in computer vision: A dynamic topic modeling approach".</p> <p><strong>The second version:</strong> Corrections to Figure 1. (fig1-> fig_v2).</p>
The data of "A comparison of citation-based clustering and topic modeling for science mapping"
<p>These files consist of the data used in "A comparison of citation-based clustering and topic modeling for science mapping". </p> <p> </p>
Topic modeling datasets
<p>Lemmatized texts of three datasets (Lenta, 20 Newsgroups, and WoS) used in the numerical experiments of work "Topic models with elements of neural networks: investigation of stability, coherence, and determining the optimal number of topics".</p>
Data from: Estimating genome-wide phylogenies using probabilistic topic modeling
Open the record for dataset details and reuse information.
Dataset: A structural topic model approach to scientific reorientation of economics and chemistry after German reunification
<p>Dataset to: A structural topic model approach to scientific reorientation of economics and chemistry after German reunification</p> <p> </p> <p>Please see readme.txt for details</p>
Characterizing Design Discussions With Semi-Supervised Topic Modeling
<p><strong>Note:</strong> Please refer to the README.md file for instructions.</p> <p><strong>Abstract:</strong> Stack Overflow is a rich source of questions and answers—discussions—about software development. One topic of discussion is software design, such as the correct use of design patterns, or best practices in data access. Since design is a more abstract topic in software engineering, researchers have long sought to characterize and model design knowledge. However, these approaches typically require significant expert input in order to contextualize the abstract design information. In this study, we explore how combining expert input with Stack Overflow might serve as an effective way to identify design topics. We first perform a qualitative analysis of design-tagged Stack Overflow questions and answers to identify the design concepts developers discuss. We report on areas where agreement was a challenge, including abstraction levels. Since inductive coding is expensive, we apply a semi-supervised (Anchored CorEx) approach. We find it performs as well as LDA but offers superior interpretability and the ability to guide the topic model. We leverage CorEx to characterize how design is discussed in Stack Overflow and on GitHub. We conclude by describing how our experience using the semi-supervised CorEx approach leads us to believe that approaches like CorEx that combine domain knowledge and scalability are key for analyzing large SE text repositories.</p>
Google Trends time series for the term "topic modeling"
<p>Dataset received from Google Trends for the phrase "topic modeling'' on 31 January 2024 using the URL <a href="https://trends.google.de/trends/explore?date=all&q=topic\%20modeling&hl=de">https://trends.google.de/trends/explore?date=all&q=topic\%20modeling&hl=de</a></p> <p><em>Data obtained from Google LLC, which is the ultimate owner of these data. Published for academic and non-commercial replication purposes only.</em></p>
Topic Modeling The Red Pill - Datasets
<p>LDA and word2vec models trained on the complete works of Return Of Kings' blog. Built using Gensim in Python. Analysis and notebooks are found here: https://github.com/ChamRoshi/Topic-Modelling-The-Red-Pill</p>
Global Transcriptomic Analysis of Topical Sodium Alginate Protection Against Peptic Damage in An In Vitro Model of Treatment-Resistant Gastroesophageal Reflux Disease
<p>PA= pepsin + Acid; "Sham + PA" means "Pretreatment + Treatment"</p> <p><span>Breakthrough symptoms </span>are thought to occur in roughly half of <span>all </span>gastroesophageal reflux disease (GERD) patients despite maximal acid suppression (proton pump inhibitor, PPI) therapy. Topical alginates have recently been shown to enhance mucosal defense against acid-pepsin insult during GERD. We aimed to examine potential alginate protection of transcriptomic changes in a cell culture model of PPI recalcitrant GERD. Immortalized normal-derived human esophageal epithelial cells underwent pretreatment with commercial alginate-based anti-reflux medications (Gaviscon Advance or Gaviscon Double Action), a matched-viscosity placebo control, or pH 7.4 buffer (sham) alone for 1 minute, followed by exposure to pH 6.0+pepsin or buffer alone for 3 minutes. RNA sequencing was conducted, and Ingenuity Pathway Analysis was performed with a false discovery rate of ≤0.01, and absolute fold-change of ≥<span>1.3. Pepsin-acid exposure disrupted gene expressions associated with epithelial barrier function, chromatin structure</span>, carcinogenesis, and inflammation<span>. Alginate formulations demonstrated protection by mitigating these changes and promoting extracellular matrix repair, downregulating proto-oncogenes, and enhancing tumor suppressor expression. </span>These data suggest molecular mechanisms by which alginates provide topical protection against injury during weakly acidic reflux and support a potential role for alginates in prevention of GERD-related carcinogenesis.</p>
Topic modeling in software engineering research
<p>Data extracted from 111 papers applying topic modeling techniques in software engineering studies.</p>
Topic model of English-language fiction, 1880-1999, with 200 topics.
<p>A topic model of 29,341 volumes of fiction, written in English and published between 1880 and 1999. The underlying corpus was organized by Ted Underwood for an experiment on period and cohort effects in cultural change. Metadata is in finalcorpus.tsv (which also has rows for 10 volumes not actually included in the model). </p> <p>To identify volumes as fiction, we relied on the NovelTM Dataset of English-Language Fiction (https://culturalanalytics.org/article/13147-noveltm-datasets-for-english-language-fiction-1700-2009). To confirm birth years of authors and publication dates of books, we compared NovelTM metadata both to the Chicago Novel Corpus and to a copy of the US Copyright Registry, digitized by the New York Public Library (https://github.com/NYPL/catalog_of_copyright_entries_project).</p> <p>The corpus itself is in cohort4.txt.gz; each line represents a roughly 10,000-word "chunk" of a document. The first 15% and last 5% of pages in each volume were discarded; the remaining pages were divided into chunks of roughly equal size. Chunk id is the first token on each line; it is formed by taking a HathiTrust volume id and adding an underscore + sequential integer (chunk number). Removing the underscore and integer produces a "document id" that can be paired to the metadata. The words in the line are not presented in original order; they are taken from HathiTrust Extracted Features, which records only page-level word counts.</p> <p>The topic model was produced using MALLET (http://mallet.cs.umass.edu/index.php), and has 200 topics.</p> <p>The top words in each topic are listed in the "keys" file; document-topic proportions are listed in "doctopics."</p> <p>For more information on the construction of the corpus and the experiment it is designed to support, see https://github.com/tedunderwood/period-cohort and/or a permanent Zenodo object created from that repository.</p>
Data for dissertation titled 'Topic modelling for the stratification of neurological patients'
<p>The uploaded zip-file entails the data needed for and obtained through the dissertation titled 'Topic modelling for the stratification of neurological patients' as part of the programme 'MSc. in Statistical Data Analysis' at Ghent University. The study aimed at exploring the applicability of hierarchical stochastic block models on resting-state functional magnetic resonance imaging (RS-fMRI) to cluster participants with known neurological disorders.</p> <p>The data is structured in different folders and aligns with the folder structure of the GitHub-repository that contains the analysis scripts (https://github.com/wvechelp/hsbm_on_fmri). The GitHub-repository already covers some example data, while additional data and results can be found in this zip-file. Additional comments on the analyses are also provided in the analysis scripts.</p> <p>The original raw data is obtained through the OpenFMRI project (<a href="http://openfmri.org/">http://openfmri.org/</a>, with label <em>ds000030</em>) and as a Stanford Digital Repository (<a href="https://purl.stanford.edu/mg599hw5271">https://purl.stanford.edu/mg599hw5271</a>) (Bilder<em> et al.</em>, 2016). It is obtained from the NIH Roadmap Initiative, as a result of the Consortium for Neuropsychiatric Phenomics (CNP) study (Poldrack<em> et al.</em>, 2016). Throughout the study, data was collected through interviews and rating scales, self-report measures, neurocognitive exams (using both paper-pencil and computerised tests), and a variety of neuroimaging data. More specific information on the selection procedure of the participants can be found in the description of the data (Bilder<em> et al.</em>, 2016) and the associated article (Poldrack<em> et al.</em>, 2016). Among the available neuroimaging data, the RS-fMRI data have been collected by asking participants to remain relaxed, while keeping their eyes open (with scans lasting 304 s and an image being collected every 2 seconds (Poldrack<em> et al.</em>, 2016)). This raw data was pre-processed by Rasero<em> et al.</em> (2019) to correct for motion and temporal alignment. Smoothing (6-mm full width at half-maximum Gaussian kernel), intensity normalisation, and a band-pass filter (between 0.01 and 0.08 Hz) were applied prior to the removal of linear and quadratic trends. Motion time courses, average CSF signal, and the average white matter signal were regressed out prior to data transformation into voxels with a volume of 3 mm x 3 mm x 3 mm. Ultimately, the functional atlas of Shen<em> et al.</em> (2013) was used to average the voxel signals per anatomical region of interest (ROI), resulting in a parcellation of 278 ROIs (and associated time series consisting of 152 measurements). From these ROI-specific time series, Rasero<em> et al.</em> (2019) generated 278 x 278 matrices with Pearson coefficients.</p> <p>References:<br> Bilder, R. M., Poldrack, R. A., Cannon, T., London, E., Freimer, N., Congdon, E., Karlsgodt, K., & Sabb, F. W. (2016). <em>UCLA Consortium for Neuropsychiatric Phenomics LA5c Study</em> Stanford Digital Repository. <a href="http://purl.stanford.edu/mg599hw5271">http://purl.stanford.edu/mg599hw5271</a> and <a href="https://openfmri.org/dataset/ds000030/">https://openfmri.org/dataset/ds000030/</a><br> Poldrack, R. A., Congdon, E., Triplett, W., Gorgolewski, K. J., Karlsgodt, K. H., Mumford, J. A., Sabb, F. W., Freimer, N. B., London, E. D., Cannon, T. D., & Bilder, R. M. (2016). A phenome-wide examination of neural and cognitive function. <em>Scientific Data</em>,<em> 3</em>(1), 160110. <a href="https://doi.org/10.1038/sdata.2016.110">https://doi.org/10.1038/sdata.2016.110</a><br> Rasero, J., Diez, I., Cortes, J. M., Marinazzo, D., & Stramaglia, S. A.-O. (2019). Connectome sorting by consensus clustering increases separability in group neuroimaging studies. <em>Network Neuroscience</em>,<em> 3</em>(2), 325-343. <a href="https://doi.org/https:/doi.org/10.1162/netn_a_00074">https://doi.org/https://doi.org/10.1162/netn_a_00074</a><br> Shen, X., Tokoglu, F., Papademetris, X., & Constable, R. T. (2013). Groupwise whole-brain parcellation from resting-state fMRI data for network node identification. <em>NeuroImage</em>,<em> 82</em>, 403-415. <a href="https://doi.org/https:/doi.org/10.1016/j.neuroimage.2013.05.081">https://doi.org/https://doi.org/10.1016/j.neuroimage.2013.05.081</a></p> <p> </p>
Supplementary material 1 from: Ancin-Murguzur FJ, Hausner VH (2020) Research gaps and trends in the Arctic tundra: a topic-modelling approach. One Ecosystem 5: e57117. https://doi.org/10.3897/oneeco.5.e57117
Supplementary table 1
Figure 7 from: Almudaris SA, Gatea FK (2024) Effects of topical Ivermectin on imiquimod-induced Psoriasis in mouse model – Novel findings. Pharmacia 71: 1-14. https://doi.org/10.3897/pharmacia.71.e114753
Figure 7 Histopathological section of mice skin (clobetasol control group) showing hyperkeratosis (black arrow), absence of parakeratosis & Munro's abscess. And epidermal granular layer (yellow arrow) with mild acanthosis, mild papillary thinning, and few rete ridges (green arrow). The dermis shows mild lymphocytic infiltrate. H&E stain (4×,10×).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.