Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Latent Dirichlet Allocation”
Sea-Ice Data Content Representation Based on Latent Dirichlet Allocation for Belgica Bank in Greenland
<p><strong>Data description</strong></p> <p>File type : -.npy (python numpy file)</p> <p>File content: Each file is a numpy array of size (number of 256x256 patches, 4096) indexed by id of the patch (each scene contains 6,400 patches, each patch has 4,096 micropatches of size 4x4, assigned one topic [1] per micropatch, resulting in 4,096 topics per patch). Each file has 4 months of observation. Array size is 25600 x 4096. We provide 6 files containing 24 months of observation (see the excel file for the Sentinel-1 ids) [2].</p> <p>Software to open with: Python</p> <p>Example code:</p> <p>import numpy</p> <p>Data= numpy.load(“filename_with_path”)</p> <p> </p> <p>Reference:</p> <p>1. C. Karmakar, C.O. Dumitru, G. Schwarz, and M. Datcu, “<em>Feature-Free Explainable Data Mining in SAR Images Using Latent Dirichlet Allocation</em>”, IEEE JSTARS, vol. 14, pp. 676-689, 2021.</p>
Data from: Public opinion in Japanese newspaper readers' posts under the prolonged COVID-19 infection spread 2019-2021: Contents analysis using Latent Dirichlet Allocation
<p><span>These data are based on free description from <span>readers' posts on Japanese hardcopy newspaper articles in the public domain</span>. </span><span>Upon searching for "coronavirus," we found 412 reader submissions during the April 7, 2020 to May 25, 2020, 521 during January 8, 2021 to March 21, 2021, 458 during the April 25, 2021 to June 6, 2021, and 519 during July 12, 2021 to September 30, 2021. </span><span>Topics of content were extracted by Latent Dirichlet Allocation, and then calculating the ratio of topic occurrence in each description.</span></p>
Data from: Public opinion in Japanese newspaper readers’ posts under the prolonged COVID-19 infection spread 2019-2021: Contents analysis using Latent Dirichlet Allocation
Open the record for dataset details and reuse information.
The Python code of the Latent Dirichlet Allocation analysis
Open the record for dataset details and reuse information.
Penerapan Latent Dirichlet Allocation dalam Analisis Konten Memasak pada Youtube Indonesia
<p>Youtube is a platform used to share videos. Since its founding in 2005, to date there are more than 20 million active users with 2 million videos uploaded every day. This is inseparable from the contribution of Indonesian content creators who also share videos with various types of content. One type of content that is loved by the Indonesian people is content about cooking or culinary belonging to food vloggers. In this study, the author wants to know what are the dominant topics related to the type of content created by food vloggers on their Youtube channel. This study uses the Latent Dirichlet Allocation (LDA) method. The research began by conducting text mining experiments on 3846 videos from the Youtube channel of nine food vloggers with more than one million subscribers. Then the LDA method is applied which is then analyzed to determine the optimal number of topics of content types by looking at the perplexity and topic coherence values. As a result, there are 5 topics of content types that often appear in videos belonging to food vloggers. Topics of this type of content include tips for making economical cakes, ingredients for making oven cakes, food business ideas, the process of cooking viral dishes, and easily available ingredients for making snacks.</p>
Fast discriminative latent Dirichlet allocation
This is the code for fast discriminative latent Dirichlet allocation, which is an algorithm for topic modeling and text classification. The related paper is at http://www-users.cs.umn.edu/~shan/icdm09_dm.pdf
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.