Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
24
datasets available to search
ShareScore release 0.7.1
Dataset results
24 results for “semantic analysis”
Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis
<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>
dataset for "basic setting", "+ binary semantic loss", "+ class weights", "+ height weights", "+ region weights", "+ elastic distortion and subsampling", "+ TreeMix" in paper Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning
<p>dataset for "basic setting", "+ binary semantic loss", "+ class weights", "+ height weights", "+ region weights", "+ elastic distortion and subsampling", "+ TreeMix" in paper Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning</p>
Dataset of KO journal paper: Semantic analysis of archival concepts in CIDOC-CRM and in RiC-CM and RiC-O
<p>Semantic analysis of archival concepts (class, relations, atributtes, and relation attributes) presents in Records in Context family (conceptual model and ontology) and its possible equivalents in CIDOC-CRM.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights
<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook </a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - images
<p>In this notebook we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains high quality versions of the images used in the analysis.</p>
Time-generalized multivariate analysis of EEG responses reveals a cascading architecture of semantic mismatch processing
<p>This entry includes the required data to conduct the analysis from our paper. It includes recordings, a look-up csv for target word onsets and montage file. Github page for analysis code: https://github.com/heikele/GAT_n4-p6</p> <p> </p> <p>Abstract from submitted paper:</p> <p>Event-related brain potentials have a strong impact on neurocognitive models, as they inform about the temporal sequence of cognitive processes. Nevertheless, their value for deciding among alternative cognitive architectures is partly limited by component overlap and the possibility of ambiguity regarding component identity. Here, we apply temporally-generalized multivariate pattern analysis – a recently-proposed machine learning method capable of tracking the evolution of neurocognitive processes over time – to constrain possible alternative architectures underlying the processing of semantic incongruency in sentences. In a spoken sentence paradigm, we replicate established N400/P600 correlates of semantic mismatch. Time-generalized decoding indicatesthat early vs. late mismatch-sensitive processes are (i) distinct in their neural substrate, arguing against recurrent or latency-shifted single process architectures, and (ii) partially overlapping in time, inconsistent withpredictions of strictly serial models. These results are in accordance withan incremental-cascading neurocognitive organization of semantic mismatch processing. We propose time-generalized multivariate decoding as a valuable tool for neurocognitive language studies.</p> <p> </p> <p>Keywords: EEG; ERP; semantic mismatch; N400; P600; multivariate pattern analysis; generalization across time decoding</p>
Exploring Korean adolescent stress on social media: A semantic network analysis
<p><strong>Korean Adolescent's Stress Semantic Network Analysis Project</strong></p> <p>Semantic Network Analysis for Korean Adolescent's Stress</p> <p>Input data file</p> <ul> <li>data_news.csv : News data collected from Naver(<a href="https://www.naver.com">https://www.naver.com</a>)</li> <li>data_blog.csv : Blog data collected from Naver(<a href="https://www.naver.com">https://www.naver.com</a>) and Daum(<a href="https://www.daum.net">https://www.daum.net</a>)</li> </ul> <p>Output files</p> <ul> <li>Frequency Table of Each word in Documents (<em><strong>freq_news.csv</strong></em>, <em><strong>freq_blog.csv</strong></em>)</li> <li>TF-IDF(Term Frequency-Inverse Document Frequency) Table of Each word in Documents (<em><strong>tfidf_news.csv</strong></em>, <em><strong>tf_idf_blog.csv</strong></em>)</li> <li>Frequency Table of 30 keywords in Documents (<em><strong>freq_news_30.csv</strong></em>, <em><strong>freq_blog_30.csv</strong></em>)</li> <li>DTM(Document Term Matrix) of 30 keywords in Documents (<em><strong>DTM_news_30.csv</strong></em>, <em><strong>DTM_blog_30.csv</strong></em>)</li> <li>COM(Co-Occurrence Matrix) of 30 keywords in Documents (<em><strong>COM_news_30.csv</strong></em>, <em><strong>COM_blog_30.csv</strong></em>)</li> <li>Binary COM of 30 keywords in Documents (<em><strong>BinaryCOM_news_30.csv</strong></em>, <em><strong>BinaryCOM_blog_30.csv</strong></em>)</li> <li>Centrality Table of 30 keywords in Documents (<em><strong>centrality_news_30.csv</strong></em>, <em><strong>centrality_blog_30.csv</strong></em>)</li> </ul>
Wikibio: a Semantic Resource for the Intersectional Analysis of Biographical Events
<p>If you use this resource please cite</p> <p> </p> <pre>@inproceedings{stranisci2023wikibio, title={WikiBio: a Semantic Resource for the Intersectional Analysis of Biographical Events}, author={Stranisci, Marco Antonio and Damiano, Rossana and Mensa, Enrico and Patti, Viviana and Radicioni, Daniele and Caselli, Tommaso and others}, booktitle={Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, volume={1}, pages={12370--12384}, year={2023}, organization={Association for Computational Linguistics} }</pre>
Word-Retrieval Treatment for Aphasia: Semantic Feature Analysis
ClinicalTrials.gov study NCT00125242. IPD Sharing: Not stated. Countries: 1. Publications: 1.
A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains
<p><strong>Content</strong></p> <p>This repository contains pre-trained computer vision models, data labels, and images used in the pre-print publication "A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains":</p> <ol> <li><em>ADPdevkit</em>: a folder containing the 50 validation ("tuning") set and 50 evaluation ("segtest") set of images from the Atlas of Digital Pathology database formatted in the VOC2012 style--the full database of 17,668 images is available for download from the original website</li> <li><em>VOCdevkit</em>: a folder containing the relevant files for the PASCAL VOC2012 Segmentation dataset, with both the trainaug and test sets</li> <li><em>DGdevkit</em>: a folder containing the 803 test images of the DeepGlobe Land Cover challenge dataset formatted in the VOC2012 style</li> <li><em>cues</em>: a folder containing the pre-generated weak cues for ADP, VOC2012, and DeepGlobe datasets, as required for the SEC and DSRG methods</li> <li><em>models_cnn</em>: a folder containing the pre-trained CNN models</li> <li><em>models_wsss</em>: a folder containing the pre-trained SEC, DSRG, and IRNet models, along with dense CRF settings</li> </ol> <p><strong>More information</strong></p> <p>For more information, please refer to the following article. <strong>Please cite this article when using the data set.</strong></p> <p>@misc{chan2019comprehensive,<br> title={A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains},<br> author={Lyndon Chan and Mahdi S. Hosseini and Konstantinos N. Plataniotis},<br> year={2019},<br> eprint={1912.11186},<br> archivePrefix={arXiv},<br> primaryClass={cs.CV}<br> }</p> <p>For the full code released on GitHub, please visit the repository at: <a href="https://github.com/lyndonchan/wsss-analysis">https://github.com/lyndonchan/wsss-analysis</a></p> <p><strong>Contact</strong></p> <p>For questions, please contact:<br> Lyndon Chan<br> lyndon.chan@mail.utoronto.ca<br> http://orcid.org/0000-0002-1185-7961</p>
Do Synthesis Centers Synthesize? A Semantic Analysis of Topical Diversity in Research
<p>Synthesis centers are a form of scientific organization that catalyzes and supports research that integrates diverse theories, methods and data across spatial or temporal scales to increase the generality, parsimony, applicability, or empirical soundness of scientific explanations. Synthesis working groups are a distinctive form of scientific collaboration that produce consequential, high-impact publications. But no one has asked if synthesis working groups synthesize: are their publications substantially more diverse than others, and if so, in what ways and with what effect? We investigate these questions by using Latent Dirichlet Analysis to compare the topical diversity of papers published by synthesis center collaborations with that of papers in a reference corpus. Topical diversity was operationalized and measured in several ways, both to reflect aggregate diversity and to emphasize particular aspects of diversity (such as variety, evenness, and balance). Synthesis center publications have greater topical variety and evenness, but less disparity, than do papers in the reference corpus. The influence of synthesis center origins on aspects of diversity is only partly mediated by the size and heterogeneity of collaborations: when taking into account the numbers of authors, distinct institutions, and references, synthesis center origins retain a significant direct effect on diversity measures. Controlling for the size and heterogeneity of collaborative groups, synthesis center origins and diversity measures significantly influence the visibility of publications, as indicated by citation measures. We conclude by suggesting social processes within collaborations that might account for the observed effects, by inviting further exploration of what this novel textual analysis approach might reveal about interdisciplinary research, and by offering some practical implications of our results.</p>
Decoding Knowledge Claims: the Evaluation of Scientific Publication Contributions through Semantic Analysis
<p>This data were used to compute the RWMD distance as described in the study submitted for the STI 2024 conference, Berlin.</p>
Figures - Semantic analysis of web archive historical data 1983 "Marche pour l'égalité et contre le racisme"
Open the record for dataset details and reuse information.
QSage: Structural and Semantic Metric Analysis for Quantum Code Smell Detection
Open the record for dataset details and reuse information.
Artifact For A Large Scale Analysis of Semantic Versioning in NPM
<p>This is the artifact for: A Large Scale Analysis of Semantic Versioning in NPM.</p> <p>The artifact contains:</p> <ul> <li>A full scrape of all metadata from NPM (package / version information, dependencies, etc.) as of October 31, 2022.</li> <li>A copy of our code, which includes the software for scraping metadata and package tarball (code) data, as well as all analysis scripts that are needed to replicate the figures from the paper.</li> </ul>
Comparing Traditional Semantic Feature Analysis (tSFA) and Semantic Feature Analysis + Metacognitive Strategy Training (SFA+MST)
ClinicalTrials.gov study NCT07036406. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.
Semantic Feature Analysis Treatment for Aphasia
ClinicalTrials.gov study NCT04215952. IPD Sharing: NO. Countries: 1. Publications: 4.
Do Synthesis Centers Synthesize? A Semantic Analysis of Topical Diversity in Research
Open the record for dataset details and reuse information.
SEMANTIC AND STRUCTURAL ANALYSIS OF THE TOURISTIC TERMS IN THE ENGLISH AND UZBEK LANGUAGES.
Open the record for dataset details and reuse information.
COMPARATIVE ANALYSIS OF SEMANTIC MEANING OF RESPECT TERMS
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.