Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

483 results for “SEMANTICS”

Learn how ShareScore rates datasets ↗
zenodo36/100

Example Datasets for Semantic SimilarityAnalysis

<p>This dataset contains a set of example data for a semantic similarity analysis tutorial.</p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

Semantic Triples from "A Collaborative, Realism-Based, Electronic Healthcare Graph: Public Data, Common Data Models, and Practical Instantiation"

<p>These RDF triples (<a href="https://zenodo.org/api/files/3d3308cb-8221-4a17-abc5-0ae32bb33f26/synthea_graph_exportable.nq.zip?versionId=7aff7c4a-bb0c-46ae-b006-7a80bcec0925">synthea_graph_exportable.nq.zip</a>) are the result of modeling electronic health records (<a href="https://zenodo.org/api/files/3d3308cb-8221-4a17-abc5-0ae32bb33f26/synthea_csv_output_turbo_cannonical.zip">synthea_csv_output_turbo_cannonical.zip), </a>that were synthesized with the Synthea software (https://github.com/synthetichealth/synthea). Anyone who loads them into a triplestore database is encouraged to provide feedback at https://github.com/PennTURBO/EhrGraphCollab/issues. The following abstract comes from a paper, describing the semantic instantiation process, and presented to the ICBO 2019 conference (https://drive.google.com/file/d/1eYXTBl75Wx3XPMmCIOZba-8Cv0DIhlRq/view).</p> <p>ABSTRACT:&nbsp; There is ample literature on the semantic modeling of biomedical data in general, but less has been published on realism-based, semantic instantiation of electronic health records (EHR). Reasons include difficult design decisions and issues of data governance. A collaborative approach can address design and technology utilization issues, but is especially constrained by limited access to the data at hand: protected health information.</p> <p>Effective collaboration can be facilitated by public EHR-like data sets, which would ideally include a large variety of datatypes mirroring actual EHRs and enough records to drive a performance assessment. An investment into reading public EHR-like data from a popular common data model (CDM) is preferable over reading each public data set&rsquo;s native format.</p> <p>In addition to identifying suitable public EHR-like data sets and CDMs, this paper addresses instantiation via relational-to-RDF mapping. The completed instantiation is available for download, and a competency question demonstrates fidelity across all discussed formats.</p>

opencc-zeroApr 2019View details →
zenodo36/100

Mining the UK Web Archive for Semantic Change Detection (Dataset)

<p>The dataset that was used and released with the RANLP 2019 paper, titled &quot;Mining the UK Web Archive for Semantic Change Detection&quot; (see&nbsp;<a href="https://github.com/adtsakal/Semantic_Change">https://github.com/adtsakal/Semantic_Change</a>).&nbsp;It contains annual word2vec representations of more than 47K words over the period 2000-2013, along with a list of 65 words with known semantic change over the same time period.&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science – data and code

<p>Data (including object annotations) and code from the following manuscript:</p> <p><em>Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science</em>.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Demo-Dataset for publication "FAIR workflows in Earth system modelling: a use case with semantic data management"

<p>This demodataset is intended to be used to test the workflow described in the publication by Lennartz &amp; Schlemmer&nbsp; "FAIR workflows in Earth System modelling: a use case with semantic data management". It contains example model output for an arbitrary biogeochemical model tracer (here: dissolved organic carbon, DOC) from an ocean model as a 4-dimensional dataset (latitude, longitude, depth, time), the corresponding grid point locations as well as a textfile specifying parameter inputs for the model. The file structure is adapted for seamless integration into the workflow described in Lennartz &amp; Schlemmer, which builds on the open source semantic research data management system LinkAhead. The dataset contains the following structure: The folder DataAnalysis stores data required for data analysis, such as the grid point locations in the file TMM_grid_v2018a.mat. The folder SimulationData stores model output in the folder 2022_TMM, containing the parameter input file nl_in.txt and the model output TR_monthly.mat. Related instructions can be accessed here: https://gitlab.com/salexan/fairworkflows-demodataset .</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Semantic2D: A Semantic Dataset for 2D Lidar Semantic Segmentation: Dataset

<p>The 2D lidar semantic segmentation datasets for the paper titled "Semantic2D: A Semantic Dataset for 2D Lidar Semantic Segmentation" by Zhanteng Xie and Philip Dames.</p> <p>The relevant code is available at: https://github.com/TempleRAIL/semantic2d</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Test Dataset for 3D semantic image segmentation of the Breast, Fibrograndular Tissue, and Breast Carcinoma

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo36/100

SynRS3D : A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery

<h1><strong>SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery</strong></h1> <h3><strong>Neural Information Processing Systems (Spotlight), 2024</strong></h3> <p>For more details, please refer to our&nbsp;<a href="https://arxiv.org/pdf/2406.18151">paper</a> and visit our <a href="https://github.com/JTRNEO/SynRS3D">GitHub repository</a>.</p> <h2><strong>Overview</strong></h2> <p><strong>TL;DR:</strong><br>SynRS3D is a comprehensive synthetic remote sensing dataset designed to improve global 3D semantic understanding from monocular high-resolution imagery. It includes data for three key tasks:</p> <ul> <li>Height estimation</li> <li>Land cover mapping</li> <li>Building change detection</li> </ul> <h2><strong>Dataset Structure</strong></h2> <p>The dataset consists of 17 folders and includes a total of 69,667 images at a resolution of 512x512. After downloading and extracting the files, ensure the directory structure follows this format:</p> <p>${DATASET_ROOT} &nbsp;# Example: /home/username/project/SynRS3D/data/grid_g05_mid_v1<br>├── opt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # RGB images (.tif), also used as post-event images for building change detection<br>├── pre_opt &nbsp; &nbsp; &nbsp; # RGB images (.tif), used as pre-event images for building change detection<br>├── gt_nDSM &nbsp; &nbsp; &nbsp; # Normalized Digital Surface Model (nDSM) images (.tif)<br>├── gt_ss_mask &nbsp; &nbsp;# Land cover mapping labels (.tif)<br>├── gt_cd_mask &nbsp; &nbsp;# Building change detection masks (.tif, 0 = no change, 255 = change area)<br>└── train.txt &nbsp; &nbsp; # List of training data filenames</p> <p>The land cover mapping labels (`gt_ss_mask`) are mapped to the following categories:</p> <ul> <li>Bareland: 1</li> <li>Rangeland: 2</li> <li>Developed Space: 3</li> <li>Road: 4</li> <li>Trees: 5</li> <li>Water: 6</li> <li>Agriculture land:&nbsp; 7</li> <li>Buildings: 8</li> </ul> <h2><strong>Image Breakdown by Folder</strong></h2> <p>The dataset is organized into grid-like and irregular terrain. It includes a range of ground sampling distances (GSDs) and variations in building heights. The folder naming convention indicates these characteristics: &nbsp;<br>- `grid` = grid-like terrain &nbsp;<br>- `terrain` = irregular terrain &nbsp;<br>- `g005`, `g05`, `g1` = GSD ranges (0.05m&ndash;0.3m, 0.3m&ndash;0.6m, and 0.6m&ndash;1m, respectively) &nbsp;<br>- `low`, `mid`, `high` = building height variations</p> <p>The dataset includes the following image counts:</p> <p>- 1,430 images &ndash; `terrain_g05_mid_v1`<br>- 10,000 images &ndash; `grid_g05_mid_v2`<br>- 2,354 images &ndash; `terrain_g05_low_v1`<br>- 3,707 images &ndash; `terrain_g05_high_v1`<br>- 880 images &ndash; `terrain_g005_mid_v1`<br>- 2,127 images &ndash; `terrain_g005_low_v1`<br>- 11,325 images &ndash; `grid_g005_mid_v2`<br>- 1,212 images &ndash; `terrain_g005_high_v1`<br>- 348 images &ndash; `terrain_g1_mid_v1`<br>- 4,285 images &ndash; `terrain_g1_low_v1`<br>- 904 images &ndash; `terrain_g1_high_v1`<br>- 3,000 images &ndash; `grid_g005_mid_v1`<br>- 2,997 images &ndash; `grid_g005_low_v1`<br>- 4,000 images &ndash; `grid_g005_high_v1`<br>- 7,000 images &ndash; `grid_g05_mid_v1`<br>- 7,098 images &ndash; `grid_g05_low_v1`<br>- 7,000 images &ndash; `grid_g05_high_v1`</p> <h2><strong>Citation</strong></h2> <p>If you find SynRS3D useful in your research, please consider citing:</p> <div> <div>@article{song2024synrs3d,</div> <div>title={SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery},</div> <div>author={Song, Jian and Chen, Hongruixuan and Xuan, Weihao and Xia, Junshi and Yokoya, Naoto},</div> <div>journal={arXiv preprint arXiv:2406.18151},</div> <div>year={2024}</div> <div>}</div> </div> <h2><strong>Contact</strong></h2> <p>For any questions or feedback, feel free to reach out via email:&nbsp;<strong> song@ms.k.u-tokyo.ac.jp</strong>.</p> <p>Enjoy using SynRS3D!</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Microservice Semantic Dependency

<p>This dataset is associated with a publication on component-based semantic dependency detection methodology in microservices-based systems.</p> <p>It includes a prototype implementation and the generated semantic clones data derived from the TrainTicket v0.1.0 benchmark.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Bugzz lightyears: To Semantic Segmentation and Bug-yond!

<p>Dataset Title:&nbsp; <em><strong>Bugzz lightyears: To Semantic Segmentation and Bug-yond!</strong></em></p> <h3>Description:</h3> <p>This dataset comprises a collection of real and robotic toy bugs designed for a small-scale semantic segmentation project. Each bug has been captured six times from various angles, ensuring comprehensive coverage of their features and details. The dataset serves as a valuable resource for exploring semantic segmentation techniques and evaluating machine learning models.</p> <h3>Dataset Details:</h3> <ul> <li>Images: Each bug is represented by six images taken from different perspectives, facilitating robust segmentation and analysis.</li> <li>Segmentation: The dataset has been meticulously segmented using Label Studio in conjunction with the SAM (Segment Anything Model), enabling precise delineation of each bug from the background.</li> <li>Diversity: The collection includes a variety of bugs, both real and robotic, providing a unique blend for training and testing segmentation models.</li> </ul> <h3>Usage: This toy dataset is ideal for researchers and developers interested in:</h3> <ul> <li>Experimenting with semantic segmentation algorithms.</li> <li>Developing and refining computer vision models for object detection and segmentation.</li> <li>Educational purposes in machine learning and computer vision courses.</li> </ul> <h3>License: This dataset is made available under [specify license type, e.g., CC BY 4.0], allowing for both academic and commercial use, with proper attribution to the creator.</h3>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Sampled sentence pairs from SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p>Each dataset consists of samples, containing two sentences with positions of one of the given target words. In every sample, first sentence is taken from corpus1 and second from corpus2. Initial sentences were taken from&nbsp;https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd/.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Data companion to Coussé & Bouma (2021) Semantic scope restrictions in complex verb constructions in Dutch

<p>The material in this archive accompanies the paper</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Evie Couss&eacute; and Gerlof Bouma,<br> &nbsp; &nbsp; Semantic scope restrictions in complex verb constructions in Dutch,</p> <p>accepted for publication in Linguistics, An Interdisciplinary Journal<br> of the Language Sciences, De Gruyter Mouton.</p> <p>The archive contains:</p> <ul> <li>the annotated data of the paper</li> <li>the code for extraction of the data from the Spoken Dutch Corpus and the Lassy Small Corpus</li> <li>documentation of the selection and annotation process that were too detailed&nbsp;to be included in the paper</li> </ul>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Digital humanities semantic Notebook / video of the presentation

<p>After a theoretical introduction on the concepts of reproducible science and the Web of Data, we will present the methodology of scientific notebooks (e.g. Jupyter Notebooks). In this context, we will present a step-by-step approach to the creation of a scientific notebook in history: we will load a dataset from the Open Data of the institute and then apply filters and calculations on the data. A presentation and analysis layer with graphics will complete our research product with enrichments from external data and vocabularies from the Semantic Web. This will make our output reproducible and documented with text enriched with semantic schemas (schema.org) and disambiguation authorities (GND, dpPedia, Wikidata...), data sources and computer code. The presentation does not require advanced technical knowledge. It is aimed at beginners. The technical demonstration, in the second part, is deliberately simple in content and will be commented on as it goes along.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

The NGI Forward semantic social network data

<p>The <a href="https://research.ngi.eu/">NGI Forward project</a> is part of the European Union&#39;s <a href="https://www.ngi.eu">Next Generation Internet Initiative</a>. It is meant to provide European institutions with policy advice for how to shape the future, human-centric Internet. As part of it, a team of ethnographers coded a specially convened online conversation, then arranged its results into a semantic social network. This dataset encodes that conversation, as well as the results of the coding exercise, in raw data form for further exploration and replication purposes. The dataset is pseudonymized.</p> <ul> <li><a href="https://exchange.ngi.eu/">Funnel website</a> of the project.</li> <li><a href="https://journals.sagepub.com/doi/10.1177/1525822X20908236">About semantic social networks</a>.</li> <li><a href="https://edgeryders.eu/t/long-term-ssna-data-storage-documentation-manual/12786">Data export and documentation process</a> (contains links to the code used to export the data)</li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo36/100

Semantically tagged Finnish parliament discussions 1991-2015

<p><strong>Semantically tagged Finnish parliament discussions 1991-2015</strong></p> <p>The original data is this (Rauh et al, 2017):</p> <p>https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/E4RSP9</p> <p>We have imported the raw text out of the original data set without speaker and party tags. The Finnish text has been first tagged with UD2 parser using the Mylly service of the Language Bank of Finland. After UD2 parse, semantic tags have been added to the text with FiST (Kettunen, 2019).</p> <p><strong>Output form</strong></p> <p># newdoc</p> <p># newpar</p> <p># sent_id = 1</p> <p># text = Arvoisa herra puhemies!</p> <p>Arvoisa arvoisa Z99 amod</p> <p>herra#herra#Noun#S2.2m S9 compound:nn</p> <p>puhemies#puhemies#Noun#G1.1/S2 root</p> <p>! PUNCT</p> <p>&nbsp;</p> <p>The output consists of the 1) token word form of the text, 2) lemma of the token, 3) POS, 4) semantic tag and 5) UD2 grammatical relation of the word in the sentence.</p> <p>Our semantic tagger does not resolve ambiguity. If the lexeme has multiple possible semantic tags, they are all included in the output (e.g., herra#herra#Noun#<strong>S2.2m S9</strong> compound:nn). Slash notation in the semantic tags (e.g., puhemies#puhemies#Noun#<strong>G1.1/S2</strong>) indicates that the word can belong to two or more categories (L&ouml;fbeg, 2017).&nbsp;If the semantic category of the word is not recognized, the word is tagged with Z99.</p> <p>The data consists of 4 036&nbsp;269 sentences and ca. 65.247 million words. Lexical coverage of FiST for the data is 86.05 %, i.e. 86% of the words are known for the tagger and marked with a semantic tag.</p> <p><strong>References</strong></p> <p>Kettunen, Kimmo (2019). FiST &ndash; towards a Free Semantic Tagger of Modern Standard Finnish. IWCLUL2019, http://aclweb.org/anthology/W19-0306</p> <p>Rauh, Christian; De Wilde, Pieter; Schwalbach, Jan, 2017, &quot;Corp_Eduskundta.Rdata&quot;, The ParlSpeech data set: Annotated full-text vectors of 3.9 million plenary speeches in the key legislative chambers of seven European states, https://doi.org/10.7910/DVN/E4RSP9/U8VZHK, Harvard Dataverse, V1.</p> <p>Lofberg, L. (2017). Creating large semantic lexical resources for the Finnish language. [Doctoral Thesis, Lancaster University]. Lancaster University. https://doi.org/10.17635/lancaster/thesis/3</p> <p>UCREL Semantic Analysis System (USAS). https://ucrel.lancs.ac.uk/usas/</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Supplementary material for a usability evaluation of a semantic search for biological datasets

<p>We conducted a usability evaluation for a semantic dataset search with 20 biodiversity scholars in June and July 2022 in Germany.</p> <p>Following the TREC guidelines (https://www-nlpir.nist.gov/projects/t9i/spec.html), we setup eight user tasks and surveys with questionnaires to guide users through the evaluation. The zip file provides questionnaires, survey templates and the original results.</p> <p>A Jupyter notebook for data analysis is provided in our GitHub repository: <a href="https://github.com/fusion-jena/semantic-search-usability-analysis">https://github.com/fusion-jena/semantic-search-usability-analysis</a></p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Annotated dataset for the semantic segmentation of radishes

<p>This folder contains pictures of radishes collected on the PMF experimental field during Spring 2017. There are two kinds of labeled images in the following folders:</p> <p><br> -&nbsp; human annotations: human annotators draw polygons around each plant and those were then refined using an active contours algorithm.<br> -&nbsp; machine annotations: An SVM trained on the human annotations was used to produce labeled images. Images with bad segmentation were manually discarded.</p> <p>Each of these folder contains an images folder containing original pictures and a labels folder containing binary segmentation masks.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Exploring Korean adolescent stress on social media: A semantic network analysis

<p><strong>Korean Adolescent&#39;s Stress Semantic Network Analysis Project</strong></p> <p>Semantic Network Analysis for Korean Adolescent&#39;s Stress</p> <p>Input data file</p> <ul> <li>data_news.csv : News data collected from Naver(<a href="https://www.naver.com">https://www.naver.com</a>)</li> <li>data_blog.csv : Blog data collected from Naver(<a href="https://www.naver.com">https://www.naver.com</a>) and Daum(<a href="https://www.daum.net">https://www.daum.net</a>)</li> </ul> <p>Output files</p> <ul> <li>Frequency Table of Each word in Documents (<em><strong>freq_news.csv</strong></em>, <em><strong>freq_blog.csv</strong></em>)</li> <li>TF-IDF(Term Frequency-Inverse Document Frequency) Table of Each word in Documents (<em><strong>tfidf_news.csv</strong></em>, <em><strong>tf_idf_blog.csv</strong></em>)</li> <li>Frequency Table of 30 keywords in Documents (<em><strong>freq_news_30.csv</strong></em>, <em><strong>freq_blog_30.csv</strong></em>)</li> <li>DTM(Document Term Matrix) of 30 keywords in Documents (<em><strong>DTM_news_30.csv</strong></em>, <em><strong>DTM_blog_30.csv</strong></em>)</li> <li>COM(Co-Occurrence Matrix) of 30 keywords in Documents (<em><strong>COM_news_30.csv</strong></em>, <em><strong>COM_blog_30.csv</strong></em>)</li> <li>Binary COM of 30 keywords in Documents (<em><strong>BinaryCOM_news_30.csv</strong></em>, <em><strong>BinaryCOM_blog_30.csv</strong></em>)</li> <li>Centrality Table of 30 keywords in Documents (<em><strong>centrality_news_30.csv</strong></em>, <em><strong>centrality_blog_30.csv</strong></em>)</li> </ul>

opencc-by-4.0Feb 2023View details →
zenodo36/100

What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study

<p><strong>What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study &nbsp;</strong></p> <p>This repository contains data and code for the paper <a href="https://arxiv.org/abs/2110.04845">What Makes Sentences Semantically Related: A Textual Relatedness Dataset and Empirical Study</a>.</p> <p>We hope that this work will spur further research on understanding sentence--sentence relatedness, methods of sentence representation, measures of semantic relatedness, and their applications.</p> <p><strong>Citing our work</strong><br> Please use the following BibTex entry to cite us if you use our dataset or any of the associated analyses:</p> <blockquote>@inproceedings{abdalla2023makes,<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; title={What Makes Sentences Semantically Related: A Textual Relatedness Dataset and Empirical Study},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; author={Abdalla, Mohamed and Vishnubhotla, Krishnapriya and Mohammad, Saif M.},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; year={2023},<br> &nbsp;&nbsp; &nbsp;&nbsp; address = {Dubrovnik, Croatia},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; publisher = &quot;Association for Computational Linguistics&quot;,<br> &nbsp;&nbsp; &nbsp;&nbsp; booktitle = &quot;Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume&quot;<br> }</blockquote> <p><strong>Dataset Description</strong></p> <p>The dataset consists of 5500 English sentence pairs that are scored and ranked on a relatedness scale ranging from 0 (least related) to 1 (most related).</p> <p><strong>Why Semantic Relatedness?</strong><br> Closeness of meaning can be of two kinds: semantic relatedness and semantic similarity. Two sentences are considered semantically similar when they have a paraphrasal or entailment relation, whereas relatedness accounts for all of the commonalities that can exist between two sentences. Semantic relatedness is central to textual coherence and narrative structure. Automatically determining semantic relatedness has many applications such as question answering, plagiarism detection, text generation (say in personal assistants and chat bots), and summarization.</p> <p>Prior NLP work has focused on semantic similarity (a small subset of semantic relatedness), largely because of a dearth of datasets. In this paper, we present the first manually annotated dataset of sentence--sentence semantic relatedness. It includes fine-grained scores of relatedness from 0 (least related) to 1 (most related) for 5,500 English sentence pairs. The sentences are taken from diverse sources and thus also have diverse sentence structures, varying amounts of lexical overlap, and varying formality.</p> <p><strong>Comparative Annotations and Best-Worst Scaling</strong><br> Most existing sentence-sentence similarity datasets were annotated, one item at a time, using coarse rating labels such as integer values between 1 and 5 @ representing coarse degrees of closeness. It is well documented that such approaches suffer from inter- and intra-annotator inconsistency, scale region bias, and issues arising due to the fixed granularity.</p> <p>The relatedness scores for our dataset were, instead, obtained using a <em>comparative</em> annotation schema. In comparative annotations, two (or more) items are presented together and the annotator has to determine which is greater with respect to the metric of interest.</p> <p>Specifically, we use Best-Worst Scaling, a comparative annotation method}, which has been shown&nbsp; to produce reliable scores with fewer annotations in other NLP tasks. We use scripts from https://saifmohammad.com/WebPages/BestWorst.html to obtain relatedness scores from our annotations.</p> <p><br> <strong>Loading the Dataset</strong><br> - The sentence pairs, and associated scores, are in the file sem_text_rel_ranked.csv in the root directory. The CSV file can be read using:</p> <pre><code class="language-python">python import pandas as pd str = pd.read_csv('sem_text_rel_ranked.csv') row = str.loc[0] sent1, sent2 = row['Text'].split("\n") score = row['Score']</code></pre> <p>- Relevant columns:</p> <p>&nbsp; - Text: Sentence pair, separated by the newline character.<br> &nbsp; - Score: The semantic relatedness score between 0 and 1.</p> <p>- Additionally:<br> &nbsp; - the SourceID column indicates the source dataset from which the sentence pair was drawn (see Table 2 of our paper)<br> &nbsp; - The SubsetID column indicates the sampling strategy used for the source dataset<br> &nbsp; - and the PairID is a unique identifier for each pair that also indicates its Source and Subset.</p> <p><br> <strong>Raw Annotations from Amazon Mechanical Turk</strong></p> <p>- The `mturk_data/` subdirectory provides the raw MTurk annotations obtained with our comparative annotation setup.<br> - Each row of `mturk_data/bws_annotations.csv` consists of four sentence pairs along with human annotations for the most related (column `BestItem`) and the least related (column `WorstItem`) pair.<br> - File `mturk_data/id2sents.csv` pairs each sentence pair with the corresponding SourceID, SubsetID, and PairID that indicates the source dataset (see Table 2 of our paper).<br> - See file `mturk_data/task_intructions.txt` for the instructions provided to annotators for our task.</p> <p><br> <strong>Datasheet for STR-2022</strong><br> The datasheet for our dataset is in the document `STR2022-datastatement.pdf` in the root folder of this repository.</p> <p><strong>Ethics Statement</strong><br> Any dataset of semantic relatedness entails several ethical considerations. We talk about this in Section 8 of our paper.</p> <p><strong>Creators</strong><br> - <a href="https://www.cs.toronto.edu/~msa/index.html">Mohamed Abdalla</a> (University of Toronto)<br> - <a href="https://priya22.github.io/">Krishnapriya Vishnubhotla</a> (University of Toronto)<br> - <a href="http://saifmohammad.com/">Saif M. Mohammad</a> (National Research Council Canada)</p> <p><strong>Contact:</strong> msa@cs.toronto.edu, vkpriya@cs.toronto.edu, saif.mohammad@nrc-cnrc.gc.ca</p> <p>&nbsp;</p>

openother-closedNov 2021View details →
zenodo36/100

Raw data for manuscript Semantic context can mask intelligibility declines at above-conversational speech levels in normal-hearing listeners

<p>Raw data for the manuscript in doc file.&nbsp;<br> Copied from the Matlab .m file. used for the analysis.</p> <p>To be updated.</p> <p>For details, contact me at mfer@health.sdu.dk</p>

opencc-by-4.0Feb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record