Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “semantic matching”
SemTab 2024: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching Data Sets - WikidataTables2024R1 and WikidataTables2024R2
<p>Data Sets from the ISWC 2024 Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, Round 1, Wikidata Tables. Links to other datasets can be found on the challenge website: https://sem-tab-challenge.github.io/2024/ as well as the proceedings of the challenge published on CEUR.</p> <p>For details about the challenge, see: http://www.cs.ox.ac.uk/isg/challenges/sem-tab/</p> <p>For 2024 edition, see: https://sem-tab-challenge.github.io/2024/</p> <p>Note on License: This data includes data from the following sources. Refer to each source for license details:<br>- Wikidata https://www.wikidata.org/</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries
<p><strong>Abstract</strong> (our paper)</p> <p>WordNet is one of the largest handcrafted concept dictionaries visualizing word connections through semantic relationships. It is widely used as a word sense inventory in natural language processing tasks. However, WordNet's fine-grained senses have been criticized for limiting its usability. In this paper, we semantically match sense definitions from Cambridge dictionaries and WordNet and develop new coarse-grained sense inventories. We verify the effectiveness of our inventories by comparing their semantic coherences with that of Coarse Sense Inventory. The advantages of the proposed inventories include their low dependency on large-scale resources, better aggregation of closely related senses, CEFR-level assignments, and ease of expansion and improvement. Our inventories are publicly available for free use.</p> <p><strong>Publication</strong></p> <p>These datasets are part of our research results. If you make use of our datasets, please cite:</p> <ul> <li>Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono. Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries. In <em>Proceedings of the 11th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA 2024)</em>. 6 pages, 2024.</li> </ul>
Semantic address matching dataset
<p>Data for our paper <strong>Lin, Y., Kang, M., Wu, Y., Du, Q. and Liu, T. (2019) A deep learning architecture for semantic address matching, <em>International Journal of Geographical Information Science</em>, DOI: 10.1080/13658816.2019.1681431</strong></p> <p>Below is an overview of each file in this dataset.</p> <ul> <li><code>train.txt</code> The training dataset</li> <li><code>train_code_a.txt</code> The index representations of the address elements (i.e., address elements represented by the corresponding indexes in the vocabulary obtained by word2vec) in <em>S<sub>a</sub></em></li> <li><code>train_code_b.txt</code> The index representations of the address elements in <em>S<sub>b</sub></em></li> <li><code>train_lable.txt</code> The labels of address pairs in the training dataset</li> </ul> <p> </p> <ul> <li><code>dev.txt</code> The development dataset</li> <li><code>dev_code_a.txt</code> The index representations of the address elements in <em>S<sub>a</sub></em></li> <li><code>dev_code_b.txt</code> The index representations of the address elements in <em>S<sub>b</sub></em></li> <li><code>dev_lable.txt</code> The labels of address pairs in the development dataset</li> </ul> <p> </p> <ul> <li><code>test.txt</code> The test dataset</li> <li><code>test_code_a.txt</code> The index representations of the address elements in <em>S<sub>a</sub></em></li> <li><code>test_code_b.txt</code> The index representations of the address elements in <em>S<sub>b</sub></em></li> <li><code>test_lable.txt</code> The labels of address pairs in the test dataset</li> </ul>
Semantic Parameter Matching in Web APIs with Transformer-based Question Answering
<p>This repository contains the evaluation results of our study, as well as datasets and model checkpoints. <br> For a detailed overview regarding the provided materials, please refer to README.md.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.