Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

18 results for “embedding evaluation”

Learn how ShareScore rates datasets ↗
zenodo44/100

Jacdac: Service-based Prototyping of Embedded Systems (PLDI 2024 Artifact Evaluation)

<p>This artifact allows others to reproduce and explore the results seen in "Jacdac: Service-based Prototyping of Embedded Systems". The artifact contains a prebuilt docker image and the Dockerfile source used to produce the prebuilt docker image. Evaluators should follow the README contained in this artifact for complete instruction.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Embedding Evaluation Data for South African Languages

<p><strong>WordSim and Simlex Data for South African Languages</strong></p> <ul> <li>Setswana</li> <li>Sepedi</li> </ul> <p><strong>Embedding Evaluation Data for South African Languages</strong></p> <p><strong>Dataset Information\</strong></p> <p>The datasets(Simlex and WordSim) contain pairs of Setswana and Sepedi words that have been assigned similarity ratings by humans to measure semantic relatedness. The word-pairs(Simlex and WordSim) are manually translated from English to Setswana and Sepedi. The evaluation task aims to find the degree of correlation between the scores provided by the model and the human rating, the score of the model is collected by computing the cosine similarity of corresponding vectors for word pairs.</p> <p>Online Repository link</p> <ul> <li><a href="https://zenodo.org/record/5673974">Zenodo Data Repository</a>&nbsp;- Link to the data repository.</li> </ul> <p>Authors</p> <ul> <li><strong>Vukosi Marivate</strong>&nbsp;-&nbsp;<a href="https://twitter.com/vukosi">@vukosi</a></li> <li><strong>Valencia Wagner</strong></li> <li><strong>Mack Makgatho</strong></li> <li><strong>Tshephisho Sefara</strong></li> </ul> <p>See also the list of&nbsp;<a href="https://github.com/dsfsi/embedding-eval-data//contributors">contributors</a>&nbsp;who participated in this project.</p> <p>Citing the dataset</p> <p>To appear in conference proceedings</p> <blockquote> <p>@article{Makgatho_Marivate_Sefara_Wagner_2022, title={Training Cross-Lingual embeddings for Setswana and Sepedi},&nbsp;<br> volume={3},&nbsp;<br> url={https://upjournals.up.ac.za/index.php/dhasa/article/view/3822},&nbsp;<br> DOI={10.55492/dhasa.v3i03.3822},&nbsp;<br> number={03},<br> journal={Journal of the Digital Humanities Association of Southern Africa },<br> author={Makgatho, Mack and Marivate, Vukosi and Sefara, Tshephisho and Wagner, Valencia},&nbsp;<br> year={2022},&nbsp;<br> month={Feb.}}</p> </blockquote>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Early Irish Analogy Dataset for Word Embedding Evaluation

<p>An embedding evaluation dataset for Early Irish described in the paper "<a href="https://aclanthology.org/2023.insights-1.10.pdf">Do not Trust the Experts: How the Lack of Standard Complicates <span>NLP</span> for Historical <span>I</span>rish</a>".</p> <p>Traditionally, analogy datasets are based on pairwise semantic proportion, and therefore every question has a single correct answer. Given the high level of variation in historical languages, such a strict definition of a correct answer seems unjustified. Therefore, Early Irish Analogy Dataset follows the <a href="https://vecto.space/projects/BATS/">Bigger Analogy Test Set (BATS)</a> and provides several correct answers to each analogy question.&nbsp;</p> <p>Morphological and spelling variation data are extracted from the <a href="https://dil.ie/">eDIL</a>, a historical dictionary of medieval Irish. Unlike BATS, no distinction is made between inflection types due to eDIL's structure. The raw data amounted to 2,370 spelling variation and 9,690 morphological variation questions, from which 150 examples were randomly selected for each of the subsets to be comparable in size with the synonym and antonym subsets. The synonym and antonym subsets are translations of the correspondent BATS parts obtained by reverse-searching the eDIL and proofread by four expert evaluators. The dataset includes 98 entries in the synonym subset and 109 entries in the antonym subset, upon which three or more experts agreed.</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Embeddings models for Buddhist Sanskrit: Evaluation Datasets

<p>Evaluation Dataset used for the study published as&nbsp;&nbsp;Embeddings models for Buddhist Sanskrit,&nbsp; <em>LREC 2022 proceedings</em>. It contains a semantic similarity dataset&nbsp;and an analogy dataset, as well as the published study and a ReadMe file containing the&nbsp;guidelines used for scoring semantic &nbsp;similarity and some notes about the manual scoring task.</p> <p>&nbsp;</p> <p>The evaluation datasets have been prepared by&nbsp;Ligeia Lugli,&nbsp; Bruno Galasek-Hul, Luis Qui&ntilde;ones and Jai Paranjape</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Tigrinya Analogy Test for evaluating Word Embeddings

<p><strong>Tigrinya Analogy Test for evaluating Word Embeddings</strong></p> <p>This is a Tigrinya version of the Google Analogy Test set, which is used to evaluate English word-embedding models.&nbsp;The analogy test is a well-established strategy to empirically evaluate the quality of word-embedding models. More information about the English task can be found at the <a href="https://aclweb.org/aclwiki/Google_analogy_test_set_(State_of_the_art)">ACL Wiki</a>.</p> <p>This data is&nbsp;&nbsp;was first machine&nbsp;translated then&nbsp;manually verified by a native speaker to reduce errors.</p> <p>Some aspects of the original analogy test is focused on English and may not transfer well to other languages, such as those related to grammar or morphology. Therefore, we have discarded examples that became irrelevant in Tigrinya when adapting the task. Finally, there are a total of <strong>18465</strong>&nbsp;entries in the Tigrinya Analogy Test set, while the source English data has <strong>19544</strong>&nbsp;entries.</p> <p>An entry is dropped if the translations led to one of the following conditions:</p> <ol> <li>If the source word pair map to one Tigrinya word, for example, lucky &amp; luckiest both correspond to ዕድለኛ.</li> <li>If the source word results in a multi-word expression. For example, grandson (ወዲ ጓል / ወዲ ወዲ), granddaughter (ጓል ጓል / ጓል ወዲ). This because the typical word-embedding approaches such as <em>word2vec</em> are not designed to predict multi-word phrases.</li> </ol> <p>&nbsp;</p> <p><strong>Test Sections</strong></p> <p>The test includes a series of semantic and syntactic analogies divided up into subsections including world capitals,&nbsp;currencies, family, tense, and plurality. The test contains the following sections:</p> <ol> <li>capital-world</li> <li>currency</li> <li>city-in-state</li> <li>family</li> <li>gram1-adjective-to-adverb</li> <li>gram2-opposite</li> <li>gram3-comparative</li> <li>gram4-superlative</li> <li>gram5-present-participle</li> <li>gram6-nationality-adjective</li> <li>gram7-past-tense</li> <li>gram8-plural</li> <li>gram9-plural-verbs</li> </ol> <p>&nbsp;</p> <p><strong>Examples:</strong></p> <ul> <li>Semantic section of World Capitals: &ldquo;ኣስመራ: ኤርትራ as ፓሪስ: ?&rdquo; and if the model responds correctly it will return: &ldquo;ፈረንሳ&rdquo;.</li> <li>Semantic section of Family section: &ldquo;ሰብኣይ: ሰበይቲ as ወዲ: ጓል&rdquo;.</li> <li>Syntax section with tense, a sample analogy might be &ldquo;Walk: Walked as Run: Ran&rdquo;.<br> &nbsp;</li> </ul> <p><strong>Evaluation</strong></p> <p>The final accuracy of a model is the proportion of the questions that the model answers correctly.<br> Generally, a better-quality model would answer more questions correctly than a model of lower quality.<br> However, note that a model with low performance on this analogy test, might still contain useful information, but may not be robust or good enough for more complex tasks.</p> <p>&nbsp;</p> <p><strong>Limitations</strong></p> <ul> <li>The analogy test could be a good indicator of the quality of word-embeddings, but it should be used with caution when comparing models trained on varying domains of data. It shall not be expected to&nbsp;generalize equally to all domains.</li> <li>The final score can be affected by the size, vocabulary, and domain of the text with which the models are trained on. For example, this may not be a good benchmark to compare models trained on news text <em>vs</em>&nbsp;posts on social media.</li> <li>Even though a manual sanity check was performed, we note that the semi-automatic construction of the Tigrinya test set might contains errors. If you discover any,&nbsp;you are welcome to contribute back by either opening an <em>Issue</em> at the GitHub repo,&nbsp;<a href="https://github.com/fgaim/tigrinya-analogy-test">https://github.com/fgaim/tigrinya-analogy-test</a>.</li> </ul> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this resource in your research, please cite it accordingly.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

MakeCode and CODAL: Intuitive and Efficient Embedded Systems Programming for Education (Artifact Evaluation)

<p>This artifact allows others to reproduce the results seen in this paper for MakeCode and CODAL, using the BBC micro:bit. The artifact contains an offline build environment for CODAL and MakeCode, allowing evaluators to test and build programs locally. In addition, we also provide espruino and micropython virtual machines to further increase repeatability of our results. Evaluators should download the virtual machine containing all pre-requisite tools, and use an oscilloscope to observe wave forms (used for timing) generated by the micro:bit, and a serial terminal to observe results reported from the micro:bit over serial.</p> <p>Full documentation is available at:&nbsp;https://lancaster-university.github.io/lctes-artefact-evaluation/</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Input data for performing a model evaluation of the sectional aerosol module SALSA embedded to PALM model system 6.0

<p>This dataset includes the input information applied to perform a model evaluation study of the PALM model system together with the sectional aerosol module SALSA.&nbsp;</p> <p>The content:</p> <ul> <li>PIDS_STATIC: building height and leaf area density data</li> <li>PIDS_AERO_&lt;simulation time&gt;_&lt;number of aerosol size bins&gt;: aerosol emission data as size bin specific surface emissions (level of detail 2) and aerosol background concentrations</li> <li>PIDS_CHEM_&lt;simulation time&gt;: emission data and background concentrations of gaseous compounds</li> </ul> <p>PIDS_STATIC contains static data and is therefore the same for all simulations.</p> <p>See the model documentation https://palm.muk.uni-hannover.de/trac/wiki/doc for further details.</p>

opencc-by-4.0Oct 2018View details →
zenodo36/100

GERBIL evaluation data of Robust and Collective Entity Disambiguation through Semantic Embeddings

<p>A table containing the SIGIR 2016 experiments performed with GERBIL in context of the SIGIR 2016 work &quot;Robust and Collective Entity Disambiguation through Semantic Embeddings&quot; by Stefan Zwicklbauer, Christin Seifert and Michael Granitzer</p> <p>It also contains the original URL to the GERBIL website</p> <p>Corresponding GitHub Repository:</p> <p>https://github.com/quhfus/</p> <p>&nbsp;</p>

opencc-by-4.0May 2016View details →
zenodo36/100

Jacdac: Service-based Prototyping of Embedded Systems (Artifact Evaluation)

<p>This artifact allows others to reproduce and explore the results seen in "Jacdac: Service-based Prototyping of Embedded Systems". The artifact contains a prebuilt docker image and the Dockerfile source used to produce the prebuilt docker image. Evaluators should follow the README contained in this artifact for complete instruction.</p>

opencc-by-4.0Mar 2024View details →
ClinicalTrials.gov32/100

'5 Rs to Rescue' A Cluster Trial With an Embedded Process Evaluation

ClinicalTrials.gov study NCT06997328. IPD Sharing: UNDECIDED. Countries: 1. Publications: 7.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Evaluation of Electronic Portal Messaging and Embedded Asynchronous Care on Physician-Assisted Smoking Quit Attempts

ClinicalTrials.gov study NCT05172219. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
geo24/100

Evaluating the restoration of DNA derived from archival formalin-fixed paraffin embedded tissues for genomic profiling by SNP-CGH analysis

GEO Series GSE43406. Homo sapiens. 43 samples. Type: Genome variation profiling by SNP array; SNP genotyping by SNP array.

openGEO-OpenJul 2013View details →
geo24/100

Systematic evaluation of RNA quality and microarray data reliability in human formalin-fixed paraffin-embedded and fresh frozen tissue samples

GEO Series GSE104562. Homo sapiens. 9 samples. Type: Expression profiling by array.

openGEO-OpenApr 2018View details →
geo24/100

Systematic evaluation of RNA quality and microarray data reliability in rat formalin-fixed paraffin-embedded and fresh frozen tissue samples

GEO Series GSE104561. Rattus norvegicus. 6 samples. Type: Expression profiling by array.

openGEO-OpenApr 2018View details →
ClinicalTrials.gov24/100

Bridging the Gap Between Theory and Practice: Evaluating Embedded Micro-Scenarios in Simulation on Academic Motivation and Clinical Decision-Making Among Nursing Students-A Randomized Controlled Trial

ClinicalTrials.gov study NCT07184112. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Clinical Evaluation of Localized Scalp Thread Embedding for Male Androgenetic Alopecia

ClinicalTrials.gov study NCT06879119. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
geo20/100

Genome-wide evaluation of change in lncRNA expression in non-functioning pituitary adenoma using formalin-fixed and paraffin-embedded tissue specimens

GEO Series GSE77517. Homo sapiens. 20 samples. Type: Expression profiling by array.

openGEO-OpenFeb 2016View details →
geo20/100

Systematic evaluation of RNA quality and microarray data reliability in formalin-fixed paraffin-embedded and fresh frozen tissue samples

GEO Series GSE104568. Homo sapiens; Rattus norvegicus. 15 samples. Type: Expression profiling by array.

openGEO-OpenApr 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record