Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

448

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

448 results for “analogy”

Learn how ShareScore rates datasets ↗
zenodo40/100

Source Data for Manuscript: "Retrievals Applied To A Decision Tree Framework Can Characterize Earth-like Exoplanet Analogs"

<p>This dataset accompanies the manuscript entitled: "Retrievals Applied To A Decision Tree Framework Can Characterize Earth-like Exoplanet Analogs", which was accepted for publication in the Planetary Science Journal. Included are the source files for all figures included in the paper.</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Early Irish Analogy Dataset for Word Embedding Evaluation

<p>An embedding evaluation dataset for Early Irish described in the paper "<a href="https://aclanthology.org/2023.insights-1.10.pdf">Do not Trust the Experts: How the Lack of Standard Complicates <span>NLP</span> for Historical <span>I</span>rish</a>".</p> <p>Traditionally, analogy datasets are based on pairwise semantic proportion, and therefore every question has a single correct answer. Given the high level of variation in historical languages, such a strict definition of a correct answer seems unjustified. Therefore, Early Irish Analogy Dataset follows the <a href="https://vecto.space/projects/BATS/">Bigger Analogy Test Set (BATS)</a> and provides several correct answers to each analogy question.&nbsp;</p> <p>Morphological and spelling variation data are extracted from the <a href="https://dil.ie/">eDIL</a>, a historical dictionary of medieval Irish. Unlike BATS, no distinction is made between inflection types due to eDIL's structure. The raw data amounted to 2,370 spelling variation and 9,690 morphological variation questions, from which 150 examples were randomly selected for each of the subsets to be comparable in size with the synonym and antonym subsets. The synonym and antonym subsets are translations of the correspondent BATS parts obtained by reverse-searching the eDIL and proofread by four expert evaluators. The dataset includes 98 entries in the synonym subset and 109 entries in the antonym subset, upon which three or more experts agreed.</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Activity cliffs with dual-atom replacements and single-atom analogs

<p>From the ChEMBL database, 852 activity cliffs (ACs) with dual-atom replacements were extracted which were formed by compounds with high-confidence activity data. Each AC captured an at least 10-fold difference in compound potency. For a subset of these ACs, analogs with corresponding single-atom replacements were identified. The dual-atom ACs and available single-atom replacement analogs were provided (SMILES representation and ChEMBL compound ID). For each AC compound and analog, targets from ChEMBL are reported (with UniProt ID). The target shared by all associated compounds represents the primary AC target&nbsp; &nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Dataset Friebus-Kardash et al, A chemerin peptide analog stimulates tumor growth in two xenograft mouse models of human colorectal carcinoma

<p>Dataset Friebus-Kardash et al, A chemerin peptide analog stimulates tumor growth in two xenograft mouse models of human colorectal carcinoma</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Tele-manipulation dataset of the Analog-1 experiment

<p>This dataset contains log data and matlab functions.</p> <p>The data was collected during the main space mission Analog-1.</p> <p>The data serves the analysis of telemanipulation with force feedback under KU-forward space communication using the Time Domain Passivity Approach for High Delays (TDPA-HD).</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Dataset for "Investigation of Mars analog candidates around Syowa Station, Antarctica"

<p>This dataset includes coordinate values of rock- and water-sampling points,&nbsp;water quality data, near-infrared spectrum data, subsurface hardness data, and photos obtained by the project AAS6304 &quot;Investigation of Mars analog candidates around Syowa Station, Antarctica&quot; during the 63rd Japanese Antarctic Research Expedition.</p> <p>Measurement and sampling points and water quality data are shown in &quot;JARE63_AAS6304_data_sample_list.xlsx&quot;.</p> <p>&quot;NIRScan_data.zip&quot; is the near-infrared data obtained using NIR-S-G1, a NIR spectrometer (InnoSpectra Corporation, http://www.inno-spectra.com/en/product).</p> <p>&quot;Aerial_photos.zip&quot; includes photos taken from the helicopter.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Tigrinya Analogy Test for evaluating Word Embeddings

<p><strong>Tigrinya Analogy Test for evaluating Word Embeddings</strong></p> <p>This is a Tigrinya version of the Google Analogy Test set, which is used to evaluate English word-embedding models.&nbsp;The analogy test is a well-established strategy to empirically evaluate the quality of word-embedding models. More information about the English task can be found at the <a href="https://aclweb.org/aclwiki/Google_analogy_test_set_(State_of_the_art)">ACL Wiki</a>.</p> <p>This data is&nbsp;&nbsp;was first machine&nbsp;translated then&nbsp;manually verified by a native speaker to reduce errors.</p> <p>Some aspects of the original analogy test is focused on English and may not transfer well to other languages, such as those related to grammar or morphology. Therefore, we have discarded examples that became irrelevant in Tigrinya when adapting the task. Finally, there are a total of <strong>18465</strong>&nbsp;entries in the Tigrinya Analogy Test set, while the source English data has <strong>19544</strong>&nbsp;entries.</p> <p>An entry is dropped if the translations led to one of the following conditions:</p> <ol> <li>If the source word pair map to one Tigrinya word, for example, lucky &amp; luckiest both correspond to ዕድለኛ.</li> <li>If the source word results in a multi-word expression. For example, grandson (ወዲ ጓል / ወዲ ወዲ), granddaughter (ጓል ጓል / ጓል ወዲ). This because the typical word-embedding approaches such as <em>word2vec</em> are not designed to predict multi-word phrases.</li> </ol> <p>&nbsp;</p> <p><strong>Test Sections</strong></p> <p>The test includes a series of semantic and syntactic analogies divided up into subsections including world capitals,&nbsp;currencies, family, tense, and plurality. The test contains the following sections:</p> <ol> <li>capital-world</li> <li>currency</li> <li>city-in-state</li> <li>family</li> <li>gram1-adjective-to-adverb</li> <li>gram2-opposite</li> <li>gram3-comparative</li> <li>gram4-superlative</li> <li>gram5-present-participle</li> <li>gram6-nationality-adjective</li> <li>gram7-past-tense</li> <li>gram8-plural</li> <li>gram9-plural-verbs</li> </ol> <p>&nbsp;</p> <p><strong>Examples:</strong></p> <ul> <li>Semantic section of World Capitals: &ldquo;ኣስመራ: ኤርትራ as ፓሪስ: ?&rdquo; and if the model responds correctly it will return: &ldquo;ፈረንሳ&rdquo;.</li> <li>Semantic section of Family section: &ldquo;ሰብኣይ: ሰበይቲ as ወዲ: ጓል&rdquo;.</li> <li>Syntax section with tense, a sample analogy might be &ldquo;Walk: Walked as Run: Ran&rdquo;.<br> &nbsp;</li> </ul> <p><strong>Evaluation</strong></p> <p>The final accuracy of a model is the proportion of the questions that the model answers correctly.<br> Generally, a better-quality model would answer more questions correctly than a model of lower quality.<br> However, note that a model with low performance on this analogy test, might still contain useful information, but may not be robust or good enough for more complex tasks.</p> <p>&nbsp;</p> <p><strong>Limitations</strong></p> <ul> <li>The analogy test could be a good indicator of the quality of word-embeddings, but it should be used with caution when comparing models trained on varying domains of data. It shall not be expected to&nbsp;generalize equally to all domains.</li> <li>The final score can be affected by the size, vocabulary, and domain of the text with which the models are trained on. For example, this may not be a good benchmark to compare models trained on news text <em>vs</em>&nbsp;posts on social media.</li> <li>Even though a manual sanity check was performed, we note that the semi-automatic construction of the Tigrinya test set might contains errors. If you discover any,&nbsp;you are welcome to contribute back by either opening an <em>Issue</em> at the GitHub repo,&nbsp;<a href="https://github.com/fgaim/tigrinya-analogy-test">https://github.com/fgaim/tigrinya-analogy-test</a>.</li> </ul> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this resource in your research, please cite it accordingly.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

DCM containing analog series

<p>The file contains 1400 analog series that consist of DCM compounds from PubChem and their bioactive analogs from ChEMBL. For all 14,796 analogs in these series, analog series ID (AS_ID), ChEMBL_ID or PubChem compound identifier (PubChem_cid) and SMILES are provided. The targets corresponding to ChEMBL compounds in the series are also provided (ChEMBL_ID_targets). In addition, DCM compounds are indicated.</p>

opencc-by-4.0Sep 2017View details →
zenodo40/100

Raw diffraction images of endothelin ETB receptor bound to clinical antagonist bosentan and its analog

<p>Diffraction images of endothelin ET<sub>B</sub> receptor bound to bosentan (PDB code 5XPR) and K-8794 (5X93).</p> <p>Bosentan is an oral medication approved for the treatment of pulmonary arterial hypertension. K-8794 is the ETB-selective high-affinity analog of bosentan.</p> <p>Two datasets of 5X93 were collected with helical method, while for 5XPR 16 small-wedge (10°/crystal) datasets were collected automatically using ZOO system. In both cases the diffraction images were collected from loop-harvested microcrystals using MX225HS CCD detector at a wavelength of 1 Å on BL32XU, SPring-8.</p> <p>For 5XPR, 14 datasets were merged at 3.6 Å resolution in the published result (Shihoya et al. NSMB 2017) using KAMO; see processing note: https://github.com/keitaroyam/yamtbx/wiki/Processing-ETBR-bonsentan-data-(5XPR)</p> <p>NOTE</p> <ul> <li> <p>Most frames have (relatively weak) lipid rings.</p> </li> </ul> <ul> <li> <p>There is the indexing ambiguity problem in 5XPR (space group P3<sub>2</sub>21) that should be resolved before merging.</p> </li> <li> <p>5xpr_bosentan/bosentan-1425-01/multi_002_000019.img.bz2 is missing (probably due to a detector problem)</p> </li> </ul>

opencc-by-4.0Aug 2017View details →
zenodo40/100

Collection of analog series-based (ASB) scaffolds

<p>The entire collection of 23,791 unique ASB scaffolds generated from compounds from Probes and Drugs Portal (PDP) and ChEMBL (version 23) is reported. Each ASB scaffold is provided in canonical SMILES representation, the database origin (DB_origin) is specified, and the number of analogs (#analogs) the scaffold represents reported. In addition, for each ASB scaffold from ChEMBL, unique target annotations of corresponding analogs are provided using UniProt target identifiers. For ASB scaffolds from PDP, collected compound annotations are provided. For scaffolds shared between PDP and ChEMBL the number of analogs is provided in the form 'X|Y' where 'X' and 'Y' denote the number of analogs in PDP and ChEMBL, respectively. Scaffolds are rank-ordered according to the number of analogs they represent. </p>

opencc-by-4.0Dec 2016View details →
zenodo40/100

Collection of analog series-based (ASB) scaffolds shared between ZINC, ChEMBL, and PubChem

<p>Analog series-based (ASB) scaffolds shared between ZINC and ChEMBL (version 22), ZINC and PubChem and all the three databases are provided as three separate files. For each ASB scaffold, the SMILES representation of ZINC compounds is provided. In addition, the number of ZINC compounds, the number and the list of targets it was annotated with is reported. A README file is also given.</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

Analog Series of Compounds with High Frequency of Activity in Screening Assays

<p>A set of 6941 analog series and associated data are provided. These series exclusively consist of compounds that are most frequently active across public screening assays.&nbsp; &nbsp;</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

ARN: Analogical Reasoning on Narratives

<p>As a core cognitive skill that enables the transferability of information across domains, analogical reasoning has been extensively studied for both humans and computational models. However, while cognitive theories of analogy often focus on narratives and study the distinction between surface, relational, and system similarities, existing work in natural language processing has a narrower focus as far as relational analogies between word pairs. This gap brings a natural question: can state-of-the-art large language models (LLMs) detect system analogies between narratives? To gain insight into this question and extend word-based relational analogies to relational system analogies, we devise a comprehensive computational framework that operationalizes dominant theories of analogy, using narrative elements to create surface and system mappings. Leveraging the interplay between these mappings, we create a binary task and benchmark for Analogical Reasoning on Narratives (ARN), covering four categories of far (cross-domain)/near (within-domain) analogies and disanalogies. We show that while all LLMs can largely recognize near analogies, even the largest ones struggle with far analogies in a zero-shot setting, with GPT4.0 scoring below random. Guiding the models through solved examples and chain-of-thought reasoning enhances their analogical reasoning ability. Yet, since even in the few-shot setting, the best model only performs halfway between random and humans, ARN opens exciting directions for computational analogical reasoners.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Supplementary figures for "Dark matter distribution in Milky Way-analog galaxies"

<p>Among the attached files, you will find:</p> <p>- All figures from the article in high-quality PDF format, ordered by name as follows: 'Figure1.pdf' corresponds to Figure 1 of the article, and so on;</p> <p>- Moment 0 (intensity), 1 (velocity), and 2 (dispersion) maps of the atomic hydrogen gas (HI) distribution for each galaxy in our sample. All moment maps were generated from our three-dimensional modeling with 3D-Barolo;</p> <p>- For each galaxy, we show the extended version of Figure 1 from the paper, which includes the intensity, velocity, and dispersion maps for the model and residual of each galaxy. For instance, 'NGC3521_kinematics.pdf' corresponds to the kinematic maps of NGC 3521;</p> <p>- Stellar distribution maps at 3.6 and 4.5 &mu;m provided by the S4G survey. For instance, 'NGC3521.phot.1.fits' corresponds to the 3.6 &mu;m image, while 'NGC3521.phot.2.fits' corresponds to the 4.5 &mu;m image.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Dataset: Analog Devices, Inc. (ADI) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
dryad40/100

Data from: Machine learning without a processor: Emergent learning in a nonlinear analog network

<p>The capabilities of digital artificial neural networks grow rapidly with their size, however the time and energy required to train them does as well. The tradeoff is far better for Brains, where the constituent parts (neurons) update their analog connections in ignorance of the actions of other neurons, eschewing centralized processing. Recently introduced analog electronic <em>contrastive local learning networks </em>(CLLNs) share this important decentralized property. However their capabilities were limited because existing implementations are linear. In this dataset we include experimental demonstrations of a nonlinear CLLN, establishing a new paradigm for scalable learning. Included here are data and scripts required to generate figures 2-6 of the manuscript titled "Machine learning without a processor: Emergent learning in a nonlinear analog network".</p>

opencc-zeroJul 2024View details →
zenodo40/100

FIGURE 4 in What are the best modern analogs for ancient South American mammal communities? Evidence from ecological diversity analysis (EDA)

FIGURE 4. Linear regression of MAP on correspondence axis 1 (CA1) score; estimated MAP for each of fossil locality based on CA1 score is indicated. Abbreviations: LV, La Venta; QH, Quebrada Honda; RU, Rümikon; SC, Santa Cruz; TG, Tinguiririca.

opencc-by-4.0Dec 2020View details →
zenodo40/100

FIGURE 6. Classification Tree results and predictions for the five fossil localities. A in What are the best modern analogs for ancient South American mammal communities? Evidence from ecological diversity analysis (EDA)

FIGURE 6. Classification Tree results and predictions for the five fossil localities. A) Results and predictions for CT1, vegetative cover. B) Results and predictions for CT2, biogeographic realm. Abbreviations: LV, La Venta; QH, Quebrada Honda; RU, Rümikon; SC, Santa Cruz; TG, Tinguiririca.

opencc-by-4.0Dec 2020View details →
zenodo40/100

FIGURE 3 in What are the best modern analogs for ancient South American mammal communities? Evidence from ecological diversity analysis (EDA)

FIGURE 3. Axes three and four of the correspondence analysis. A) Positions of the fossil localities and 179 modern ecoregions; B) Positions of the 22 variables. Note that the scale is not the same in the two graphs.

opencc-by-4.0Dec 2020View details →
zenodo40/100

FIGURE 2 in What are the best modern analogs for ancient South American mammal communities? Evidence from ecological diversity analysis (EDA)

FIGURE 2. Axes one and two of the correspondence analysis. A) Positions of the fossil localities and 179 modern ecoregions; B) Positions of the 22 variables. Note that the scale is not the same in the two graphs.

opencc-by-4.0Dec 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record