Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,641

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,641 results for “similarity”

Learn how ShareScore rates datasets ↗
zenodo36/100

SAS: Semantic Artist Similarity Dataset

<p>The Semantic Artist Similarity dataset consists of two datasets of artists entities with their corresponding biography texts, and the list of top-10 most similar artists within the datasets used as ground truth. The dataset is composed by a corpus of 268 artists and a slightly larger one of 2,336 artists, both gathered from Last.fm in March 2015. The former is mapped to the MIREX Audio and Music Similarity evaluation dataset, so that its similarity judgments can be used as ground truth. For the latter corpus we use the similarity between artists as provided by the Last.fm API. For every artist there is a list with the top-10 most related artists. In the MIREX dataset there are 188 artists with at least 10 similar artists, the other 80 artists have less than 10 similar artists. In the Last.fm API dataset all artists have a list of 10 similar artists.&nbsp;</p> <p>There are 4 files in the dataset.</p> <p><strong>mirex_gold_top10.txt</strong>&nbsp;and&nbsp;<strong>lastfmapi_gold_top10.txt</strong>&nbsp;have the top-10 lists of artists for every artist of both datasets. Artists are identified by MusicBrainz ID. The format of the file is one line per artist, with the artist mbid separated by a tab with the list of top-10 related artists identified by their mbid separated by spaces.</p> <p>artist_mbid \t artist_mbid_top10_list_separated_by_spaces \n</p> <p><strong>mb2uri_mirex</strong>&nbsp;and&nbsp;<strong>mb2uri_lastfmapi.txt</strong>&nbsp;have the list of artists. In each line there are three fields separated by tabs. First field is the MusicBrainz ID, second field is the last.fm name of the artist, and third field is the DBpedia uri.</p> <p>artist_mbid \t lastfm_name \t dbpedia_uri \n</p> <p>There are also 2 folders in the dataset with the biography texts of each dataset. Each .txt file in the biography folders is named with the MusicBrainz ID of the biographied artist. Biographies were gathered from the Last.fm wiki page of every artist.</p> <p><strong>Using this dataset</strong></p> <p>We would highly appreciate if scientific publications of works partly based on the Semantic Artist Similarity dataset quote the following publication:</p> <blockquote> <p>Oramas, S.,&nbsp;Sordo M.,&nbsp;Espinosa-Anke L., &amp;&nbsp;Serra X.&nbsp;(In Press).&nbsp;&nbsp;<a href="http://mtg.upf.edu/node/3316">A Semantic-based Approach for Artist Similarity</a>.&nbsp;16th International Society for Music Information Retrieval Conference.</p> </blockquote> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p>&nbsp;</p> <p>https://www.upf.edu/web/mtg/semantic-similarity</p>

opencc-by-4.0Oct 2015View details →
zenodo36/100

alawinia/provClustering: Discovering Similar Workflows via Provenance Clustering

<p>Several workflow management systems and scripting languages have adopted provenance tracking, yet many researchers choose to manually capture or instrument their processing scripts to write provenance information to files. The Next Generation Sequencing (NGS) project we are associated with is tracking provenance in such manner. The NGS project is a collaboration between multiple groups at different sites, where each group is collecting and processing samples using an agreed-upon workflow. The workflow contains many stages with varying degrees of complexity. Over time workflow stages are modified, but data samples are only comparable when processed with identical versions of the workflow. However, for various reasons (including the distributed nature of the collaboration) it is not always clear which samples have been processed with which version of the workflow. In this paper, we introduce new techniques for clustering provenance datasets and attempt to discover the ones that are likely to be generated by same workflow. Based on the clustering result, users can identify similar provenance and would be able to categorize them into different clusters for debugging and zoom-in/zoom-out viewing.</p>

openother-openJul 2018View details →
zenodo36/100

A crop yield change emulator for use in GCAM and similar models: Persephone v1.0

<p>This is an archive of the raw data and analysis source code for the paper &quot;A crop yield change emulator for use in GCAM and similar models: Persephone v1.0&quot;.&nbsp; The archive contains:</p> <ul> <li><strong>data.zip:</strong>&nbsp;All source code for analysis, input data for analysis, and results of analysis</li> <li><strong>persephone.proj&nbsp;:&nbsp;</strong>R project for ease of reproducing analysis</li> </ul>

opencc-by-4.0Sep 2018View details →
zenodo36/100

SeSaMe: A Data Set of Semantically Similar Java Methods

<p>This is the data set presented in the paper</p> <p>Kamp, M., Kreutzer P., Philippsen M.: SeSaMe: A Data Set of Semantically<br> Similar Java Methods. 16th International Conference on Mining Software<br> Repositories (MSR 2019), Montreal, QC, Canada. 2019</p>

opencc-by-4.0Feb 2019View details →
zenodo36/100

The RESPECT Trade Similarity Index

<p>We use detailed product-level trade data from Eurostat to construct a Trade Similarity Index for each Member State of the European Union and each third country. The Trade Similarity Index takes values between 0 and 1. A higher value indicates that the trade pattern of the Member State is more similar to the average of the European Union. A value of 0 means that there is no product overlap between the exports (or imports) of the Member State and the rest of the EU. A value of 1 means that the Member State exports (or imports) every product in the same proportion as the rest of the EU.</p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

Measuring similarity between gene interaction profiles

<p>This repository contains the following:</p> <p>Original genetic interaction score matrices for the genes. Also contains a binarized version of this data based on a threshold of 0.05, for use in the binary distance measures.</p> <p>The distance measure matrices that were calculated based on the genetic interaction score matrices; Pearson, Maryland bridge, Ochiai, and Braun-Blanquet matrices for the one-square case and the same for the two-squares case.</p> <p>Lists of all the genes used.</p> <p>Lists of modules for each of the eight distance measure matrices, with the genes for each.</p> <p>Summary statistics for the two-squares distance matrices, network and module summary statistics for the two-squares modules, and stats on how many modules for each of the eight different module clusterings could be mapped to subnetworks found in Tong et al. (reference 14 in the paper).</p> <p>A figure containing distribution of distance measure values for the two-square distance matrices.</p> <p>View the readme file for full details.</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

Fig. 1 in Description of two new species similar to Anolis insignis (Squamata: Iguanidae) and resurrection of Anolis (Diaphoranolis) brooksi

Fig. 1. Anolis insignis, male, Pocosol, Alajuela, Costa Rica.

opencc-by-4.0Jul 2017View details →
zenodo36/100

Fig. 4. Metastrongylus spp. eggs from a in Lungworms (Metastrongylus spp.) and intestinal parasitic stages of two separated Swiss wild boar populations north and south of the Alps: Similar parasite spectrum with regional idiosyncrasies

Fig. 4. Metastrongylus spp. eggs from a wild boar faecal sample.

opencc-by-4.0Apr 2021View details →
zenodo36/100

LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models - Replication Package

<p>LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models</p> <p>This is the replication package associated with the paper "LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models".</p> <p><strong>Replication Package Contents:</strong></p> <p>This replication package contains all the necessary data and code required to reproduce the results reported in the paper. We provide the results of the Fault Detection Rate (FDR), Total Minimization Time (MT), Time Saving Rate (TSR) , statistical tests for all the minimization budgets (i.e., 25%, 50%, and 75%), results for the preliminary study, results for UniXcoder/Cosine with preprocessed code on 16 projects.</p> <p><strong>Data:</strong></p> <p>We provide in the <em><strong>Data</strong></em> directory the data used in our experiments, which is the source code of test cases (Java test methods) of 17 projects collected from Defects4J.</p> <p><strong>Code:</strong></p> <p>We provide in the<em> <strong>Code</strong> </em>directory the code (Python) and bash files required to run the experiments and reproduce the results.</p> <p><strong>Results:</strong></p> <p>We provide in the<em> <strong>Results </strong></em>directory the detailed results for our approach (called LTM). We also provide the summarized results of LTM and a baseline (ATM) for comparison purposes. Additional technical details about ATM can be found at https://zenodo.org/record/7455766.</p> <p><strong>_________________________________</strong></p> <p><strong>LTM's Similarity Measurement:</strong></p> <p>The source code of this step is in the <strong><em>Code/LTM/Similarity</em></strong> directory.</p> <p><strong>Requirements:</strong></p> <p>To run this step, Python 3 is required (we used Python 3.10). Also, the required libraries in the <em><strong>Code/LTM/Similarity/requirements.txt</strong></em> file should be installed, as follows:</p> <p>cd Code/LTM/Similarity</p> <p>pip install -r requirements.txt</p> <p><strong>Input:</strong></p> <ul> <li>Data/LTM/TestMethods</li> </ul> <p><strong>Output:</strong></p> <ul> <li>Data/LTM/similarity_measurements</li> </ul> <p><strong>Running the experiment:</strong></p> <p>To measure the similarity between all pairs of test cases, the following bash script should be executed:</p> <p>bash measure_similarity.sh</p> <p>The source code of test methods of each project in the <strong><em>Data/LTM/TestMethods</em></strong> is parsed to generate pairs of test cases. This steps includes test methods tokenization, test methods embeddings extraction and similarity calculation. Then, all similarity scores are stored in <em><strong>Data/LTM/similarity_measurements</strong></em> folder. Due to the large size of the calculated similarity scores (60 GB), they were not uploaded on Zenodo, but they can be available upon request.</p> <p><strong>LTM's Test Suite Minimization:</strong></p> <p>The source code of this step is in the Code/LTM/Search directory.</p> <p><strong>Requirements:</strong></p> <p>To run this step, Python 3 is required (we used Python 3.10). Also, the required libraries in the <strong><em>Code/LTM/Search/requirements.txt</em></strong> file should be installed, as follows:</p> <p>cd Code/LTM/Search</p> <p>pip install -r requirements.txt</p> <p><strong>Input:</strong></p> <p>Data/LTM/similarity_measurements</p> <p><strong>Output:</strong></p> <p>Results/LTM/minimization_results</p> <p><strong>Running the experiments:</strong></p> <p>To minimize the test suite for each project version, the following bash script should be executed:</p> <p>bash minimize.sh</p> <p>The similarity scores of all test case pairs per project version are parsed by the search algorithm (Genetic Algorithm). Each experiment runs ten times using three minimization budgets (25%, 50%, and 75%). The results are stored in the <em><strong>Results/LTM/minimization_results</strong></em> directory.</p> <p><strong>LTM's Evaluation:</strong></p> <p>To evaluate the minimization results for each version and each project, the following bash script should be executed:</p> <p>cd Code/LTM/Evaluation</p> <p>bash evaluate_per_version.sh</p> <p>cd Code/LTM/Evaluation</p> <p>bash evaluate_per_project.sh</p> <p>This will evaluate the FDR, MT and TSR results for each version and each project for each minimization budget. These results are stored in the <em><strong>Results/LTM</strong></em> directory.</p> <p>Note that for each version, the FDR is either 1 or 0. For each project, the FDR ranges from 0 to 1.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Updated dataset from "Exploring seed density and limiting similarity to reduce invasive grass performance for grassland restoration purposes"

<p>Here we aimed to compare the effect of two seed mixes and three density sowing treatments on the performance of the invasive grass Eragrostis plana, one of the major threats to the Campos Sulinos grasslands, at South Brazil. The experiment was carried out in a greenhouse experiment in Porto Alegre, Brazil. The seed mixes have the same species, but differ in terms of species abundance.</p> <p>Paper Thomas et al. (2024) "Exploring seed density and limiting similarity to reduce invasive grass performance for grassland restoration purposes", published at Applied Vegetation Science. <a href="https://doi.org/10.1111/avsc.12804">https://doi.org/10.1111/avsc.12804</a></p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Similarity of CRISPR genes grouped by chromosome location with chromosome arm correction

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo36/100

Similarity of CRISPR genes grouped by chromosome location without chromosome arm correction

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo36/100

Divergence time and environmental similarity predict the strength of morphological convergence in stick and leaf insects

<p>This uploads contains the datasets, phylogenetic tree and associated R code used to generate the results reported in the article: "Divergence time and environmental similarity predict the strength of morphological convergence in stick and leaf insects" published in Proceedings of the National Academy of Sciences USA (2024).<br>A detailed explanation of datasetS1 can be found in the supplementary data of the article.&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

How to assess similarities and differences between mantle circulation models and Earth using disparate independent observations: Data and Analysis

<p>Dataset includes simulation output produced by a TERRA simulation for `How to assess similarities and differences between mantle circulation models and Earth using disparate independent observations'.&nbsp;</p> <p>&nbsp;</p> <h3><strong>Description of data file contents</strong></h3> <ul> <li><strong>NC*comp.tar.gz</strong> - compressed archives containing NetCDF files (file-per-process) with TERRA grid data including temperature, velocity, interpolated bulk composition, denisty, and voscosity fields. Can be read using <a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>. *dump number</li> <li><strong>NC_seis_037.tar.gz</strong> - compressed archive containing NetCDF files (file-per-process) with predicted seismic properties at the resolution of the TERRA grid generated from the present day state of the simulated mantle, including elastic and anelastic Vs and Vp, bulk sound velocity, and predicted density from mineral phyiscs tables. Can be read using&nbsp;<a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>.</li> <li><strong>NC_hpes_037.tar.gz</strong> - compressed archive containing NetCDF files (file-per-process) with interpolated abundances at the resolution of the TERRA grid for isotopes including the heat-producing elements ^40^K, ^232^Th, ^235^U and ^238^U.&nbsp;Can be read using&nbsp;<a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>.</li> <li><strong>P_files_037.tar.gz</strong> - compressed archive of TERRA P-files (particle files).</li> <li><strong>C_files_037.tar.gz</strong> - compressed archive of TERRA C-files (grid state files) - together with the P-files describe the full present day state of the simulation.&nbsp;</li> <li><strong>seis_filtered_037.tar.gz</strong> - compressed archive (file-per-layer) with reparameterised and seismically filtered (against S40RTS) present day Vs field.</li> <li><strong>seis_tables.tar.gz</strong> - compressed archive contianing lookup tables of seismic properties for the 3 principal lithologies assumed in the TERRA simulation (harzburgite, lherzolite and basaltic crust).</li> <li><strong>density_037.sph</strong> - Spherical harmonic coefficients for the density field in format to be read by the <a title="propagator" href="https://zenodo.org/records/12696774" target="_blank" rel="noopener">propagator matrix code.</a></li> <li><strong>plumes.pkl, ridges.pkl</strong> - Files containing tracer particle information for particles associated with plumes and ridges.&nbsp;</li> <li><strong>plumes_ridges.py</strong> - Python script containing example code for reading and plotting plumes.pkl and ridges.pkl files.&nbsp;</li> <li><strong>ptcls_rdgs_plms.py, interrogate_particles.py</strong> - Python script and module containing required functions for carrying out post processing routine generating the plumes.pkl and ridges.pkl files. Requires&nbsp;<a title="terratools" href="https://github.com/mantle-convection-constrained/terratools" target="_blank" rel="noopener">terratools</a>.</li> <li><strong>hst.dat </strong>- Time series of key simulation properties including mantle temperature profile used for calcualting CMB heat flux.&nbsp;</li> <li><strong>terra, interra</strong> - TERRA executable and input parameter file.</li> <li><strong>pyflowng.zip</strong> - Compressed directory containing version of the `pyflowng` code used in this work.</li> <li><strong>mode_splitting_methods.zip</strong> - Compressed directory contianing synthetic splitting function predictions and maps.</li> </ul> <p>&nbsp;</p> <h3><strong>Dump Numbers</strong></h3> <p>Below is a table of dump numbers (final three digits of file names) and the corresponding model times.</p> <table> <tbody> <tr> <td><strong>Dump number&nbsp;</strong></td> <td><strong>Model time (Ma)</strong></td> </tr> <tr> <td>037</td> <td>0 (present day)</td> </tr> <tr> <td>027</td> <td>10</td> </tr> <tr> <td>026</td> <td>20</td> </tr> <tr> <td>025</td> <td>30</td> </tr> <tr> <td>024</td> <td>40</td> </tr> <tr> <td>023</td> <td>50</td> </tr> <tr> <td>022</td> <td>60</td> </tr> <tr> <td>021</td> <td>70&nbsp;</td> </tr> <tr> <td>020</td> <td>80</td> </tr> <tr> <td>019</td> <td>90</td> </tr> <tr> <td>018</td> <td>100</td> </tr> </tbody> </table>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Measurement, self-similarity, and TNT equivalence of blasts from exploding wires

<p>This dataset contains raw overpressure time histories from blasts generated from exploding aluminum wires, recorded under controlled conditions at five energy levels and six standoff distances. The time histories, referenced in the accompanying study, provide detailed information about the shock wave profiles specific to aluminum-wire explosions.</p>

opencc-by-nc-1.0Nov 2024View details →
zenodo36/100

Forager mobility and lithic discard probability similarly affect the distance of raw material discard from source - Supplemental Material

<p>The data, R code and Netlogo model used in &quot;Forager mobility and lithic discard probability similarly affect the distance of raw material discard from source&quot;.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Dataset related to the article "Cardiac Biomarkers and Autoantibodies in Endurance Athletes: Potential Similarities with Arrhythmogenic Cardiomyopathy Pathogenic Mechanisms"

<p>This record contains raw data related to the article &ldquo;Cardiac Biomarkers and Autoantibodies in Endurance Athletes: Potential Similarities with Arrhythmogenic Cardiomyopathy Pathogenic Mechanisms&quot;.&nbsp;</p> <p>The &quot;Extreme Exercise Hypothesis&quot; states that when individuals perform training beyond the ideal exercise dose, a decline in the beneficial effects of physical activity occurs. This is due to significant changes in myocardial structure and function, such as hemodynamic alterations, cardiac chamber enlargement and hypertrophy, myocardial inflammation, oxidative stress, fibrosis, and conduction changes. In addition, an increased amount of circulating biomarkers of exercise-induced damage has been reported. Although these changes are often reversible, long-lasting cardiac damage may develop after years of intense physical exercise. Since several features of the athlete&#39;s heart overlap with arrhythmogenic cardiomyopathy (ACM), the syndrome of &quot;exercise-induced ACM&quot; has been postulated. Thus, the distinction between ACM and the athlete&#39;s heart may be challenging. Recently, an autoimmune mechanism has been discovered in ACM patients linked to their characteristic junctional impairment. Since cardiac junctions are similarly impaired by intense physical activity due to the strong myocardial stretching, we propose in the present work the novel hypothesis of an autoimmune response in endurance athletes. This investigation may deepen the knowledge about the pathological remodeling and relative activated mechanisms induced by intense endurance exercise, potentially improving the early recognition of whom is actually at risk.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Data for: Simultaneous GPS-tracking of parents reveals a similar parental investment within pairs, but no immediate co-adjustment on a trip-to-trip basis

<p>This repository contains data for the paper: Kavelaars et al. 2021. Simultaneous GPS-tracking of parents reveals a similar parental investment within pairs, but no immediate co-adjustment on a trip-to-trip basis.&nbsp;<strong>Movement Ecology</strong>.&nbsp;https://doi.org/10.1186/s40462-021-00279-1</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Figure 30. A–I, Serpula nudiradiata n in Descriptions of New Serpulid Polychaetes from the Kimberleys of Australia and Discussion of Australian and Indo-West Pacific Species of Spirobranchus and Superficially Similar Taxa

Figure 30. A–I, Serpula nudiradiata n.sp., from holotype, AM W202942: (A–E)

opencc-by-4.0Nov 2009View details →
zenodo36/100

Figure 25. A–C, Hydroides trihamulatus n in Descriptions of New Serpulid Polychaetes from the Kimberleys of Australia and Discussion of Australian and Indo-West Pacific Species of Spirobranchus and Superficially Similar Taxa

Figure 25. A–C, Hydroides trihamulatus n.sp.—an older specimen from AM W202943: (A) anterior end

opencc-by-4.0Nov 2009View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record