Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,641

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,641 results for “similarity”

Learn how ShareScore rates datasets ↗
dryad40/100

Optimal sequence similarity thresholds for clustering of molecular operational taxonomic units in DNA metabarcoding studies

<p><span>Clustering approaches are pivotal to handle the many sequence variants obtained in DNA metabarcoding datasets, therefore they have become a key step of metabarcoding analysis pipelines. Clustering often relies on a sequence similarity threshold to gather sequences in Molecular Operational Taxonomic Units (MOTUs), each of which ideally representing a homogeneous taxonomic entity, e.g. a species or a genus. However, the choice of the clustering threshold is rarely justified, and its impact on MOTU over-splitting or over-merging even less tested. Here, we evaluated clustering threshold values for several metabarcoding markers under different criteria: limitation of MOTU over-merging, limitation of MOTU over-splitting, and trade-off between over-merging and over-splitting. We extracted sequences from a public database for nine markers, ranging from generalist markers targeting Bacteria or Eukaryota, to more specific markers targeting a class or a subclass (e.g. Insecta, Oligochaeta). Based on the distributions of pairwise sequence similarities within species and within genera, and on the rates of over-splitting and over-merging across different clustering thresholds, we were able to propose threshold values minimizing the risk of over-splitting, that of over-merging, or offering a trade-off between the two risks. For generalist markers, high similarity thresholds (0.96-0.99) are generally appropriate, while more specific markers require lower values (0.85-0.96). These results do not support the use of a fixed clustering threshold. Instead, we advocate a careful examination of the most appropriate threshold based on the research objectives, the potential costs of over-splitting and over-merging, and the features of the studied markers.</span></p>

opencc-zeroOct 2021View details →
dryad40/100

Simulated pollinator decline has similar effects on seed production of female and hermaphrodite Lobelia siphilitica, but different effects on selection on floral traits

<p><span>PREMISE:</span><span> Pollinator decline, by reducing seed production, is predicted to strengthen natural selection on floral traits. However, the effect of pollinator decline on gender dimorphic species (such as gynodioecious species, where plants produce female or hermaphrodite flowers) may differ between the sex morphs: if pollinator decline reduces the seed production of females more than hermaphrodites, then it should also have a larger effect on selection on floral traits in females than in hermaphrodites.</span></p> <p><span>RESULTS: </span><span>Experimentally reducing pollination decreased seed production of both females and hermaphrodites by ~21%. Reducing pollination also strengthened selection on floral traits, but this effect was not larger in females than in hermaphrodites. Instead, reducing pollination intensified selection for taller inflorescences in hermaphrodites, but did not intensify selection on any floral trait in females.</span></p> <p><span>CONCLUSIONS:</span><span> Our results suggest that pollinator decline will not have a larger effect on either seed production or selection on floral traits of female plants. As such, any effect of pollinator decline on seed production may be similar for gender dimorphic and monomorphic species. However, the potential for floral traits of females (and thus of gender dimorphic species) to evolve in response to pollinator decline could be limited.</span></p>

opencc-zeroNov 2022View details →
dryad40/100

Data for: Similar environmental cues guide timing of breeding and seasonal shifts in songbird social structure

<p>Seasonally breeding animals often exhibit different social structures during non-breeding and breeding periods that coincide with seasonal environmental variation. Therefore, ongoing climate change may play an important role in determining the future structure of animal societies, especially if climate determines when seasonal shifts in social structure occur. However, we know little about the environmental cues that determine the timing of seasonal shifts in social structure, a lack of knowledge that contrasts with our well-defined knowledge of the environmental cues that trigger a shift to breeding physiology in seasonally breeding species. Here we tested whether the environmental cues that drive seasonal shifts in social structure are similar to those that determine timing of breeding in the red-backed fairywren (<em>Malurus melanocephalus</em>), an Australian songbird. Social network analyses revealed that social groups, which are highly territorial during the breeding season, interact in social "communities" on larger ranges during the non-breeding season. Interactions among non-breeding groups were related to rainfall, with more rainfall leading to reductions in home range size and fewer interactions among non-breeding social groups. Similarly, onset of breeding was also determined by rainfall during the non-breeding season, with greater rainfall leading to earlier breeding. These findings reveal that for some species, the cues that determine the timing of shifts in social structure across seasonal boundaries can be similar to those that determine timing of breeding. This study increases our understanding of how social structure and the selection pressures that result from different social structures might respond to changing climates.</p>

opencc-zeroNov 2022View details →
zenodo40/100

Semantic Similarity of IT Support Tickets

<p>Collection of 300 support tickets manually labeled for semantic similarity, obtained from a IT support company in the Florian&oacute;polis (Brazil) region. Each ticket is represented by an unstructured text field, which is typed by the user that opened the call. The labeling process was performed in 2022 by three IT support professionals. The corpus contains tickets in many languages, mainly English, German, Portuguese and Spanish.</p> <p>All Personal Identifiable Information (PII) and sensitive information were removed (substituted by a tag indicating the original content, for instance: the sentence &quot;this text was written by Leonardo&quot; is converted to &quot;this text was written by [NAME]&quot;). The removal was performed in three steps: first, the automated machine learning-based tool AWS Comprehend PII Removal was used; then, a sequence of custom regular expressions was applied; last, the entire corpus was manually verified.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Figure data to "Continuous similarity transformation for critical phenomena: easy-axis antiferromagnetic XXZ model"

<p>This collection of data is complementary to the publication &quot;Continuous similarity transformation for critical phenomena: easy-axis antiferromagnetic XXZ model&quot;, Matthias R. Walther, Dag-Bj&ouml;rn Hering, G&ouml;tz S. Uhrig, Kai P. Schmidt, arXiv:2211.05689 (https://arxiv.org/abs/2211.05689).</p> <p>It contains the data points calculated by the method of Continuous Similarity Transformation(CST) used in Figures 3,5,6,7 and 8 in the CSV-Format.</p> <p>For details on the CST, the used error estimates and physical quantities we refer the the publication.</p> <p>For details on the format we recommend the README.md file.</p>

opencc-by-4.0Jan 2023View details →
dryad40/100

Detecting frequency-dependent selection through the effects of genotype similarity on fitness components

<p>Frequency-dependent selection (FDS) is an evolutionary regime that can maintain or reduce polymorphisms. Despite the increasing availability of polymorphism data, few effective methods are available for estimating the gradient of FDS from the observed fitness components. We modeled the effects of genotype similarity on individual fitness to develop a selection gradient analysis of FDS. This modeling enabled us to estimate FDS by regressing fitness components on the genotype similarity among individuals. We detected known negative FDS on the visible polymorphism in a wild <em>Arabidopsis</em> and damselfly by applying this analysis to single-locus data. Further, we simulated genome-wide polymorphisms and fitness components to modify the single-locus analysis as a genome-wide association study (GWAS). The simulation showed that negative or positive FDS could be distinguished through the estimated effects of genotype similarity on simulated fitness. Moreover, we conducted the GWAS of the reproductive branch number in <em>Arabidopsis thaliana</em> and found that negative FDS was enriched among the top-associated polymorphisms of FDS. These results showed the potential applicability of the proposed method for FDS on both visible polymorphism and genome-wide polymorphisms. Overall, our study provides an effective method for selection gradient analysis to understand the maintenance or loss of polymorphism.</p>

opencc-zeroFeb 2023View details →
zenodo40/100

Dataset for: Similarity scores of vibrational spectra reveal the atomistic structure of pentapeptides in multiple basins

<p>This dataset provides input/output files and scripts for the publication: Similarity scores of vibrational spectra reveal the atomistic structure of pentapeptides in multiple basins.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

US Patent Similarity Data

<p>Pairwise semantic similarity measures for US utility patents. Includes measures for citing/cited patent pairs, 100 most-similar patents for each patent, and doc2vec vectors for each patent. Second edition includes .npy file needed to generate new text embeddings using the pre-trained model.</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

Dataset Similarity Artikel Topic Modeling

<p>Dataset Similarity Artikel Topic Modeling.</p> <p>Data artikel diperoleh dari Google Scholar menggunakan Publish or Perish. Query yang digunakan &quot;Topic modeling&quot; dengan batas waktu publikasi sejak 2019. Keyword diisi dengan Mendeley auto update dan sisanya dicari secara manual. Data berupa buku dan data tanpa keyword dibuang. Data dipasangkan secara kombinatorik dari 50 data menjadi 1225. Labelling manual.</p>

openother-openApr 2023View details →
zenodo40/100

FIG. 7 in Oospore features among morphologically similar and closely related charophyte species: consistency and variability

FIG. 7. — The relationship between oospore parameters of Chara baueri A.Braun and Chara braunii C.C.Gmel., and variable that represents these two species from all sampling localities. Redundancy analysis (RDA). Abbreviations: see Material and methods.

opencc-zeroNov 2022View details →
zenodo40/100

FIG. 5 in Oospore features among morphologically similar and closely related charophyte species: consistency and variability

FIG. 5. — The relationship between oospore parameters of Chara "connivens" P.Salzmann ex A.Braun and Chara globularis Thuil., and nominal variable referring to these two species: A, considering all sampling localities; B, referring to these two species considering Dulin pond locality only. Redundancy analysis (RDA). Abbreviations: see Material and methods.

opencc-zeroNov 2022View details →
zenodo40/100

FIG. 6 in Oospore features among morphologically similar and closely related charophyte species: consistency and variability

FIG. 6. — Oospore parameters of Chara braunii C.C.Gmel. in relation to localities where this species was found. Redundancy analysis (RDA). Abbreviations: see Material and methods.

opencc-zeroNov 2022View details →
zenodo40/100

FIG. 3 in Oospore features among morphologically similar and closely related charophyte species: consistency and variability

FIG. 3. — Oospores of selected charophyte species: A, Chara globularis Thuil.; B, Chara "connivens" P.Salzmann ex A.Braun; C, Chara baueri A.Braun; D, Chara braunii C.C.Gmel. Scale bars: 200 µm.

opencc-zeroNov 2022View details →
zenodo40/100

FIG. 4 in Oospore features among morphologically similar and closely related charophyte species: consistency and variability

FIG. 4. — Oospore parameters of Chara globularis Thuil. in relation to localities where this species was found. Redundancy analysis (RDA). Abbreviations: see Material and methods.

opencc-zeroNov 2022View details →
zenodo40/100

DBsimilarity of Natural Products to aid compound identification on MS and NMR pipelines, similarity networking and more

<p>DBsimilarity is proposed for organizing structure databases into Similarity Networks to assist researchers with analyzing and making sense of chemical data available. It converts SDF files into CSV files when needed, adds chemoinformatics data, constructs a MZMine custom database file and a NMRfilter candidate list of compounds for rapid dereplication of MS and 2D NMR data, calculates similarities, and constructs CSV files for Similarity Networks using Cytoscape. DBsimilarity aims to bridge chemoinformatics to laboratory-focused natural products researchers and students.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Supplementary data for a systematic literature review on source code similarity measurement and clone detection: techniques, applications, and challenges

<p>The Microsoft Excel files containing the supplementary data and diagrams for the paper:</p> <p><strong>A systematic literature review on source code similarity measurement and clone detection: techniques, applications, and challenges</strong></p> <p>The article is under review in the Journal of Systems and Software.</p> <p>In this version, the literature search has been performed between April and May 2023, and studies published until that date has been mentioned in the attached Excel files.</p> <p>This version (v3.4.0) corresponds to the third revision (R3) of the manuscript.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
dryad40/100

Data from: Maximum mutational robustness in genotype-phenotype maps follows a self-similar blancmange-like curve

<div class="section abstract"> <p>Phenotype robustness, defined as the average mutational robustness of all the genotypes that map to a given phenotype, plays a key role in facilitating neutral exploration of novel phenotypic variation by an evolving population. By applying results from coding theory, we prove that the maximum phenotype robustness occurs when genotypes are organised as bricklayer's graphs, so called because they resemble the way in which a bricklayer would fill in a Hamming graph. The value of the maximal robustness is given by a fractal continuous everywhere but differentiable nowhere sums-of-digits function from number theory. Interestingly, genotype-phenotype (GP) maps for RNA secondary structure and the HP model for protein folding can exhibit phenotype robustness that exactly attains this upper bound. By exploiting properties of the sums-of-digits function, we prove a lower bound on the deviation of the maximum robustness of phenotypes with multiple neutral components from the bricklayer's graph bound, and show that RNA secondary structure phenotypes obey this bound. Finally, we show how robustness changes when phenotypes are coarse-grained and derive a formula and associated bounds for the transition probabilities between such phenotypes.</p> </div>

opencc-zeroJul 2023View details →
dryad40/100

Similar parasite communities but dissimilar infection patterns in two closely related chickadee species

<p>Haemosporidian parasite communities are broadly similar in Boulder County, CO between two common songbirds –– the Black-capped Chickadee (<em>Poecile</em> <em>atricapillus</em>) and Mountain Chickadee (<em>Poecile</em> <em>gambeli</em>). However, Mountain Chickadees appear more likely to be infected with <em>Plasmodium</em> and potentially experience higher infection burdens with <em>Leucocytozoon</em> in contrast to Black-capped Chickadees. We found that elevation change (and associated ecology) drives the distributions of these parasite genera. For Boulder County chickadees, environmental factors play a more important role in structuring haemosporidian communities than host evolutionary differences. However, evolutionary differences are likely key to shaping the probability of infection, infection burden, and whether an infection remains detectable over time. We found that for recaptured birds, their infection status (i.e., presence or absence of detectable parasite infection) tends to remain consistent across capture periods. We sampled 234 chickadees between 2017–2021 across a ~1500-meter elevation gradient from low elevation (i.e., the city of Boulder) to comparatively high elevation (i.e., the CU Boulder Mountain Research Station). It is unknown whether long-term haemosporidian abundance trends have changed over time in our sampling region. However, we ask whether potentially disparate patterns of <em>Plasmodium</em> susceptibility and <em>Leucocytozoon</em> infection burden could be playing a role in the negative population trends of Mountain Chickadees.</p>

opencc-zeroJul 2023View details →
zenodo40/100

Codon similarity data in ATTED-II ver 8.0 (Ptr, Zma)

<p>Codon similarity data in ATTED-II ver 8.0</p> <p>The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).</p> <p>Protein-coding sequences utilized in this study were retrieved from NCBI&#39;s RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.</p>

opencc-by-4.0Aug 2015View details →
dryad40/100

Data from: An environmental habitat gradient and within-habitat segregation enable co-existence of ecologically similar bird species

<p>Niche theory predicts that ecologically similar species can co-exist through multidimensional niche partitioning. However, due to the challenges of accounting for both abiotic and biotic processes in ecological niche modelling, the underlying mechanisms that facilitate co-existence of competing species are poorly understood. In this study, we evaluated potential mechanisms underlying the co-existence of ecologically similar bird species in a biodiversity-rich transboundary montane forest in east-central Africa by computing niche overlap indices along an environmental elevation gradient, diet, forest strata, activity patterns, and within-habitat segregation across horizontal space. We found strong support for abiotic environmental habitat niche partitioning, with 55% of species pairs having separate elevation niches. For the remaining species pairs that exhibited similar elevation niches, we found that within-habitat segregation across horizontal space and to a lesser extent vertical forest strata provided the most likely mechanisms of species co-existence. Co-existence of ecologically similar species within a highly diverse montane forest was determined primarily by abiotic factors (e.g., environmental elevation gradient) that characterize the Grinnellian niche and secondarily by biotic factors (e.g., vertical and horizontal segregation within habitats) that describe the Eltonian niche. Thus, partitioning across multiple levels of spatial organization is a key mechanism of co-existence in diverse communities.</p>

opencc-zeroJul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record