Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
753
datasets available to search
ShareScore release 0.9.0
Dataset results
753 results for “metrics”
Fig. 5 in Do different sampling designs produce differences in the metrics of curimba, Prochilodus lineatus (Characiformes: Prochilodontidae)?
Fig. 5. Relationship between standard length (SL) and capture distance for curimba at fixed and variable sampling sites of Volta Grande (VGR) and Jaguara (JR) reservoirs.
Fig. 1 in Do different sampling designs produce differences in the metrics of curimba, Prochilodus lineatus (Characiformes: Prochilodontidae)?
Fig. 1. Maps of the regional location of the study area (1), of Grande River (2) and Volta Grande (3) and Jaguara (4) reservoirs, indicating the zones of the reservoirs (A, B and C) and fixed (diamond) and variable (black circle) sampling sites.
Fig. 2 in Do different sampling designs produce differences in the metrics of curimba, Prochilodus lineatus (Characiformes: Prochilodontidae)?
Fig. 2. Catch per unit effort (CPUE) of curimba for fixed and variable sampling sites by zone of Volta Grande (VGR) and Jaguara (JR) reservoirs.
Metric Extended Reduced monthly MSL data UK 1958-2018
<p>Extended and corrected Monthly mean sea level values at each tide gauge site in UK 1958 to 2018, corrected for levelling datum (RLR-equivalent), and also corrected for datum steps + seasonal + barotropic variability, as well as the Final Common Mode (average for UK).</p>
WEGE: A NEW METRIC FOR RANKING LOCATIONS FOR BIODIVERSITY CONSERVATION
<p>Here you will find all the maps produced in the paper "WEGE: A NEW METRIC FOR RANKING LOCATIONS FOR BIODIVERSITY CONSERVATION" that didn't make it into the main MS.</p> <p>It includes global grided maps at 100km x 100km and 20km x 20km of KBA triggering cells, EDGE, ED, WE, ER and our newly developed WEGE index for all amphibians, mammals and birds assessed by IUCN.</p>
Researchers Hidden Preferences for Metrics - Datasets
<p>Data collected during conjoint analysis experiments conducted as part of the *metrics-project. Detailed information on the experiments' background, conduction and results can be found in the following article from the Journal of the Association for Information Science and Technology: https://asistdl.onlinelibrary.wiley.com/doi/10.1002/asi.24445</p>
On the Importance and Shortcomings of Code Readability Metrics: A Case Study on Reactive Programming - replication package
<p>This is the replication package for the conference paper submission "On the Importance and Shortcomings of Code Readability Metrics: A Case Study on Reactive Programming"</p> <p><strong>Contents:</strong></p> <ul> <li>measurements.zip <ul> <li>DATASET_ORIGINAL.csv</li> <li>DATASET_REACTIVE.csv</li> </ul> </li> <li>source_code.zip <ul> <li> source_code_orig <ul> <li>Client.java</li> <li>Connection.java</li> <li>Server.java</li> <li>TcpConnection.java</li> <li>UdpConnection.java</li> </ul> </li> <li> source_code_rx <ul> <li>Client.java</li> <li>Connection.java</li> <li>Server.java</li> <li>TcpConnection.java</li> <li>UdpConnection.java</li> </ul> </li> </ul> </li> </ul> <p> </p>
Shortcomings of Event-Based Metrics [Supplement to PhD "Machine-Actionable Assessment of Research Data Products"]
<p>This deposition includes a tabular overview of all publications analyzed by my PhD to identify and classify shortcomings of event-based metrics for research data products. The tabular overview is linked via its field "Key" to the bibtex file which holds the bibliographic information to replicate the results shown in the table.</p> <p>The deposition is supplementary material to the dissertation "Machine-Actionable Assessment of Research Data Products" (Tobias Weber, yet unpublished), especially chapter 3.</p>
Replication package with data used in the study: The effect of code smells and design patterns on two change-related metrics: An exploratory study
<p>This is a replication package with data used in a study by T. Alkhaeir and B. Walter "The effect of code smells and design patterns on two change-related metrics: An exploratory study"</p> <p>This dataset contains the following folders:</p> <ul> <li>Aggregated Results Per System <ul> <li> For each subject system (AOI, Jedit, JHotDraw), we identify the following datasets: DP, nDP, S, nS ,SDP, nSDP, SnDP, and nSnDP. Each dataset is represented by a separate csv file.</li> <li> Those csv files include raw data about every class in every release, the csv files also include columns which represent: <ul> <li>- CHURN (CLPLPR(C)*100): defined as the sum of added and deleted lines in a class in a release, adjusted to the size of the class and to the number of revisions in the release;</li> <li>- and FREQ (MTPR(C)*100): defined as the average number of changes made to a class in a release, adjusted to the number of revisions in the release</li> </ul> </li> </ul> </li> <li>Detailed Results Per Smell Or Pattern <ul> <li> For each specific code smell (S) in each public release (Rel) of all subject systems, we identify SDP and SnDP datasets. Each dataset is in a separate .csv file</li> <li> For each specific design pattern (DP) in each public release (Rel) of all subject systems, we identify SDP and nSDP </li> </ul> </li> <li>Plots<br> We also include QQ plots for CHURN, FREQ values for every dataset in every system, that could serve as a supplementary data for the paper.</li> </ul>
Primary metric measurements of USA tree rings in one dataset in JSON format.
<p>The International Tree Rings Data Bank (ITRDB) is the most comprehensive tree growth database (https://www1.ncdc.noaa.gov/pub/data/paleo/treering).</p> <p>Shoudong Zhao, et al. (2019, 2018) analyzes the representativity of dendrochronological data (ITRDB) and proposes a corrected database with error indications. One of the bottlenecks of data use (ITRDB) is that the data is loaded as a collection of separate files in the Tucson positional format.</p> <p>The purpose of our data presentation is to change the Tucson data format to JSON format and combine the separate files into one.</p> <p>We convert the initial data for the USA of 2298 rwl-files into Json format of data on tree growth in one file. The data was converted using the R programming language and the dplR program library Bunn, A. (2008)</p> <p>The experience of developing the structure of dendroclimatic data in JSON format is described in the works of Kachaev A. (2016, 2017, 2020).</p> <p>Description of the structure of JSON data format is attached in the file ReadMe.pdf</p> <p> </p> <p><em>References</em></p> <p><em>Bunn, A. G. (2008). A dendrochronology program library in R (dplR). Dendrochronologia, 26, 115-124. https://doi.org/10.1016/j.dendro.2008.01.002</em></p> <p><em>Kachaev, Alexander (2020), "Compact dataset of dendrochronological data of pri-mary metric characteristics of tree rings of Asia.", Mendeley Data, V1, doi: 10.17632 / p9zhpmzgtk.1</em></p> <p><em>Kachaev A. V. (2017) Model for describing the structure of dendroclimatic data In the collection: Regional problems of remote sensing of the Earth Materials of the IV international scientific conference. Siberian Federal University, Institute of Space and Information Technologies. p. 120-122. (Russia)</em></p> <p><em>Kachaev A. V. (2016) NOSQL Approach for Development of Dendroclimatic Data Bank. In the collection: Regional problems of remote sensing of the Earth. Materials of the III International Scientific Conference. p. 89-91. (Russia)</em></p> <p><em>Shoudong Zhao, et al. (2019). The International Tree-Ring Data Bank (ITRDB) revisited: Data availability and global ecological representativity. Journal of Biogeography, 46 (2), 355-368. doi: 10.1111 / jbi.13488</em></p> <p><em>Zhao, Shoudong et al. (2018), Data from: The International Tree-Ring Data Bank (ITRDB) revisited: data availability and global ecological representativity, Dryad, Dataset, https://doi.org/10.5061/dryad.kh0qh06</em></p>
Data from: Validation of network communicability metrics for the analysis of brain structural networks.
Computational network analysis provides new methods to analyze the brain's structural organization based on diffusion imaging tractography data. Networks are characterized by global and local metrics that have recently given promising insights into diagnosis and the further understanding of psychiatric and neurologic disorders. Most of these metrics are based on the idea that information in a network flows along the shortest paths. In contrast to this notion, communicability is a broader measure of connectivity which assumes that information could flow along all possible paths between two nodes. In our work, the features of network metrics related to communicability were explored for the first time in the healthy structural brain network. In addition, the sensitivity of such metrics was analysed using simulated lesions to specific nodes and network connections. Results showed advantages of communicability over conventional metrics in detecting densely connected nodes as well as subsets of nodes vulnerable to lesions. In addition, communicability centrality was shown to be widely affected by the lesions and the changes were negatively correlated with the distance from lesion site. In summary, our analysis suggests that communicability metrics that may provide an insight into the integrative properties of the structural brain network and that these metrics may be useful for the analysis of brain networks in the presence of lesions. Nevertheless, the interpretation of communicability is not straightforward; hence these metrics should be used as a supplement to the more standard connectivity network metrics.
Data from: Social rank, color morph, and social network metrics predict oxidative stress in a cichlid fish
Dominance hierarchies are a fundamental part of social systems in many species and social rank can influence access to resources and impact health and physiology. While social subordination is a profound stressor, few studies consider the social stress experienced by dominant males due to constantly needing to defend their dominance status through costly aggressive displays. Recent studies suggest that in species that use body coloration to signal status, these costs may also be color morph-specific. Our study examines the link between the social rank, intensity of territorial defense, body coloration, and oxidative stress in males of the color polymorphic cichlid fish Astatotilapia burtoni where males are either blue or yellow. We studied behavior in naturalistic communities and examined circulating reactive oxygen metabolites and antioxidant defenses. We found that dominant males experience higher concentrations of circulating reactive oxygen metabolites without notably increasing their antioxidant defenses, but this effect was not related to color morph. Aggression and social network ties predicted oxidative stress in a morph-specific manner, with yellow but not blue males showing signs of increased oxidative damage with increasing agonistic effort. In contrast to expectation, oxidative stress was not influenced by cortisol or testosterone levels. We conclude that oxidative stress is instrumental to understanding the costs and benefits of high social rank.
Data from: Development and application of a novel metric to assess effectiveness of biomedical data
Objective: Design a metric to assess the comparative effectiveness of biomedical data elements within a study that incorporates their statistical relatedness to a given outcome variable as well as a measurement of the quality of their underlying data. Materials and methods: The cohort consisted of 874 patients with adenocarcinoma of the lung, each with 47 clinical data elements. The p value for each element was calculated using the Cox proportional hazard univariable regression model with overall survival as the endpoint. An attribute or A-score was calculated by quantification of an element's four quality attributes; Completeness, Comprehensiveness, Consistency and Overall-cost. An effectiveness or E-score was obtained by calculating the conditional probabilities of the p-value and A-score within the given data set with their product equaling the effectiveness score (E-score). Results: The E-score metric provided information about the utility of an element beyond an outcome-related p value ranking. E-scores for elements age-at-diagnosis, gender and tobacco-use showed utility above what their respective p values alone would indicate due to their relative ease of acquisition, that is, higher A-scores. Conversely, elements surgery-site, histologic-type and pathological-TNM stage were down-ranked in comparison to their p values based on lower A-scores caused by significantly higher acquisition costs. Conclusions: A novel metric termed E-score was developed which incorporates standard statistics with data quality metrics and was tested on elements from a large lung cohort. Results show that an element's underlying data quality is an important consideration in addition to p value correlation to outcome when determining the element's clinical or research utility in a study.
Data from: Practical performance of tree comparison metrics
The phylogenetic literature contains numerous measures for assessing differences between two phylogenetic trees. Individual measures have been criticized on various grounds, but little is known about their comparative performance in typical applications. We evaluate the performance of nine tree distance measures on two tasks: (1) distinguishing trees separated by lesser versus greater numbers of recombinations, and (2) distinguishing trees inferred with lower versus higher quality data. We find that when the trees being compared are similar, measures which make use of branch lengths are superior, with the branch-length version of the Robinson-Foulds metric (Robinson & Foulds, 1979) performing best. In contrast, for dissimilar trees topology-only measures are superior, with the Alignment metric of Nye et al. (2006) performing best. We also apply the measures to a mammalian data set and observe that the best metric depends on whether branch-length information is of interest. We give practical recommendations for choosing a tree distance metric in different applications.
Data from: U-Index, a dataset and an impact metric for informatics tools and databases
Measuring the usage of informatics resources such as software tools and databases is essential to quantifying their impact, value and return on investment. We have developed a publicly available dataset of informatics resource publications and their citation network, along with an associated metric (u-Index) to measure informatics resources' impact over time. Our dataset differentiates the context in which citations occur to distinguish between 'awareness' and 'usage', and uses a citing universe of open access publications to derive citation counts for quantifying impact. Resources with a high ratio of usage citations to awareness citations are likely to be widely used by others and have a high u-Index score. We have pre-calculated the u-Index for nearly 100,000 informatics resources. We demonstrate how the u-Index can be used to track informatics resource impact over time. The method of calculating the u-Index metric, the pre-computed u-Index values, and the dataset we compiled to calculate the u-Index are publicly available.
Data from: Joint allelic effects on fitness and metric traits
Theoretical explanations of empirically observed standing genetic variation, mutation, and selection suggest that many alleles must jointly affect fitness and metric traits. However, there are few direct demonstrations of the nature and extent of these pleiotropic associations. We implemented a mutation accumulation (MA) divergence experimental design in Drosophila serrata to segregate genetic variants for fitness and metric traits. By exploiting naturally occurring MA line extinctions as a measure of line-level total fitness, manipulating sexual selection, and measuring productivity we were able to demonstrate genetic covariance between fitness and standard metric traits, wing size and shape. Larger size was associated with lower total fitness and male sexual fitness, but higher productivity. Multivariate wing shape traits, capturing major axes of wing shape variation among MA lines, evolved only in the absence of sexual selection, and to the greatest extent in lines that went extinct, indicating that mutations contributing wing shape variation also typically had deleterious effects on both total fitness and male sexual fitness. This pleiotropic covariance of metric traits with fitness will drive their evolution, and generate the appearance of selection on the metric traits even in the absence of a direct contribution to fitness.
Data from: Resident species with larger size metrics do not recruit more offspring from the seed bank in old-field meadow vegetation
1. According to the traditional 'Size Advantage' (SA) hypothesis, plant species with larger body size are expected to be more successful when competition is intense, i.e. within severely crowded vegetation. Recent studies in old-field habitats, however, have shown that those species with greater numerical abundance as resident plants generally have a relatively small minimum reproductive threshold size (MIN), not a relatively large maximum potential body size (MAX). 2. In this study, we test for a size advantage in terms of species abundance representation in the soil seed bank, and we extend the SA hypothesis to include two additional size metrics: leaf size and seed size. Specifically, we ask, for resident species within a crowded old-field meadow: is larger seed size, leaf size, and/or body size associated with greater reproductive / recruitment success (i.e. number of germinable seeds within — and establishing plants emerging from — the soil seed bank)? We collected soil cores for a greenhouse experiment to record relative species abundances of germinable seeds in the seed bank, and we used a field experiment to record local abundances of species emerging from the resident seed bank within denuded plant neighbourhoods over three subsequent field seasons. 3. We found no general support for the SA hypothesis involving any of the size metrics, and none of the latter was a strong predictor of the number of germinable seeds emerging from soil cores in the greenhouse experiment. However, for species establishing in the field experiment from the seed bank over the three-year survey period, more abundant species in years 2 and 3 tended to be those with smaller MIN, and thus smaller MAX. In addition, within more crowded neighbourhoods, representation of reproductive plants was generally greater for species with relatively small MIN (and hence small MAX). 4. Synthesis. Our results extend support for the 'Reproductive Economy Advantage' hypothesis in old field habitats, to include not just established, largely undisturbed vegetation, but also very early stages of recruitment from seed within locally crowded plant neighbourhoods. Specifically, more successful species here are not those with relatively large potential body size (MAX); they are species capable of producing at least some offspring despite severe body size suppression, because they have a relatively small MIN.
Data from: Visual landmarks sharpen grid cell metric and confer context specificity to neurons of the medial entorhinal cortex
Neurons of the medial entorhinal cortex (MEC) provide spatial representations critical for navigation. In this network, the periodic firing fields of grid cells act as a metric element for position. The location of the grid firing fields depends on interactions between self-motion information, geometrical properties of the environment and nonmetric contextual cues. Here, we test whether visual information, including nonmetric contextual cues, also regulates the firing rate of MEC neurons. Removal of visual landmarks caused a profound impairment in grid cell periodicity. Moreover, the speed code of MEC neurons changed in darkness and the activity of border cells became less confined to environmental boundaries. Half of the MEC neurons changed their firing rate in darkness. Manipulations of nonmetric visual cues that left the boundaries of a 1D environment in place caused rate changes in grid cells. These findings reveal context specificity in the rate code of MEC neurons.
Supplementary data for the manuscript, "Comprehensive Assessment of Physiochemical Metrics for the Clustering of Adaptive Immune Repertoires"
<p>This repository contains data and code used in the manuscript "Comparative Assessment of Physiochemical Metrics for the Clustering of Adaptive Immune Receptor Repertoires" by Girgis et al. For additional information regarding how these data were used, please refer to our manuscript.</p> <p>Code is organized per figure in the main text. Each figure folder contains a 'script-inputs' folder and 'script-outputs' folder. The outputs folder is empty and may be populated with graphs and results tables by executing code within the directory. The inputs folder contains some pre-formatted data which may be used in executing code. Most of these inputs may be generated from scratch using raw data (ie the results of Homolig clustering on simulated repertoires) but may require substantial time and/or computational resources. Raw patient data used in Figure 6 (Pancreatic cancer anti-mKRAS TCRB repertoire clustering) and Figure 7 (Rheumatoid arthritis TCRB and IGH repertoires) are not included here. However, several graphs may be reproduced stripped of sequence-specific data. Pancreatic cancer patient repertoire data will be made available on dbGaP,study accession number phs003425.v1.p1. Rheumatoid arthritis patient data will be made available on ImmuneAccess, accession pending.</p> <p>To browse repo, first unzip all subdirectories: </p> <blockquote> <p><code>for f in *.zip; do</code><br><code> unzip "$f"</code><br><code>done</code></p> </blockquote> <p>All non-code files have been compressed to .gz format. To decompress, use: <code>gzip -dr *</code>, or <code>pigz -dr ./raw-data/* </code>for parallel decompression (recommended). </p> <p>Alexander Girgis <br>agirgis3@jhmi.edu <br>July 2025 </p>
Reduction steps for CPS metrics
<p>Reduction steps for CPS metrics</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.