Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
753
datasets available to search
ShareScore release 0.9.0
Dataset results
753 results for “metrics”
Effects of Charring on Squash (Cucurbita L ) Seed Morphology and Compression Strength: Implications for Paleoethnobotany Metric Data
<p>Metric data resulting from a series of charring experiments on seeds from three species of squash: <em>Cucurbita pepo</em>, <em>Cucurbita moschata</em>, and <em>Cucurbita maxima</em>.</p>
Indicators and metrics in local climate adaptation plans
<p>This dataset gathers information related to indicators and metrics collected from 11 local climate adaptation plans in worldwide cities - Athens (Greece), Auckland (USA), Barcelona (Spain), Glasgow (UK), Istanbul (Turkey), Lima (Peru), Los Angeles (USA), Montreal (Canada), Nagoya (Japan), New York City (USA), Portland (USA), Tokyo (Japan) and Vancouver (Canada). The dataset describes the use and characteristics of adaptation indicators and metrics across climate adaptation-related planning documents. Although the sample is relatively small, it is a reflection of the global embryonic stage of adaptation metrics practice.</p> <p>The results of the analysis of this database have been published in:</p> <p>Goonesekera, S. M., & Olazabal, M. (2022). Climate adaptation indicators and metrics: State of local policy practice. <em>Ecological Indicators</em>, <em>145</em>, 109657. <a href="https://doi.org/10.1016/j.ecolind.2022.109657">https://doi.org/10.1016/j.ecolind.2022.109657</a> (OPEN ACCESS)</p>
A benchmark of gene expression tissue-specificity metrics
<p>Supplementary figures and data to the paper "A benchmark of gene expression tissue-specificity metrics"</p> <p><em>Briefings in Bioinformatics</em>, Volume 18, Issue 2, March 2017, Pages 205–214, <a href="https://doi.org/10.1093/bib/bbw008">https://doi.org/10.1093/bib/bbw008</a></p> <p>Previously published at FigShare, republishing because of access problems for some researchers.</p>
Microservice Security Detectors & Metrics & Detection Strategies: Dataset
<p>This is the dataset for replicability for the article "Detection Strategies for Microservice Security Tactics." It provides the code needed to replicate the study in the article and the model data set of 10 system models and 20 variants of those models.</p> <p> </p> <p>The abstract of the article is:</p> <p> </p> <p>Microservice architectures are widely used today to implement distributed systems. Securing microservice architectures is challenging because of their polyglot nature, continuous evolution, and various security concerns relevant to such architectures. This article proposes a novel, model-based approach providing detection strategies to address the automated detection of security tactics (or patterns and best practices) in a given microservice architecture decomposition model. Our novel detection strategies are metrics-based rules that decide conformance to a security recommendation based on a statistical predictor. The proposed approach models this recommendation using Architectural Design Decisions (ADDs). We apply our approach for four different security-related ADDs on access management, traffic control, and avoiding plaintext sensitive data in the context of microservice systems. We then apply our approach to a model data set of 10 open-source microservice systems and 20 variants of those systems. Our results are detection strategies showing a very low bias, a very high correlation, and a low prediction error in our model data set.<br> <br> The dataset is based on a dataset from a previous article: https://zenodo.org/record/6424722<br> </p>
Supporting dataset for tertiary study on source code metrics
<p>The dataset includes search results (SearchResultsFromScopusIEEEACM June23.xlsx) and compiled results from the included secondary studies (Extracted Data For Tertiary Study June23 15 SS.xlsx). </p> <p>Search results description: Please use Figure 1 from paper to trace the data made available. The excel file contains the following data</p> <p>1) validation set of studies used to quasi-gold standard validation of search string. </p> <p>2) known set of papers used to formulate the search string</p> <p>3) search results from databased Scopus, IEEE, and ACM</p> <p>4) secondary studies from tertiary study on code smells by Lacerda et al.</p> <p>5) Combined results from all sources</p> <p>6) papers removed after duplicates removal</p> <p>7) Preliminary search results</p> <p>8) Search Strings used for Scopus, IEEE, ACM</p> <p>Extracted Data Description: Extracted data contains the following:</p> <p>1) meta data and characteristics of secondary studies (title, author, year of publication)</p> <p>2) Raw data on evidence on maintainability, reliability and security from included set of studies</p> <p>3) compiled data on evidence on maintainability, reliability and security from included set of studies </p> <p>4) Description of prediction models used</p> <p>5) Description of source code metrics</p> <p>6) Secondary studies that used metrics but reported no evidence and secondary studies removed due to low DARE score</p> <p>7) Quality assessment score of included secondary studies</p> <p>8) Quality assessment questions for (additional quality assessment score)AQAS</p> <p> </p> <p> </p>
Sound stimuli: Considerations for the perceptual evaluation of steady-state and time-varying sounds using psychoacoustic metrics
<p>The current dataset provides all the sound stimuli (wav files) to recreate the figures from the study with the title "Considerations for the perceptual evaluation of steady-state and time-varying sounds using psychoacoustic metrics" by the same authors. There are three datasets: <strong>dataset 1</strong> for fly-by aircraft sounds (directory <em>1-Aircraft-fly-by</em>), <strong>dataset 2</strong> for pass-by train sounds (directory <em>2-Train-pass-by</em>), and <strong>dataset 3</strong> for sounds from a resonating tube called hummer (directory <em>3-Hummer-resonances</em>). A brief description of the datasets is included. The reproduction of figures requires the installation of the sound quality analysis toolbox (SQAT) for MATLAB.</p> <p><strong>Use these data</strong>:<br> Download all these data, locate them in a local directory of your computer. If you have MATLAB and you downloaded a local copy of the <strong>SQAT toolbox</strong> (open access at: <a href="https://github.com/ggrecow/SQAT">https://github.com/ggrecow/SQAT</a>) you can recreate the figures of our paper. After downloading and initialising the toolbox (type 'startup_SQAT;', without quotation marks in MATLAB), run the script <strong>copy_sounds_for_pub_Osses2023c_to_SQAT.m</strong>. This last step is not required but recommended. Subsequently, you can run the toolbox script pub_Osses2023c_Forum_Acusticum_SQAT.m and follow the instructions on screen.</p>
The eco-conscious wind turbine: design beyond purely economic metrics
<p>Figures from the publication <em>The eco-conscious wind turbine: design beyond purely economic<br> metrics</em>.</p>
Understanding the robustness of spectral-temporal metrics across the global Landsat archive from 1984-2019 – a quantitative evaluation: extended material
<p>This dataset contains extended material for the paper:</p> <p>Frantz, D., Rufin, P., Janz, A., Ernst, S., Pflugmacher, D., Schug, F., Hostert, P.<strong>: Understanding the robustness of spectral-temporal metrics across the global Landsat archive from 1984-2019 – a quantitative evaluation. </strong><em>In revision.</em></p>
Data from: Scoutknife: A naïve, whole genome informed phylogenetic robusticity metric
<p>The phylogenetic bootstrap, first proposed by Felsenstein in 1985, is a critically important statistical method in assessing the robusticity of phylogenetic datasets. Core to its concept was the use of pseudosampling - assessing the data by generating new replicates derived from the initial dataset that was used to generate the phylogeny. In this way, phylogenetic support metrics could overcome the lack of perfect, infinite data. With infinite data, however, it is possible to sample smaller replicates directly from the data to obtain both the phylogeny and its statistical robusticity in the same analysis. Due to the growth of whole genome sequencing, the depth and breadth of our datasets have greatly expanded and are set to only expand further. With genome-scale datasets comprising thousands of genes, we can now obtain a proxy for infinite data. Accordingly, we can potentially abandon the notion of pseudosampling and instead randomly sample small subsets of genes from the thousands of genes in our analyses. Here, we introduce Scoutknife, a jackknife-style subsampling implementation that generates 100 datasets by randomly sampling a small number of genes from an initial large-gene dataset to jointly establish both a phylogenetic hypothesis and assess its robusticity. Using 18 previously published datasets and 100 simulation studies, we show that Scoutknife is conservative and informative as to conflicts and incongruence across the whole genome, without the need for subsampling based on traditional model selection criteria.</p>
Extending the application of connectivity metrics within the framework of the characterization of the dynamic behaviour of a WDS subjected to users' activity
<p>Water distribution networks (WDNs) are complex combinations of nodes and links, and the current tendency is to modify their topological structure through the closure of isolation valves for monitoring and water quality reasons. For their analysis, several approaches based on graph theory have recently been proposed, mainly considering steady-state flow conditions. However, in their real functioning, WDNs are continuously subjected to pressure transients generated by manoeuvres on regulation devices or by users’ activity. This study investigates the application of some metrics from graph theory, already used in the context of steady-state analysis, for assessing the effects of changes in the topological structure of a network ‒ due for example to sectorization or branching operations ‒ on its transient response when subjected to manoeuvres on devices such as hydrants, pumps, etc. or users’ activity. The analysis shows that some connectivity metrics can effectively reflect the dynamic pressure behaviour of the network and, thus, provide useful indications for design and management operations taking into account unsteady flow features.</p>
Dataset of UAI 2021 Paper "An Unsupervised Video Game Playstyle Metric via State Discretization"
<p>This is a part of dataset of the paper published in UAI 2021 (37th Conference on Uncertainty in Artificial Intelligence).</p> <p>Including training datasets, testing datasets, and HSD models of three game platforms used in the paper: TORCS, RGSK, and Atari.</p> <p>The example program for using this file will be put on the author's github repo branch: <a href="https://github.com/DSobscure/cgi_drl_platform/tree/playstyle_uai2021">https://github.com/DSobscure/cgi_drl_platform/tree/playstyle_uai2021</a></p>
PoqueiraOccupancy: Dataset and metrics from occupancy sensors of urban areas and establishments in the region of Barranco del Poqueira in the Alpujarra Granadina
<p>This dataset is linked to the analysis of different aspects related to the conservation of the Sierra Nevada National Park through advanced digital systems. The devices have been deployed in the municipalities of Pampaneira and Capileira in the Alpujarra region of the province of Granada. The data is collected by 11 BOSCH fixed cameras 11.00 387.4900 Interior IR 5.3 MP and 4 TURRET type cameras Interior IR Lens 2.8 mm 5.3 MP 100º H.265 multi-streaming (H.265; H.264; M-JPEG).</p> <p>The devices have been installed as follows: 3 devices have been placed in establishments in Capileira, 8 in establishments in Pampaneira, and 4 in urban passage areas in the municipality of Pampaneira. All devices are capable of measuring the entry and exit to the establishment or area they are designated for. In some cases, there is also a metric which measures the number of people present within that area. The information related to the establishments has been anonymized to ensure the privacy of the collaborating companies in the project and the flow of customers during the studied period.</p> <p>The data attached in the CSV files DATA_OCCUPANCY_2022 and DATA_OCCUPANCY_2023 contain information about individuals detected by the cameras in the years 2022 (from February to December) and 2023 (from January to August). The calculation of the number of people is done cumulatively in hourly intervals. The collected variables include:</p> <ul> <li> <p>device_ID: The name of the device recording the value.</p> </li> <li> <p>type: The metric measuring the recording, which can be ENTRADA (entry), SALIDA (exit), or AFORO (occupancy).</p> </li> <li> <p>date: The date and time at which the cumulative people count is recorded for the specific metric.</p> </li> <li> <p>counter: The number of people counted for a specific metric in that time period.</p> </li> </ul>
Example calculation of E1[h1] contribution to the source for second-order metric perturbations of a Schwarzschild black hole
<p>This repository contains data for the h1 an dr0/h1 perturbations that can be used to compute a piece of the source for the second-order metric perturbation. The Mathematica notebook 'SecondOrderE1h1.nb' shows how to combine the data to compute E1[h1].</p> <p>The h1 data was computed using the h1Lorenz code that is available in the Black Hole Perturbation Toolkit (https://github.com/BlackHolePerturbationToolkit/h1Lorenz). The dr0/dh1 data was computed by Leanne Durkan following the method detailed in "Slow evolution of the metric perturbation due to a quasicircular inspiral into a Schwarzschild black hole" by Leanne Durkan and Niels Warburton, arXiv:2206.08179</p> <p>Authors: Leanne Durkan, Niels Warburton</p>
Assessing Computational Notebook Understandability through Code Metrics Analysis
<p>Computational notebooks have become the primary coding environment for data scientists. Despite their popularity, research on the code quality of these notebooks is still in its infancy, and the code shared in these notebooks is often of poor quality. Considering the importance of maintenance and reusability, it is crucial to pay attention to the comprehension of the notebook code and identify the notebook metrics that play a significant role in their comprehension. The level of code comprehension is a qualitative variable closely associated with the user's opinion about the code. Previous studies have typically employed two approaches to measure it. One approach involves using limited questionnaire methods to review a small number of code pieces. Another approach relies solely on metadata, such as the number of likes and user votes for a project in the software repository. In our approach, we enhanced the measurement of the understandability level of notebook code by leveraging user comments within a software repository. As a case study, we started with 248,761 Kaggle Jupyter notebooks introduced in previous studies and their relevant metadata. To identify user comments associated with code comprehension within the notebooks, we utilized a fine-tuned DistillBERT transformer. We established a \emph{user comment based criterion} for measuring code understandability by considering the number of code understandability-related comments, the upvotes on those comments, the total views of the notebook, and the total upvotes received by the notebook. This criterion has proven to be more effective than alternative methods, making it the ground truth for evaluating the code comprehension of our notebook set. In addition, we collected a total of 34 metrics for 10,857 notebooks, categorized as script-based and notebook-based metrics. These metrics were utilized as features in our dataset. Using the Random Forest classifier, our predictive model achieved 85% accuracy in predicting code comprehension levels in computational notebooks, identifying developer expertise and markdown-based metrics as key factors.</p> <p> </p>
PoqueiraAR: Dataset and Metrics from Augmented Reality App in Barranco del Poqueira in the Alpujarra Granadina
<p>This dataset is related to an Augmented Reality application implemented in the Sierra Nevada National Park, specifically in the Barranco del Poqueira. This application has been designed to enhance the experience of visitors walking a circular path that connects the three villages of the ravine: Pampaneira, Bubión and Capileira, offering seven points of interest.</p><p>The dataset covers the period from June 2022 to October 2023 and includes information on the number of downloads, download dates, types of devices used to download the application, geographic origin of the devices, attractions visited and most demanded audiovisual resources.</p><p>These data provide an objective view of the effectiveness of the application in improving the visitor experience and its impact on the tourism promotion of Barranco del Poqueira. In addition, they provide relevant information about user preferences, which can guide future updates and improvements in the Augmented Reality application. The application contributes to enrich the visit to these villages and provides useful data for the sustainable management of tourism in this region of Sierra Nevada.</p>
Data from: Resilience metrics are robust across data qualities but sensitive to community size models
Open the record for dataset details and reuse information.
Repository Analytics and Metrics Portal (RAMP) 2021 data
Open the record for dataset details and reuse information.
The Dayhoff Exchange Score: A new metric to quantify site saturation in amino acid datasets prior to phylogenetic analysis
Open the record for dataset details and reuse information.
Data from: Effects of taxon sampling and tree reconstruction methods on phylodiversity metrics
Open the record for dataset details and reuse information.
Group and individual social network metrics are robust to changes in resource distribution in experimental populations of forked fungus beetles
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.