Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
753
datasets available to search
ShareScore release 0.9.0
Dataset results
753 results for “metrics”
Quantifying representativeness in RCTs using ML fairness metrics - Data and codes
<p>The "Quantifying representativeness in RCTs using ML fairness metrics - Data and codes" is used to quantify representativeness in randomized clinical trials (RCTs) and provide insights to improve the clinical trial equity and health equity. We developed RCT representativeness metrics based on Machine Learning (ML) Fairness Research. Visualizations and statistical tests based on proposed metrics enable researchers and physicians to rapidly visualize and assess subgroup representation in RCTs. The approach enables users to determine underrepresentation, absence, or other misrepresentation of subgroups indicating potential limitations of RCTs. The method could help support generalizability evaluation of existing RCT cohorts, enrollment target decisions for new RCTs (if eligibility criteria are included), and monitoring of RCT enrollment, ultimately contributing to more equitable public health outcomes. We apply the proposed RCT representativeness metrics to three landmark clinical trials released in the last decade: Action to Control Cardiovascular Risk in Diabetes (ACCOD), Antihypertensive and Lipid-Lowering Treatment to Prevent Heart Attack Trial (ALLHAT), and Systolic Blood Pressure Intervention Trial (SPRINT). This dataset contains the processed data and results for the experiments and visualization codes in the paper titled "Quantifying representativeness in randomized clinical trials using machine learning fairness metrics."</p>
Radiomics and Artificial Intelligence Analysis with Textural Metrics Extracted by Contrast-Enhanced Mammography in the Breast Lesions Classification.
<p>We uploaded the database of 104 lesions included in the manuscript: Fusco R, Piccirillo A, Sansone M, Granata V, Rubulotta MR, Petrosino T, Barretta ML, Vallone P, Di Giacomo R, Esposito E, Di Bonito M, Petrillo A. Radiomics and Artificial Intelligence Analysis with Textural Metrics Extracted by Contrast-Enhanced Mammography in the Breast Lesions Classification. Diagnostics (Basel). 2021 Apr 30;11(5):815. doi: 10.3390/diagnostics11050815. PMID: 33946333; PMCID: PMC8146084.</p>
Online program metrics and evaluation for FMNP and NATA
<p>Restrictions on public gatherings in early 2020 due to the COVID-19 pandemic resulted in cancellation of in-person outreach programs offered by the Florida Master Naturalist Program and Natural Areas Training Academy, two successful University of Florida extension programs that provide natural history and resource management training to lay and professional audiences. In response, both programs rapidly transitioned to blended or 100% online educational methods to continue offering courses and maintain program operations. To assess participant responses to these changes, we used surveys and course registry data to evaluate and compare course enrollment, satisfaction, and outcomes among courses with new online formats to courses offered prior to the COVID-19 pandemic. We also examined logistical challenges and key programmatic elements that facilitated the transition of both programs to increased reliance on online education. Course participants responded favorably to classes offered online. Our results revealed an audience exists for online programming, that satisfaction with online courses was high and comparable to that measured for in-person courses, and that online approaches effectively transferred knowledge and promoted behavior change in participants. The transition to online programming required investments of time, energy, and in some cases, direct costs. However, this transition was greatly facilitated by the existence of well-defined program protocols, educational curricula, strong partnerships, and feedback mechanisms for both programs. Long-term investments in program structure, partnerships, and support systems enabled both programs to be resilient and adaptable and successfully implement online programming in response to the COVID-19 pandemic.</p>
MIRA-Datasets: Datasets from Metrics for Intercomparison of Remapping Algorithms
<p>The Metrics for Intercomparison of Remapping Algorithms (<a href="https://github.com/CANGA/MIRA">MIRA</a>) project provides the Python drivers for the intercomparison study to enable the computation of metrics for different remapping algorithms of interest in ESM.</p> <p>The dataset repository contains three groups of artifacts: the original test cases used in the study and the output metrics data from four different remapping algorithms, along with some helpful scripts to compare the metrics data. Details are provided below.</p> <ol> <li> <p>All of the input meshes, sampled reference data on the meshes for several uniformly refined resolutions, and regionally refined cases are contained within the <code>Meshes</code> directory.</p> <ul> <li>The uniformly refined meshes for Cubed-Sphere (CS), polygonal quasi-uniform MPAS and Regular Latitude-Longitude (RLL) meshes along with sampled field data for five different fields are provided in <code>Meshes/UniformlyRefined/</code> directory.</li> <li>The regionally refined meshes for CS and MPAS meshes around continental-US (CONUS) region with the sampled reference field data is available under <code>Meshes/RegionallyRefined</code> directory.</li> </ul> </li> <li> <p>The input meshes provided under <code>Meshes</code> directory were used to perform a remapping intercomparison study that analyzed the key numerical metrics to gain better insight into the behavior of remapping algorithms, and to compare several key properties under a unified framework. Four different remapping algorithms were considered in this study.</p> <ul> <li> <p>Earth System Modeling Framework (ESMF) Regrid</p> </li> <li> <p>TempestRemap high-order conservative maps</p> </li> <li> <p>Generalized Moving-Least-Squares (GMLS) algorithm</p> <ul> <li>A variation with the Clip-And-Assured-Sum (CAAS) algorithm to enforce bounds preservation</li> </ul> </li> <li> <p>Weighted-Least-Squares Essentially Non-oscillatory Remap (WLS-ENOR) scheme</p> <p>The metrics data collected for each of the cases and remapping algorithms are stored under the <code>MetricsData</code> directory. The metrics CSV files include details about:</p> <ul> <li>Error convergence data in global norms $L_1, L_2, L_{\inf}, H_1$ and $\left|H_1\right|$</li> <li>Global bounds preservation for determining monotonicity</li> <li>Local feature preservation through repeated remapping cycles</li> <li>Grid independence by using test cases with different mesh types and (uniformly refined/regionally refined) resolutions</li> </ul> </li> </ul> </li> <li> <p>A set of helpful Python scripts have also been provided to easily compare different aspects of the metrics data to gain more insight into the behavior of the remapping algorithms. These are under <code>Scripts</code> directory.</p> </li> </ol>
Can Proposed Service Interface Metrics Effectively Evaluate the Quality of RESTful APIs? Dataset and Filter Criteria
<p>The csv datasets contains repositories with API metrics and software quality metrics.</p> <p>We use linear regression to predict a software quality metric with the API metrics.</p> <p> </p> <p>The .txt files contain the lists for the filter criteria to filter out unsuitable repositories.</p>
Zalophus hindflipper turning metrics and angle of attack
<p>California sea lions (<em>Zalophus californianus</em>) are a highly maneuverable species of marine mammal. During uninterrupted, rectilinear swimming, sea lions oscillate their foreflippers to propel themselves forward without aid from the collapsed hindflippers, which are passively trailed. During maneuvers such as turning and leaping (porpoising), the hindflippers are spread into a delta-wing configuration. There is little information defining the role of otarrid hindflippers as aquatic control surfaces. To examine Z. californianus hindflippers during maneuvering, trained sea lions were video recorded underwater through viewing windows performing porpoising behaviors and banking turns. Porpoising by a trained sea lion was compared with sea lions executing the maneuver in the wild. Anatomical points of reference (ankle and hindflipper tip) were digitized from videos to analyze various performance metrics and define the use of the hindflippers. During a porpoising bout, the hindflippers were considered to generate lift when surfacing with a mean angle of attack of 14.6±6.3 deg. However, while performing banked 180 deg turns, the mean angle of attack of the hindflippers was 28.3±7.3 deg, and greater by another 8-12 deg for the maximum 20% of cases. The delta-wing morphology of the hindflippers may be advantageous at high angles of attack to prevent stalling during high-performance maneuvers. Lift generated by the delta-shaped hindflippers, in concert with their position far from the center of gravity, would make these appendages effective aquatic control surfaces for executing rapid turning maneuvers<span>.</span></p>
Elephant agricultural use metrics in Mara-Serengeti ecosystem
<p>Agricultural use metrics were calculated for 66 elephants as part of a study to characterize crop use tactics in the Mara-Serengeti ecosystem in Kenya and Tanzania. Metrics were calculated to capture mean agricultural use, maximum use from a moving average, and the difference between mean and max use. These metrics were used to classify agricultural use tactics for each elephant using Gaussian mixture models. Tables are provided with metrics and tactic classifications for the lifetime track (TableS3) and individual years (TableS4). Data contained in these files can be used to reproduce and further investigate Guassian mixture model clustering and cutpoint calculation, agricultural use linear mixed models, and tactic change generalized logistic mixed models in Hahn et al. 2021. </p>
Country AI Activity Metrics
<p>This <a href="https://eto.tech" target="_blank" rel="noopener">Emerging Technology Observatory</a> dataset includes national-level metrics for AI-related research, patents, and private-market investment. Metrics are presented for AI as a whole as well as select subfields, such as natural language processing and computer vision. For schemas and a detailed methodological description, visit the <a href="https://eto.tech/dataset-docs/country-ai-activity-metrics" target="_blank" rel="noopener">full documentation</a>. To browse the data visually, visit ETO's <a href="https://cat.eto.tech" target="_blank" rel="noopener">Country Activity Tracker</a>.</p> <p>Research subject classifications are based on work supported in part by the Alfred P. Sloan Foundation under Grant No. G-2023-22358.</p>
ETO Cross-Border Tech Research Metrics
<p>ETO's Cross-Border Tech Research Metrics dataset includes metrics for cross-border research in emerging technology domains, such as AI, robotics, and cybersecurity. For complete documentation and caveats, visit the <a href="https://eto.tech/dataset-docs/cross-border-tech-research-metrics">ETO website</a>.</p> <p>Research subject classifications are based on work supported in part by the Alfred P. Sloan Foundation under Grant No. G-2023-22358.</p>
Revisiting process versus product metrics: A large scale analysis
<p>Numerous methods can build predictive models from software data. However, what methods and conclusions should we endorse as we move from analytics in-the-small (dealing with a handful of projects) to analytics in-the-large (dealing with hundreds of projects)? To answer this question, we recheck prior small-scale results (about process versus product metrics for defect prediction and the granularity of metrics) using 722,471 commits from 700 Github projects. We find that some analytics in-the-small conclusions still hold when scaling up to analytics in-the-large. For example, like prior work, we see that process metrics are better predictors for defects than product metrics (best process/product-based learners respectively achieve recalls of 98%/44% and AUCs of 95%/54%, median values).</p> <p>That said, we warn that it is unwise to trust metric importance results from analytics in-the-small studies since those change dramatically when moving to analytics in-the-large. Also, when reasoning in-the-large about hundreds of projects, it is better to use predictions from multiple models (since single model predictions can become confused and exhibit a high variance).</p>
Spark Data containing logs and metrics (KPIs) for Hades
<p>Please make sure to cite our paper whenever you use the data in your research:</p><p>@inproceedings{DBLP:conf/icse/LeeYCSYL23, author = {Cheryl Lee and Tianyi Yang and Zhuangbin Chen and Yuxin Su and Yongqiang Yang and Michael R. Lyu}, title = {Heterogeneous Anomaly Detection for Software Systems via Semi-supervised Cross-modal Attention}, booktitle = {45th {IEEE/ACM} International Conference on Software Engineering, {ICSE} 2023, Melbourne, Australia, May 14-20, 2023}, pages = {1724--1736}, publisher = {{IEEE}}, year = {2023}, url = {https://doi.org/10.1109/ICSE48619.2023.00148}, doi = {10.1109/ICSE48619.2023.00148}, timestamp = {Wed, 19 Jul 2023 10:09:12 +0200}, biburl = {https://dblp.org/rec/conf/icse/LeeYCSYL23.bib}, bibsource = {dblp computer science bibliography, https://dblp.org} }</p><p> </p><p> </p>
Risk and Equity Metrics for the NYC Flood Risk Digital Twin
<p>These datasets contain the Risk and Equity metrics used to quantify the impact of pluvial flooding in NYC. They are obtained combining several sources (US Census data, New York State Traffic data, etc.) with the NYC Stormwater Flood Map corresponding to an extreme rain event.</p>
Bitcoin volatility in bull vs. bear market - insights from analyzing on-chain metrics and Twitter posts
<p>On-Chain Metrics.xlsx contains a description of the on-chain metrics.<br> Merged_df.xlsx is the main data source containing the BTC prices, the on-chain metrics and the sentiment scores.<br> btc_twets_new.csv and training.1600000.processed.noemoticon.csv are the data sources for calculating the sentiment scores.<br> Sentiment_Analysis.py contains the code to calculate the sentiment scores. The scores are in Merged_df.xlsx<br> BTC_Prediction.py contains the implementation of the main approach described in the paper, especially in Fig. 11.</p>
An Exploratory Factor Analysis of Code Quality Metrics
<p>Replication package for the paper "An Exploratory Factor Analysis of Code Quality Metrics"</p>
FIG. 2. Nonmetric multidimensional scaling calculated with the Bray-Curtis distance metric using a in Spatial Variation of False Map Turtle (Graptemys pseudogeographica) Bacterial Microbiota in the Lower Missouri River, United States
FIG. 2. Nonmetric multidimensional scaling calculated with the Bray-Curtis distance metric using a square root transformation and Wisconsin double-standardization. Location is represented by color, and sex is represented by shape. Stress of fit for the ordination is reported at 0.145. Axis titles represent the two dimensions to which the data have been ordinated.
V-RMS and Temperature metrics of the ball bearing units of an industrial blower
<p>Sensor devices were attached to the two mounted ball bearing units of an industrial blower and collected v-rms and temperature measurements throughout several months of operation.</p> <p>The v-rms is a vibration metric representing the average velocity of the bearings' vibrations and here is expressed in <em>inch/sec</em>, while the temperature is a self explanatory metric and is expressed in degrees of <em>Fahrenheit</em>.</p> <p>The data were collected during a normal and an encumbered operational production period and were used for the training and evaluation of an Autoencoder model for the purpose of predictive maintenance and the task of anomaly detection.</p>
Comparison Metrics Microscale Simulation Challenge for Wind Resource Assessment - Perdigão
<p>Simulation results of the "Comparison Metrics Microscale Simulation Challenge for Wind Resource Assessment". Simulations with various models and tools were performed of the Perdigão site. The wind speed profiles for nine metmast positions and the AEP values for two metmast positions are made available. At each position the wind speed profiles and AEP values are given for twelve wind direction sectors.</p>
Data for: Functional response metrics explain and predict high but differing ecological impacts of juvenile and adult lionfish
<p>Recent accumulation of evidence across taxa indicates that the ecological impacts of invasive alien species are predictable from their Functional Response (FR; e.g. the maximum feeding rate) and Functional Response Ratio (FRR; the FR attack rate/handling time ratio). Here, we experimentally derive these metrics to predict the ecological impacts of both juvenile and adult lionfish (<em>Pterois volitans</em>), one of the world's most damaging invaders, across representative and likely future prey types. Potentially prey-population destabilising Type II FRs were exhibited by both life stages of lionfish towards four prey species: <em>Artemia salina</em>, <em>Gammarus oceanicus</em>, <em>Palaemonetes varians</em> and <em>Nephrops norvegicus</em>. FR magnitudes revealed ontogenetic shifts in lionfish impacts, while lionfish FRR values were substantially higher than mean FRR values across known damaging invasive taxa. Thus, both life stages of lionfish are predicted to contribute to differing but high ecological impacts across prey communities, including commercially important species. With lionfish invasion ranges currently expanding across multiple regions globally, efforts to reduce lionfish numbers and population size structure, and provision of prey refugia through habitat complexity, might reduce their impacts. However, early detection and complete eradication of individuals located in new regions is advised.</p>
Dataset, metrics and indicators used for Horizon Results Booster analysis
<p>Input data used for the HRB analysis. Criteria and indicators to determine what can be assessed as best practice and/or innovative output in the context of the HRB analysis. The HRB proposal for harmonizing project data collection (instead of manually retrieving this information looking in documents).</p>
Metrics Datasets
<p>Metrics dataset</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.