Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
753
datasets available to search
ShareScore release 0.9.0
Dataset results
753 results for “metrics”
SLA Metrics
<p>This dataset represents the SLA metrics for the service level objectives that were identified in the SUNFISH requirements.</p>
Datasets for "Reading Order Independent Metrics for Information Extraction in Handwritten Documents"
<p>This repository includes the five datasets used for our paper entitled <em>Reading Order Independent Metrics for Information Extraction in Handwritten Documents</em>, in which we compare various metrics to evaluate end-to-end information extraction from scanned documents.</p> <h2>Datasets</h2> <p>Five datasets are released following the BIO format:</p> <ul> <li>IAM</li> <li>Simara</li> <li>POPP</li> <li>Esposalles</li> <li>French Military Records</li> </ul> <p>For each dataset, we provide the following data (on test sets):</p> <ul> <li>Ground truth annotations (<code>gt/</code>)</li> <li>Automatic predictions (<code>dan/</code>)</li> <li>Automatic predictions with entities appearing in random order (<code>dan_shuffled/</code>)</li> </ul> <p>The data is organized as follows:</p> <p><code>├── Dataset name/</code><br><code>│ ├── gt/</code><br><code>│ ├── dan/</code><br><code>│ └── dan_shuffled/</code></p> <h2>Metrics</h2> <p>To install the <a href="https://pypi.org/project/ie-eval/"><code>ie-eval</code></a> package, run <code>pip install ie-eval</code>.</p> <p>To compute all metrics on a specific dataset, run:<br><br><code>ie-eval all --label-dir IAM_paragraph/gt/ --prediction-dir IAM_paragraph/dan/</code><br><br></p> <p>To learn more about the various options, use the <code>--help</code> argument or read the <a href="https://ie-eval-ner-metrics-050f40e80b04480e2310d39ad338de778f6bec80e18.pages.teklia.com/">documentation</a>.</p> <p> </p>
Remotely sensed temperature metrics data for Ningaloo Coast, Western Australia
<p>Here we made available four datasets of remotely sensed temperature metrics (daily SST, weekly SST, SSTA frequency, TSA DHW) from different NOAA and IMOS satellite products for the Ningaloo Coast (Western Australia) bounding box with coordinates [-23.5654,-21.66538], [113.4847,114.318], which can be used to assess thermal history and trends across coral reef areas within this World Heritage Area. </p> <ul> <li><a target="_blank" rel="noopener noreferrer">ningaloo_reef_crw_sst_1985-2022.nc:</a> Daily SST (1985-01-01 - 2022-12-29), NOAA Coral Reef Watch (CRW) Version 3.1 global 5 km daily nighttime SST (aka CoralTemp) </li> <li>ningaloo_reef_cortadv6_filled_sst_1985_2022.nc: Weekly SST (1985-12-31 - 2022-12-20), NOAA Coral Temperature Anomaly Database (CoRTAD) Version 6 global 4.6 km from weekly averaged daytime and nighttime SST;</li> <li>ningaloo_reef_cortadv6_ssta_freq_1985_2022.nc: Frequency of SSTA (1985-12-31 - 2022-12-20), NOAA Coral Temperature Anomaly Database (CoRTAD) Version 6 global 4.6 km from weekly averaged daytime and nighttime SST;</li> <li>ningaloo_reef_cortadv6_tsa_dhw_1985_2022.nc: TSA DHW (1985-12-31 - 2022-12-20), NOAA Coral Temperature Anomaly Database (CoRTAD) Version 6 global 4.6 km from weekly averaged daytime and nighttime SST;</li> <li>ningaloo_reef_imos_hinawari-8_L3C_sst_1_hour_2016-2022.nc: Hourly SST (2016-01-01 - 2022-12-13) , IMOS Hinawari-8 L3C 1 km 1 hour SST from 2016 to 2022</li> </ul> <p> </p> <p> </p>
A database of Defra statutory biodiversity metric unit values for terrestrial habitat samples across England, with plant, butterfly and bird species data
<p>Policies requiring biodiversity no net loss or net gain as an outcome of environmental planning have become more prominent worldwide, catalysing interest in biodiversity offsetting as a mechanism to compensate for development impacts on nature. Offsets rely on credible and evidence-based methods to quantify biodiversity losses and gains. Following the introduction<span> of the United Kingdom's Environment Act in November 2021, all new developments requiring planning permission in England are expected to demonstrate a 10% biodiversity net gain from 2024, calculated using the statutory biodiversity metric framework (Defra, 2023). </span><span>The metric is used to calculate both baseline and proposed post-development biodiversity units, and is </span>set to play an increasingly prominent role in nature conservation nationwide.<span> </span><span>The metric has so far </span>received limited scientific scrutiny.</p> <p><span>This dataset comprises a database of statutory biodiversity metric unit values for terrestrial habitat samples across England. For each habitat sample, we present </span><span>biodiversity units alongside five long-established single-attribute proxies for biodiversity (</span><span>species richness, individual abundance, number of threatened species, mean species range or population, mean species range or population change)</span><span>. </span><span>Data were compiled </span><span>for species from three taxa (vascular plants, butterflies, birds), from sites across England. The dataset includes 24 sites within </span>grassland, wetland, woodland and forest, sparsely vegetated land, cropland, heathland and shrub, i.e. <span>all terrestrial broad habitats except urban and individual trees. Species data were reused from long-term ecological change monitoring datasets</span> (mostly in the public domain), whilst biodiversity units were calculated following field visits. Fieldwork was carried out in April-October 2022 to calculate biodiversity units for the samples. <span>Sites were initially assessed using metric version 3.1, which was current at the time of survey, and were subsequently updated to the statutory metric for analysis using field notes and species data. </span>Species data <span>were derived from </span>24 <span>long-term ecological change monitoring</span> sites across the Environmental Change Network (ECN), Long Term Monitoring Network (LTMN) and Ecological Continuity Trust (ECT), collected between 2010 and 2020.</p>
QSage: Structural and Semantic Metric Analysis for Quantum Code Smell Detection
Open the record for dataset details and reuse information.
A data set from a survey investigating the smart approach to selecting good cyber security metrics
Open the record for dataset details and reuse information.
Iowa herptile detection histories and landcover metrics
<p>Predictions of species occurrence allow land managers to focus conservation efforts on locations where species are most likely to occur. Such analyses are rare for herpetofauna compared to other taxa, despite increasing evidence that herptile populations are declining because of land cover change and habitat fragmentation. Our objective was to create predictions of occupancy and colonization probabilities for 15 herptiles of greatest conservation need in Iowa. From 2006–2014, we surveyed 295 properties throughout Iowa for herptile presence using timed visual-encounter surveys, coverboards, and aquatic traps. Data were analyzed using robust design occupancy modeling with landscape-level covariates. Occupancy ranged from 0.01 (95% CI = -0.01, 0.03) for prairie ringneck snake (<em>Diadophis punctatus arnyi</em>) to 0.90 (95% CI = 0.898, 0.904) for northern leopard frog (<em>Lithobates pipiens</em>). Occupancy for most species correlated to landscape features at the 1-km scale. General patterns of species' occupancy included the negative effects of agricultural features and the positive effects of water features on turtles and frogs. Colonization probabilities ranged from 0.007 (95% CI = 0.006, 0.008) for spiny softshell turtle (<em>Apalone spinifera</em>) to 0.82 (95% CI = 0.62, 1.0) for western fox snake (<em>Pantherophis ramspotti</em>). Colonization probabilities for most species were best explained by the effects of water and grassland landscape features. Predictive models had strong support (AUC > 0.70) for six out of 15 species (40%), including all three turtles studied. Our results provide estimates of occupancy and colonization probabilities and spatial predictions of occurrence for herptiles of greatest conservation need across the state of Iowa.</p>
Replication Package for "Early Career Developers' Perceptions of Code Understandability. A Study of Complexity Metrics"
<div> <div> <div> <div> <div> <h2>Authors</h2> <ul> <li>Matteo Esposito, University of Oulu, Finland</li> <li>Andrea Janes, <span>Free University of Bozen-Bolzano</span>, Italy</li> <li>Terhi Kilamo, University of Tampere, Finland</li> <li>Valentina Lenarduzzi, University of Oulu, Finland</li> </ul> <h2>Content Overview</h2> <p>This replication package contains the following materials:</p> <ul> <li><strong>Tables:</strong> Excel files that include all hypothesis testing data, including normality tests.</li> <li><strong>Data:</strong> RAW Questionarie datasett.</li> </ul> <h2>Contact Information</h2> <p>For any issues, questions, or further assistance, please do not hesitate to contact the authors of the paper. We are here to help!</p> </div> </div> </div> </div> </div>
24-hour movement behaviors regarding different accelerometer metrics and cardiometabolic variables of Belgian adults
<p>Datase of 213 adults with 24-hour movement behaviors features and cardiometabolic health variables</p> <p>Sociodemographic information</p> <ul> <li>age</li> <li>sex </li> <li>educational level</li> <li>smoking status</li> <li>pathology (having T2DM or not)</li> </ul> <p>Cardiometabolic variables </p> <ul> <li>BMI</li> <li>Waist circumference</li> <li>waist to hip ratio</li> <li>Fat percentage</li> <li>glucose</li> <li>HbA1c</li> <li>HDL-cholesterol</li> <li>LDL-cholesterol</li> <li>Total cholesterol</li> <li>Triglycerides</li> <li>Systolic Blood pressure</li> <li>Diastolic blood pressure</li> </ul> <p>24-hour movement behaviors (Actigraph GT3X+):</p> <ul> <li>Cut-point dependent time usse estimates for sleep, sedentary behavior, light phyiscal activity, moderate to vigorous physical activity</li> <li>Cut-point independent average acceleration</li> <li>Cut-point independent intensity gradient</li> <li>These cut-off points are available for ENMO, MAD, CMP VA (neish) and CMP VM metric (sasaki) as mentioned in the paper</li> </ul> <p>For more information please contact willems.iris@ugent.be</p>
Predicting Bug-Inducing Commits Using Software Quality Metrics
Open the record for dataset details and reuse information.
Impact of Methodological Choices on the Analysis of Code Metrics and Maintenance
<p>The repo-data folder contains 53 <code>.json</code> files, each corresponding to one of the 53 Java open-source projects. Each file contains various metrics for methods in the project.</p> <p> </p> <pre><code> { "hawtio-3976.json": { "Age": 794, "sloc": [11,11,11], "slocAsItIs": [11,14,14], "slocNoCommentPretty": [11,11,11], "diffSizes": [0,7,0 ], "bodychanges": [0,1,0], "newAdditions": [0,5,0], "isGetter": [false,false,false], "isSetter": [false,false,false], "changeDates": [0,3,794], "isEssentialChange": [false,true,false], "isBuggy": [false,false,false], "changeTypes": ["Yintroduced","Ybodychange","Yfilerename"], "filename": "hawtio-3976.json", "authors": ["X","Y","Z"], "editDistance": [0, 68, 0], "repo": "hawtio" }, "method_id": {...}, "method_id": {...} } </code></pre> <p> </p> <p>The above method with id <code>hawtio-3976.json</code> has total 3 revisions which is why the array of values for a particular metric (e.g., sloc: <code>[11,11,11]</code>) are of length 3. <code>Index 0</code> of the array represents the introduction value of a particular metric for the above method.</p> <div> <h3>Description of the metrics</h3> </div> <ul> <li><code>Age</code>: Age of the method in days</li> <li><code>sloc</code>: Source line of code of a method without comment and blank lines</li> <li><code>slocAsItIs</code>: Source line of code of a method with comment and blank lines</li> <li><code>slocNoCommentPretty</code>: Source line of code pretty printed without comment and blank lines</li> <li><code>diffSizes</code>: Total number of lines added + removed in git <code>diff</code></li> <li><code>bodychanges</code>: Contains value 0 or 1; where 1 implies occurrence of body change</li> <li><code>newAdditions</code>: Total number of lines added in git <code>diff</code></li> <li><code>isGetter</code>: Contains <code>true</code> or <code>false</code>; where <code>true</code> indicates it is a <code>get</code> method</li> <li><code>isSetter</code>: Contains <code>true</code> or <code>false</code>; where <code>true</code> indicates it is a <code>set</code> method</li> <li><code>changeDates</code>: Contains the date difference in days from when the method was introduced. Index <code>0</code> is always 0 which indicates the introduction date</li> <li><code>isEssentialChange</code>: Contains <code>true</code> or <code>false</code>; where <code>true</code> indicates it is an essential change. Essential change includes: <code>Ybodychange</code>, <code>Ymodifierchange</code>, <code>Yexceptionschange</code>, <code>Yrename</code>, <code>Yparameterchange</code>, <code>Yreturntypechange</code> and <code>Yparametermetachange</code> detected by <code>CodeShovel</code></li> <li><code>isBuggy</code>: Contains <code>true</code> or <code>false</code>; where <code>true</code> indicates the method bug was fixed at a particular revision</li> <li><code>changeTypes</code>: All transformations applied to the method at each revision. The full list of transformation that is detected by <code>CodeShovel</code> are: <code>Ybodychange</code>, <code>Ymodifierchange</code>, <code>Yexceptionschange</code>, <code>Yrename</code>, <code>Yparameterchange</code>, <code>Yreturntypechange</code>, <code>Yparametermetachange</code>, <code>Yannotationchange</code>, <code>Ydocchange</code>, <code>Yformatchange</code>, <code>Yfilerename</code> and <code>Ymovefromfile</code></li> <li><code>filename</code>: It is the method <code>id<br></code> <p><strong>bugData</strong> folder contains 53 .json files with bug information, each belonging to one of the 53 Java open-source projects. The sample JSON schema of a file is given below:</p> <pre><code>{ "hawtio-3976.json":{ "exactBug0Match": [false, false, false], "exactBug1Match": [false, false, false], "exactBug2Match": [false, false, false], "exactBug3Match": [false, false, false], "regExBug0": [false, false, false], "regExBug1": [false, false, false], "regExBug2": [false, false, false], "regExBug3": [false, false, false] }, "method_id": {...}, "method_id": {...}, }</code></pre> <p>The above method can be mapped to its metrics dataset using the method_id. For e.g., the above method with id hawtio-3976.json in bugData/hawtio.jsonthat has 3 revision can be found in the metric dataset using the same id hawtio-3976.json in the file repo-data/hawtio.json.<br>Description of bug dataset</p> <p>Each key in the above example contains value true or false indicating if a method was buggy or not at each revision. The "hawtio-3976.json method has 3 revisions (including method's introduction) which is why the array length is 3. The keys in the above json output represent bug-fix classification based on buggy keywords adopted from prior work.</p> <p>We identified bug-fix commit using two approaches:</p> <p> Exact case insensitive match of buggy keywords from the commit message (keys prefix wih exact represent this)<br> Partial case insensitive substring match (using regular expression) excluding words that ends with fix or bug. (keys prefix with regEx represent this)</p> <p> Bug0: This is the approach that we have used for classifying bug-fix commit. Buggy keyword list: <strong>["error", "bug", "fixes", "fixing", "fix", "fixed", "mistake", "incorrect", "fault", "defect", "flaw"]</strong><br> Bug1: Same keyword list as exactBug0Match with the addition of keyword issues<br> Bug2: Buggy keyword list from prior work: <strong>["bug", "fix", "error", "issue", "crash", "problem", "fail", "defect", "patch"]</strong><br> Bug3: Buggy keyword list from prior work: <strong>["error", "bug", "fix", "issue", "mistake", "incorrect", "fault", "defect", "flaw", "type"]</strong></p> </li> </ul>
Supplementary material 1 from: Ávila MP, Carvalho RN, Casatti L, Simião-Ferreira J, de Morais LF, Teresa FB (2018) Metrics derived from fish assemblages as indicators of environmental degradation in Cerrado streams. Zoologia 35: 1-8. https://doi.org/10.3897/zoologia.35.e12895
Table S1. Species identity, total number of individuals and species classification according to trophic guilds and habitat use. Terins: terrestrial invertivorous; Aquins: aquatic invertivorous; Det-Per: detritivorous/periphytivorous; Pis: piscivorous; Omni: omnivorous; WC: water column; Ben: benthic; Nectb: nectobenthic; Bank: bank-dwelling species; Rheo: rheophilic. Figure S1. Species accumulation (black) and rarefaction curve (gray) based on samples. : Data type: measurement
Text-fig. 16. Metric comparisons of upper and lower fourth premolars and third molars of Hippopotamodon erymanthius from Mahmutgazi (open squares) and Akkaşdaği (dots), Turkey. The two samples are metrically closely similar. in Hippopotamodon erymanthius (Suidae, Mammalia) from Mahmutgazi, Denizli-Çal basin, Turkey
Text-fig. 16. Metric comparisons of upper and lower fourth premolars and third molars of Hippopotamodon erymanthius from Mahmutgazi (open squares) and Akkaşdaği (dots), Turkey. The two samples are metrically closely similar.
WWBS Metrics
<p>The dataset consists of recordings acquired from a sensorized WWBS smart-vest garment (developed by the project's partner SMARTEX), which the older adults wear during their daily routine.</p> <p>The garment is equipped with electrodes for ECG monitoring, a piezoresistive sensor for respiration monitoring, and Inertial Measurement Units (IMUs) for movement and posture monitoring.</p> <p>Each recording consists of a series of channels, and each recording channel is a pair of [Timestamp, Value] columns.</p> <p>The full list of recording channels is described below:</p> <p><strong>- Part_id</strong>: User ID, a 4-digit number, serving as the file name for each user's data</p> <p>- <strong>ECG</strong>: Electric signal measuring the ECG (value: 0.8 mV - sampling rate: 250Hz)</p> <p>- <strong>ECGquality</strong>: ECG signal quality (value: 0-255 where 0=poor and 255=excellent - sampling rate: 1/5 sec)</p> <p>- <strong>ECGHR</strong>: Heart rate (value: Beats/minute - sampling rate: 1/5 sec)</p> <p>- <strong>ECGRR</strong>: R-R intervals (value: number of samples between R-R peaks - sampling rate:1/5 sec)</p> <p>- <strong>ECGHRV</strong>: Heart rate variability (value: ms - sampling rate: 1/60 sec)</p> <p>- <strong>AccX-Y-Z</strong>: Accelerometer in X-Y-Z axes (value: 0.97 10-3 g - sampling rate: 25 Hz)</p> <p>- <strong>GyroX-Y-Z</strong>: Gyroscope in X-Y-Z axes (value: 0.122 °/s - sampling rate: 25 Hz)</p> <p>- <strong>MagX-Y-Z</strong>: Magnetometer in X-Y-Z axes (value: 0.6 µT - sampling rate: 25 Hz)</p> <p>- <strong>RespPiezo</strong>: Electric signal measuring the chest pressure on the piezoelectric point (value: 0.8 mV - sampling rate: 25 Hz)</p> <p>- <strong>RespQuality</strong>: Respiration signal quality (value: 0-255 where 0=poor and 255=excellent - sampling rate: 1/5 sec)</p> <p>- <strong>BR</strong>: Breathing rate (value: Breaths/minute - sampling rate: 1/5 sec)</p> <p>- <strong>BA</strong>: Breathing Amplitude (value: logic levels - sampling rate: 1/15 sec)</p> <p>- <strong>Activityenergy</strong>: estimation of energy activity (value: estimation where 0=no activity and 255=max of activity - sampling rate: 1/5 sec)</p> <p>- <strong>Activityclass</strong>: Activity performed (value: 0-4 where 0=other, 1=lying, 2=standing/sitting, 3=walking and 4=running - sampling rate: 1/5 sec)</p> <p>- <strong>Activity1Pace</strong>: Step period (value: ms- sampling rate: 1 Hz)</p> <p>- <strong>ActivityPace</strong>: Pace (value: steps/min - sampling rate: 1/5 sec)</p> <p>- <strong>Q0-Q1-Q2-Q3</strong>: Quaternions from main electronic device, i.e. Q0, Q1, Q2, Q3 components (value: Q14 format - sampling rate: 25 Hz)</p> <p>- <strong>QEL0-QEL1-QEL2-QEL3</strong>: Quaternions from external left arm device, i.e. Q0, Q1, Q2, Q3 components (value: Q14 format - sampling rate: 25 Hz)</p> <p>- <strong>QER0-QER1-QER2-QER3</strong>: Quaternions from external right arm device, i.e. Q0, Q1, Q2, Q3 components (value: Q14 format - sampling rate: 25 Hz)</p>
*Metrics Survey on Researchers' Usage Purposes for Social Media B
<p>Response data from part B of the second online survey of the *metrics-project. Main goal of the survey was to determine researchers' motivations for using various social media platforms that are potential sources for altmetrics, as well as inquiring about their perceptions of various metrics for research evaluation. During dissemination, a strong focus lay on researchers from social sciences and economics. </p> <p>The questionnaire had originally been implemented in LimeSurvey. Dissemination started on 25 June 2018; the survey was finally taken offline on 13 February 2019. The call for participation was disseminated via a combination of direct non-personalized mails and mailing lists. A total of ~27,000 recipients has been targeted this way. </p>
*Metrics Survey on Researchers' Usage Purposes for Social Media A
<p>Response data from part A of the second online survey of the *metrics-project. Main goal of the survey was to determine researchers' motivations for using various social media platforms that are potential sources for altmetrics, as well as inquiring which target groups researchers usually plan to reach by being active on these platforms. During dissemination, a strong focus lay on researchers from social sciences and economics. </p> <p>The questionnaire had originally been implemented in LimeSurvey. Dissemination started on 25 June 2018; the survey was finally taken offline on 13 February 2019. The call for participation was disseminated via a combination of direct non-personalized mails and mailing lists (see "Coping with Altmetrics' Heterogeneity - A Survey on Social Media Platforms' Usage Purposes and Target Groups for Researchers" by Lemke & Peters [2019] for further details). A total of ~27,000 recipients has been targeted this way. </p>
*Metrics Survey on Usage of Social Media Services in Science
<p>Response data from the first online survey of the *metrics-project. Aim of the survey was to determine the status quo of online platform usage by researchers, with a strong focus lying on researchers from social sciences and economics. </p> <p>The questionnaire had originally been implemented in LimeSurvey. Dissemination started on 31 March 2017; the survey was finally taken offline on 13 February 2019. The call for participation was disseminated via a combination of direct personalized mails (containing recipients' first and surnames), direct non-personalized mails and mailing lists (see "Are There Different Types of Online Research Impact?" by Lemke, Mehrazar, Mazarakis, & Peters [2018] for further details). A total of ~54,000 recipients has been targeted this way. </p>
Metrics: BOLDS data coverage
Includes species-level taxa.
Metrics: Data Hubs data coverage - family level
NCBI, GGBN, BHL, BOLDS For family-level taxa.
Metrics: Data Hubs data coverage - genus level
NCBI, GGBN, BHL, BOLDS For genus-level taxa.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.