Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

753

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

753 results for “metrics”

Learn how ShareScore rates datasets ↗
zenodo36/100

Alternative metrics and social impact of research about Social Sciences in Cuba

<p>The evaluation of social impact of research is a subject demanded by the scientific and social community. The present research is developed with the objective of describing the social impact of the results of scientific research in the field of Social Sciences in Cuba. 5 dimensions of analysis and 16 alternative indicators were used, through the use of altmetric tools and data sources. The sample collection for the study was carried out through the Scopus database and the altmetric data provider PlumX Metrics. For the analysis, statistical techniques of trend and correlation between indicators, data visualization and scientific information were used. The results show that the indicators with the greatest presence were citations in Scopus and CrossRef, Views count, Full Text Views, Abstract Views, Readers in Mendeley captures and the social network metrics Facebook and Twitter. The research results with the greatest social impact are related to climate change and environmental policy, scientific production about COVID-19, higher education, sustainable development, gender studies, legislation, and tourism.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Supplementary material for 'The MAP metric in Information Retrieval Fault Localization'

<pre># map_bench4bl This is the supplementary material, data, and evaluation source code for the paper &quot;The MAP metric in Information Retrieval Fault Localization&quot; by Thomas Hirsch and Birgit Hofer. ## Preliminaries ### Python environment - Python 3.8 - pandas - numpy - matplotlib ## Datasets The [Bench4BL](<em>https://github.com/exatoa/Bench4BL</em>) dataset has been used in this evaluation, with the addition of intermediate files taken from the [SABL](<em>http://dx.doi.org/10.5281/zenodo.4681242</em>) experiment performed on this Bench4BL dataset. All data used in our evaluation is included in this repository. However, if the data is to be re-imported directly from these benchmark and datasets they have to be downloaded first and their local paths have to be set in [paths.py](<em>paths.py</em>). ### Bench4BL The Bench4BL dataset was published with the paper &quot;Bench4BL: Reproducibility study on the performance of IR-based bug localization&quot; by Lee, J., Kim, D., Bissyand&eacute;, T.F., Jung, W. and Le Traon, Y.. The dataset can be obtained [here](<em>https://github.com/exatoa/Bench4BL</em>). Follow the steps described in the corresponding [README](<em>https://github.com/exatoa/Bench4BL/blob/master/README.md</em>) to set up the dataset. The Bench4BL dataset contains the _old subjects_ subdataset, containing 558 bugs from AspectJ, JDT, PDE, SWT, and ZXing that have been widely used in older IRFL studies. This _old subjects_ subdataset was used in answering our RQ1, as discussed below, the corresponding scripts use _old subjects_ in their name to highlight this. #### SABL The SABL dataset is the online appendix of the paper &quot;An Extensive Study of Smell-Aware Bug Localization&quot; by TTakahashi, A., Sae-Lim, N., Hayashi, S. and Saeki, M.. The dataset can be downloaded [here](<em>http://dx.doi.org/10.5281/zenodo.4681242</em>). The experiments in this dataset build on top of Bench4BL and intermediate files are provided in the datapackage. #### Rankings Rankings for BLIA, BRTracer, and BugLocator were produced by running these tools on Bench4BL locally. Rankings for AmaLgam and BLUiR were taken from the SABL experiment dataset. ## Structure ### Folders Bench4BL ground truths: - bench4bl_old_subjects_summary - bench4bl_summary Localization results of the included tools in Bench4BL: - bench4bl_localization_results - bench4bl_localization_results_sabl Target projects size metrics: - cloc_results - cloc_results_old_subjects Utility functions: - utils Output folders containing results, generated figures and tables: - results - results_old_subjects ### Scripts Scripts for re-importing data from Bench4BL and SABL datasets: - data_preparation_step_1_cloc_bench4bl.py - data_preparation_step_1_cloc_old_subjects_bench4bl.py - data_preparation_step_2_import_ground_truth_from_bench4bl.py - data_preparation_step_2_import_ground_truth_from_old_subjects_bench4bl.py - data_preparation_step_3_import_bench4bl_ranking_results.py - data_preparation_step_3_import_sabl_ranking_results.py Utilities: - paths.py - utils/bench4bl_utils.py - utils/Logger.py ### Evaluation scripts for the corresponding research questions: **Dataset analysis:** - rq_0_dataset_analysis_bench4bl_issues.py **RQ1: How big is the average ground truth in Bench4BL datasets, and what proportion of bugs have a ground truth containing multiple files?** - rq_1_bench4bl_ground_truth_size.py - rq_1_old_subjects_bench4bl_ground_truth_size.py RQ2: Do the IRFL tools included in Bench4BL truncate their results? - rq_2_ranking_lengths.py **RQ3: How strong is $AP_{asrd}$ overestimating $AP_{mb}$ for truncated BugLocator retrieval results on the Bench4BL dataset? RQ3a: How strong is $AP_{asrd}$ overestimating $AP_{mb}$ for truncated BugLocator retrieval results when considering the bloated ground truth issue found in Bench4BL?** - rq_3_truncating_BugLocator_rankings_bench4bl.py **RQ3b: How strong is $AP_{asrd}$ overestimating $AP_{mb}$ for truncated BugLocator retrieval results when undefined $AP$ values are simply ignored?** - rq_3b_undefined_ap_BugLocator_rankings_bench4bl.py ## Licence All code and results are licensed under [CCA v4](<em>https://creativecommons.org/licenses/by/4.0/</em>), according to LICENSE file. Other licences may apply for some tools and datasets contained in this repo: [cloc-1.92.pl](<em>https://github.com/AlDanial/cloc</em>) under GPL v2, [Bench4BL](<em>https://github.com/exatoa/Bench4BL</em>) and [SABL](<em>http://dx.doi.org/10.5281/zenodo.4681242</em>) under CCA 4.0.</pre>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data for 'Assessing the value of biodiversity-specific footprinting metrics linked to South American soy trade'

<p>Underlying data for publication &#39;Assessing the value of biodiversity-specific footprinting metrics linked to South American soy trade&#39;.&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Satellite remote sensing dataset of Sentinel-2 for phenology metrics extraction from sites in Bulgaria and France

<p><strong>Site Description:</strong></p> <p>In this dataset, there are seventeen production crop fields in Bulgaria where winter rapeseed and wheat were grown and two research fields in France where winter wheat &ndash; rapeseed &ndash; barley &ndash; sunflower and winter wheat &ndash; irrigated maize crop rotation is used. The full description of those fields is in the database &quot;In-situ crop phenology dataset from sites in Bulgaria and France&quot; (doi.org/10.5281/zenodo.7875440).</p> <p>&nbsp;</p> <p><strong>Methodology and Data Description:</strong></p> <p>Remote sensing data is extracted from Sentinel-2 tiles 35TNJ for Bulgarian sites and 31TCJ for French sites on the day of the overpass since September 2015 for Sentinel-2 derived vegetation indices and since October 2016 for HR-VPP products. To suppress spectral mixing effects at the parcel boundaries, as highlighted by Meier et al., 2020, the values from all datasets were subgrouped per field and then aggregated to a single median value for further analysis.</p> <p>Sentinel-2 data was downloaded for all test sites from CREODIAS (https://creodias.eu/) in&nbsp;L2A processing level using a maximum scene-wide cloudy cover threshold of 75%. Scenes before 2017 were available in L1C processing level only. Scenes in L1C processing level were corrected for atmospheric effects after downloading using Sen2Cor (v2.9) with default settings. This was the same version used for the L2A scenes obtained intermediately&nbsp;from CREODIAS.&nbsp;</p> <p>Next, the data was extracted from the Sentinel-2 scenes for each field parcel where only SCL classes 4 (vegetation) and 5 (bare soil) pixels were kept. We resampled the 20m band B8A to match the spatial resolution of the green and red band (10m) using nearest neighbor interpolation. The entire image processing chain was carried out using the open-source Python Earth Observation Data Analysis Library (EOdal) (Graf et al., 2022).</p> <p>Apart from the widely used Normalized Difference Vegetation Index (NDVI) and Enhanced Vegetation Index (EVI), we included two recently proposed indices that were reported to have a higher correlation with photosynthesis and drought response of vegetation: These were the Near-Infrared Reflection of Vegetation (NIRv) (Badgley et al., 2017)&nbsp; and Kernel NDVI (kNDVI) (Camps-Valls et al., 2021). We calculated the vegetation indices in two different ways:&nbsp;</p> <p>First, we used <strong>B08</strong> as&nbsp;near-infrared (NIR) band which comes in a native spatial resolution of 10 m. <strong>B08</strong> (central wavelength 833 nm) has a relatively coarse spectral resolution with a bandwidth of 106 nm.</p> <p>Second, we used <strong>B8A</strong> which is available at 20 m spatial resolution. <strong>B8A</strong> differs from B08 in its central wavelength (864 nm) and has a narrower bandwidth (21 nm or 22 nm in the case of Sentinel-2A and 2B, respectively) compared to B08.</p> <p>&nbsp;</p> <p>The High Resolution Vegetation Phenology and Productivity (<strong>HR-VPP</strong>) dataset from Copernicus Land Monitoring Service (CLMS) has three 10-m set products of Sentinel-2: vegetation indices, vegetation phenology and productivity parameters and seasonal trajectories (Tian et al., 2021). Both vegetation indices, Normalized Vegetation Index (NDVI) and Plant Phenology (PPI) and plant parameters, Fraction of Absorbed Photosynthetic Active Radiation (FAPAR) and Leaf Area Index (LAI) were computed for the time of Sentinel-2 overpass by the data provider.&nbsp;</p> <p>NDVI is computed directly from B04 and B08 and PPI is computed using Difference Vegetation Index (DVI = B08 - B04) and its seasonal maximum value per pixel. FAPAR and LAI are retrieved from B03 and B04 and B08 with neural network training on PROSAIL model simulations. The dataset has a quality flag product (QFLAG2) which is a 16-bit that extends the scene classification band (SCL) of the Sentinel-2 Level-2 products. A &ldquo;medium&rdquo; filter was used to mask out QFLAG2 values from 2 to 1022, leaving land pixels (bit 1) within or outside cloud proximity (bits 11 and 13) or cloud shadow proximity (bits 12 and 14).&nbsp;</p> <p>The <strong>HR-VPP</strong> daily raw vegetation indices products are described in detail in the user manual (Smets et al., 2022) and the computations details of PPI are given by Jin and Eklundh (2014).&nbsp;Seasonal trajectories refer to the 10-daily smoothed time-series of PPI used for vegetation phenology and productivity parameters retrieval with TIMESAT (J&ouml;nsson and Eklundh 2002, 2004).</p> <p>HR-VPP data was downloaded through the WEkEO Copernicus Data and Information Access Services (DIAS) system with a Python 3.8.10 harmonized data access (HDA) API 0.2.1. Zonal statistics [&rsquo;min&rsquo;, &rsquo;max&rsquo;, &rsquo;mean&rsquo;, &rsquo;median&rsquo;, &rsquo;count&rsquo;, &rsquo;std&rsquo;, &rsquo;majority&rsquo;] were computed on non-masked pixel values within field boundaries with rasterstats Python package 0.17.00.</p> <p>&nbsp;</p> <p>The Start of season date (SOSD), end of season date (EOSD) and length of seasons (LENGTH) were extracted from the annual Vegetation Phenology and Productivity Parameters (<strong>VPP</strong>) dataset as an additional source for comparison. These data are a product of the Vegetation Phenology and Productivity Parameters, see (https://land.copernicus.eu/pan-european/biophysical-parameters/high-resolution-vegetation-phenology-and-productivity/vegetation-phenology-and-productivity) for detailed information.</p> <p>&nbsp;</p> <p><strong>File Description:</strong></p> <p>4 datasets:</p> <p>1_senseco_data_S2_B08_Bulgaria_France; 1_senseco_data_S2_B8A_Bulgaria_France; 1_senseco_data_HR_VPP_Bulgaria_France; 1_senseco_data_phenology_VPP_Bulgaria_France</p> <p>3 metadata:</p> <p>2_senseco_metadata_S2_B08_B8A_Bulgaria_France; 2_senseco_metadata_HR_VPP_Bulgaria_France; 2_senseco_metadata_phenology_VPP_Bulgaria_France</p> <p>&nbsp;</p> <p>The dataset files&nbsp;&ldquo;1_senseco_data_S2_B8_Bulgaria_France&rdquo; and &ldquo;1_senseco_data_S2_B8A_Bulgaria_France&rdquo; concerns all vegetation indices (EVI, NDVI, kNDVI, NIRv) data values and related information, and metadata file &ldquo;2_senseco_metadata_S2_B08_B8A_Bulgaria_France&rdquo; describes all the existing variables. Both&nbsp;&ldquo;1_senseco_data_S2_B8_Bulgaria_France&rdquo; and &ldquo;1_senseco_data_S2_B8A_Bulgaria_France&rdquo; have the same column variable names and for that reason, they share the same metadata file&nbsp;&ldquo;2_senseco_metadata_S2_B08_B8A_Bulgaria_France&rdquo;.</p> <p>The dataset file &ldquo;1_senseco_data_HR_VPP_Bulgaria_France&rdquo; concerns vegetation indices (NDVI, PPI) and plant parameters (LAI, FAPAR) data values and related information, and metadata file &ldquo;2_senseco_metadata_HRVPP_Bulgaria_France&rdquo; describes all the existing variables.&nbsp;</p> <p>The dataset file &ldquo;1_senseco_data_phenology_VPP_Bulgaria_France&rdquo; concerns the vegetation phenology and productivity parameters (LENGTH, SOSD, EOSD)&nbsp;values and related information, and metadata file &ldquo;2_senseco_metadata_VPP_Bulgaria_France&rdquo; describes all the existing variables.</p> <p>&nbsp;</p> <p><strong>Bibliography</strong></p> <p>G. Badgley, C.B. Field, J.A. Berry, Canopy near-infrared reflectance and terrestrial photosynthesis, Sci. Adv. 3 (2017) e1602244. https://doi.org/10.1126/sciadv.1602244.</p> <p>G. Camps-Valls, M. Campos-Taberner, &Aacute;. Moreno-Mart&iacute;nez, S. Walther, G. Duveiller, A. Cescatti, M.D. Mahecha, J. Mu&ntilde;oz-Mar&iacute;, F.J. Garc&iacute;a-Haro, L. Guanter, M. Jung, J.A. Gamon, M. Reichstein, S.W. Running, A unified vegetation index for quantifying the terrestrial biosphere, Sci. Adv. 7 (2021) eabc7447. https://doi.org/10.1126/sciadv.abc7447.</p> <p>L.V. Graf, G. Perich, H. Aasen, EOdal: An open-source Python package for large-scale agroecological research using Earth Observation and gridded environmental data, Comput. Electron. Agric. 203 (2022) 107487. https://doi.org/10.1016/j.compag.2022.107487.</p> <p>H. Jin, L. Eklundh, A physically based vegetation index for improved monitoring of plant phenology, Remote Sens. Environ. 152 (2014) 512&ndash;525. https://doi.org/10.1016/j.rse.2014.07.010.</p> <p>P. Jonsson, L. Eklundh, Seasonality extraction by function fitting to time-series of satellite sensor data, IEEE Trans. Geosci. Remote Sens. 40 (2002) 1824&ndash;1832. https://doi.org/10.1109/TGRS.2002.802519.</p> <p>P. J&ouml;nsson, L. Eklundh, TIMESAT&mdash;a program for analyzing time-series of satellite sensor data, Comput. Geosci. 30 (2004) 833&ndash;845. https://doi.org/10.1016/j.cageo.2004.05.006.</p> <p>J. Meier, W. Mauser, T. Hank, H. Bach, Assessments on the impact of high-resolution-sensor pixel sizes for common agricultural policy and smart farming services in European regions, Comput. Electron. Agric. 169 (2020) 105205. https://doi.org/10.1016/j.compag.2019.105205.</p> <p>B. Smets, Z. Cai, L. Eklund, F. Tian, K. Bonte, R. Van Hoost, R. Van De Kerchove, S. Adriaensen, B. De Roo, T. Jacobs, F. Camacho, J. S&aacute;nchez-Zapero, S. Else, H. Scheifinger, K. Hufkens, P. J&ouml;nsson, HR-VPP Product User Manual Vegetation Indices, 2022.</p> <p>F. Tian, Z. Cai, H. Jin, K. Hufkens, H. Scheifinger, T. Tagesson, B. Smets, R. Van Hoolst, K. Bonte, E. Ivits, X. Tong, J. Ard&ouml;, L. Eklundh, Calibrating vegetation phenology from Sentinel-2 using eddy covariance, PhenoCam, and PEP725 networks across Europe, Remote Sens. Environ. 260 (2021) 112456. https://doi.org/10.1016/j.rse.2021.112456.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Journal metrics as predictors of Research Excellence Framework 2021 results: Comparison of impact factor quartiles and Finnish expert-ratings - dataset

<p>This dataset accompanies the conference submission &#39;Journal metrics as predictors of Research Excellence Framework 2021 results: Comparison of impact factor quartiles and Finnish expert-ratings&#39;. It contains the data in CSV format, one file per unit of analysis (Units of Assessment, Higher Education Institutions, and Subject Areas).</p> <p>The format of the UoA file is as follows (the other two files are analogous):</p> <ul> <li> <p>institution_name: name of higher education institution (e.g. university)</p> </li> <li> <p>unit_of_assessment_name: UoA name in REF (https://www.ref.ac.uk/panels/units-of-assessment/)</p> </li> <li> <p>main_panel: main panel in REF</p> </li> <li> <p>multiple_submission_letter: blank unless submitted to multiple panels.Exceptionally HEIs may have requested permission to make <a href="https://ref.ac.uk/publications-and-reports/invitation-to-make-requests-for-multiple-submissions-exception-from-submission-for-small-units-and-for-impact-case-studies-requiring-security-clearance/">two submissions from the same UoA to different panels</a>.</p> </li> <li> <p>multiple_submission_name: blank unless submitted to multiple panels</p> </li> <li> <p>non_english: number of articles in language other than English</p> </li> <li> <p>jufo_score_uoa: JUFO score (see paper for calculation details) of UoA</p> </li> <li> <p>jif_score_uoa: JIF score (see paper for calculation details) of UoA</p> </li> <li> <p>ref_score_uoa: REF score (see paper for calculation details) of UoA</p> </li> </ul>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data from: Metrics for quantifying the contributions of different threats to Red Lists

<p>The dataset contains data from four Norwegian Red Lists. Data included are the Red List Categories, reasons for change, and threats. These data were used to evaluate metrics for quantifying the contributions of different threats to Red Lists, described by <a href="https://doi.org/10.1111/cobi.14105">Sandvik &amp; Pedersen (2023)</a>.</p> <p>The dataset contains six files:</p> <ol> <li>species.csv (semicolon-delimited plain-text file with Red Lists for species)</li> <li>Species.pdf (explanations of species.csv)</li> <li>Species.xlsx (microsoft excel spreadsheet workbook with Red Lists for species)</li> <li>ecosyst.csv (semicolon-delimited plain-text file with the Red List for ecosystems)</li> <li>Ecosyst.pdf (explanations of ecosyst.csv)</li> <li>Ecosyst.xlsx (microsoft excel spreadsheet workbook with the Red List for ecosystems)</li> </ol> <p>The excel workbooks contain the same information as the respective csv and pdf files combined.</p> <p>Columns, abbreviations etc. are explained in the excel and pdf files.</p> <p>Data were derived from the following sources, all published by the <a href="http://www.biodiversity.no">Norwegian Biodiversity Information Centre</a>:</p> <ul> <li>&nbsp; <a href="http://www.artsportalen.artsdatabanken.no/">2010 Norwegian Red List for species</a></li> <li>&nbsp; <a href="https://www.artsdatabanken.no/Rodlista2015">2015 Norwegian Red List for species</a></li> <li>&nbsp; <a href="https://artsdatabanken.no/rodlistefornaturtyper">2018 Norwegian Red List for ecosystems and habitat types</a></li> <li>&nbsp; <a href="https://artsdatabanken.no/lister/rodlisteforarter/2021">2021 Norwegian Red List for species</a></li> </ul> <p><strong>R</strong> code to analyse the dataset and reproduce the results of the paper is available on Zenodo via <a href="https://doi.org/10.5281/zenodo.7843806">doi:10.5281/zenodo.7843806</a>.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

A metadata-based approach for research discipline prediction using machine learning techniques and distance metrics

<p>The dataset is based on&nbsp;the paper:&nbsp;</p> <p>Hoang-Son Pham, Hanne Poelmans&nbsp;and Amr Ali-Eldin &lsquo;&rsquo;A metadata-based approach for research discipline prediction using machine learning techniques and distance metrics&rsquo;&rsquo;, IEEE Access (2023).</p> <p>The dataset includes:&nbsp;</p> <p>1. a list of project metadata extracted from FRIS portal</p> <p>2. a list of VODS disciplines</p> <p>3. a distance matrix</p> <p>&nbsp;</p> <p>* Kindly refer to our paper for more details on the dataset.</p> <p>https://ieeexplore.ieee.org/document/10156853</p>

opencc-by-4.0May 2023View details →
dryad36/100

The Mammal Dental Metrics Database: a compilation of fossil and extant mammal tooth measurements

<p class="MsoNormal"><span>Fossil relative abundance data in combination with specimen-level measurements can reveal community-level changes underpinning long-term ecological and evolutionary dynamics. However relative abundance data is not consistently reported in the paleontological literature, and measurement data is presented in individual studies in varying formats. Here, we compiled dental measurements (tooth crown length, width, and height) from the paleontological literature (or otherwise available sources) in a single database, simultaneously permitting analyses of size and relative abundance for the fossil record. This first version of the database focuses on large mammals and on the African and Arabian Neogene fossil record. Measurements of teeth of extant species were also collected. Each row gives the available measurements for a single tooth. Specimens with multiple teeth reported are divided among several rows. Taxonomic information was entered largely as given in the original source, without standardization or revision of taxonomic identifications. Geographic, stratigraphic, and assemblage information was also collected for fossil specimens, and specimens were assigned to individual 'computational localities' (normally combinations of geological unit and location) that can be treated as paleo-communities (metacommunities). Numeric ages and geographic and geological information for these localities are given in a corresponding spreadsheet. Sources are given for all measurement and age data. This dataset provides a powerful new resource for investigations of changes in relative abundance, body mass, or mass-abundance distributions in the fossil record.</span></p>

opencc-zeroMay 2023View details →
zenodo36/100

Replication Package for the paper "AI-based Fault-proneness Metrics for Source Code Changes"

<p>This is the replication package for the paper &quot;<em>AI-based Fault-proneness Metrics for Source Code Changes</em>&quot;, submitted at the <em>IWSM-Mensura &#39;23 </em>conference.</p> <p>The archive is a <em>Docker&nbsp;</em>image file with a fully setup and working environment to re-execute the experiments involved in the manuscript. We pre-loaded all libraries and codeBERT models to ease the replication process and avoid compatibility issues, as the environment cannot be easily managed using <em>Dockerfile</em>s.</p> <p>To run the image, a <em>Docker</em>&nbsp;installation is needed. Once downloaded, from the command line type:</p> <pre><code>docker load -i &lt;/path/to/downloaded/ai-proneness-replication.tar&gt;</code></pre> <p>After the loading process, you can run the container by typing:</p> <pre><code>docker run -it mensura/ai-proneness-replication:1.0</code></pre> <p>All the source code and the dataset to re-execute the experiment is located into the&nbsp;<em>/Replication</em>&nbsp;folder. The folder contains the results of our experimentation in CSV and MS Excel format, along with the following subdirectories:</p> <ul> <li><em>dataset</em>: a replication of the used dataset. The file&nbsp;<em>dataset.csv</em>&nbsp;gives information on all the entries, while the&nbsp;<em>code </em>folder contains a subdirectory for each sample, named by its id. In the folder, the file <em>old.txt&nbsp;</em>and<em>&nbsp;</em><em>new.txt&nbsp;</em>refers to the older and newer version of the method, respectively;&nbsp;<em>gitdiff.txt </em>stores the raw <em>git-diff</em>&nbsp;command output, while&nbsp;<em>diff.html</em>&nbsp;stores a more human-readable version of the differences.</li> <li><em>ai-fault-proneness-tk-replication</em>: the Java code used to apply Tree Kernel techniques on the dataset (we used JDK-11, embedded within the container). To build and execute the package, refer to the file&nbsp;<em>README.md</em>&nbsp;in the folder. For convenience, we also provided an executable&nbsp;JAR file&nbsp;<em>ai-fault-proneness-tk-replication-1.0-jar-with-dependencies.jar </em>that can be run directly and saves the output in a CSV file in the&nbsp;<em>results</em>&nbsp;folder of the replication package.</li> <li><em>code-embeddings-and-analysis</em>: python scripts to execute the <em>codeBERT</em>-based approaches and to extract the&nbsp;<em>diff</em>&nbsp;statistics. To execute all the steps, a convenience shell script&nbsp;<em>execute.sh</em>&nbsp;has been pre-loaded and can be executed to automatize all the process.</li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Associated data underlying the article "Assessment of the Croatian Open Data Portal Using User-Oriented Metrics"

<p>Associated data underlying the article&nbsp;</p> <p>Miletić, A.; Kuveždić Divjak, A.; Welle Donker, F. Assessment of the Croatian Open Data Portal Using User-Oriented Metrics.&nbsp;<em>ISPRS Int. J. Geo-Inf.</em>&nbsp;<strong>2023</strong>,&nbsp;<em>12</em>, 185. <a href="https://doi.org/10.3390/ijgi12050185">https://doi.org/10.3390/ijgi12050185</a></p> <p>Article Abstract:<br> Open data portals are web services that serve as a central access point for all government-published open data and can exist at local, regional, national, and international levels. They are an important element of most open data initiatives that have enabled a large amount of government data to be widely available. However, data quantity and quality are not the only aspects that should be considered when publishing data. To improve the reusability of data and to achieve greater impact and benefits from open data, it is important to consider user-oriented aspects of the portal management, discovery, and use of data (e.g., organizing the portal in a user-centric way, providing accurate metadata, using a standardized and open data format, etc.). In this paper, we adopted the metrics proposed by the European Commission to assess compliance of the Croatian Open Data Portal with 10 user-oriented principles that open data portals should implement in terms of sustainability and added value. While the results show the government&rsquo;s efforts in publishing data, some aspects such as better collaboration with data providers and other data portals, offering different visualization tools, etc. need to be improved to achieve active use and impact.</p> <p>Keywords: open data; open data portal; assessment; user experience; data reuse</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Dataset for: Novel Physics Informed-Neural Networks for Estimation of Hydraulic Conductivity of Green Infrastructure as a Performance Metric by Solving Richards-Richardson PDE

<p><strong>Based on the Github respostitory:&nbsp;<a href="https://github.com/Khadrawi/Physics-Informed-Neural-Networks-for-Estimation-of-Hydraulic-Conductivity/tree/main">https://github.com/Khadrawi/Physics-Informed-Neural-Networks-for-Estimation-of-Hydraulic-Conductivity/tree/main</a></strong></p> <p>This repository contains the data used for the paper &quot;Novel Physics Informed-Neural Networks for Estimation of Hydraulic Conductivity of Green Infrastructure as a Performance Metric by Solving Richards-Richardson PDE&quot;<br> You&#39;ll find the csv files for the three simulated (Hydrus 1D) scenarios explained in the paper.&nbsp;These files were processed from the &#39;Nod_Inf.out&#39; files to csv format.</p> <p><strong>Acknowledgments</strong><br> The publicly available data used for this study (scenario 1 &amp; 2) as well as the code for the second PINN architecture (based on Dr. Maziar Raissi PINN code) and the code used to transform &ldquo;Nod_inf.out&rdquo; files from Hydrus 1D to csv files created by Dr. Toshiyuki Bandai and Dr. Teamrat A. Ghezzehei were helpfulfor this study.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

An Interpretable Framework to Characterize Compound Treatments on Filamentous Fungi using Cell Painting and Deep Metric Learning

<p>This deposit contains:</p> <ul> <li>{train,test,val}.csv:&nbsp; &nbsp; Meta-data files</li> <li>images.tar.gz:&nbsp; &nbsp; &nbsp;Archive of images</li> <li>checkpoint.pth.tar:&nbsp; &nbsp; &nbsp;Model weights</li> <li>code.tar.gz:&nbsp; &nbsp; Code necessary to reproduce our experiments.</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Dataset: Evolution Of Computational Ontologies: Assessing Development Processes Using Metrics

<p>Ontologies facilitate meaning between human and computational actors. On the one hand, the underlying technology can be considered mature. It has a standardized language, established tools for editing and sharing, and broad adoption in practice and research. On the other hand, we still know little about how these artifacts evolve over their lifetime, even though knowledge of the development process could influence quality control. It would enable us to give knowledge engineers better modeling or selection guidelines.</p> <p>This paper examines the evolution of computational ontologies using ontology metrics. First, we gathered hypotheses on the ontology development process. We assume that groups of ontologies follow a similar development pattern and that a stereotypical development process exists. Afterward, these hypotheses are tested against historical metric data from 7053 versions from 69 dormant ontologies.</p> <p>We will show that ontology development processes are highly heterogeneous. While the made hypotheses are partly true for a slight majority of ontologies, concluding the bigger picture of ontology development down to the individual ontologies is mostly not possible. Further, the data revealed that most ontologies have disruptive change events for most of the measures attributes. These disruptive events are further examined regarding their occurrences, combinations, and sizes.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Data from: Evaluating predator control using two non-invasive population metrics: a camera trap activity index and density estimation from scat genotyping

<p>Includes datasets from the Wimmera and Mallee, Victoria, Australia:</p> <p>- Fox camera trap data used to model activity</p> <p>- Fox scat SECR capture and trap files used to model density</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

2025 Competition on Electric Energy Consumption Forecast Adopting Multi-criteria Performance Metrics

<p>&nbsp;</p> <p>This dataset is the second release of data for the <a href="https://www.gecad.isep.ipp.pt/ERM-competitions/2025-energy-forecast/">2025 Competition on Electric Energy Consumption Forecast Adopting Multi-criteria Performance Metrics</a></p> <p>The competition is open and welcomes everyone who wishes to participate and to anyone who can benefit from these data.</p> <h2>Competition Outline</h2> <p>Forecasting of electric energy consumption can be a very difficult tasks when handling building-level data. However, an accurate forecast is needed to boost the potential of energy management systems. The need to forecast energy consumption grows as our reliance on renewable energy sources, such as solar and wind power, grows. This means that to meet consumer demand with renewable energy generation, energy management systems must operate based on accurate energy forecasting models for both short and long-term periods. Energy consumption forecasting techniques that can manage a variety of scenarios, including varying prediction timeframes, accessible data, data frequency, and even data quality, have been the subject of intense research. There is no one-size-fits-all approach, where certain situations call for different approaches. The goal of this competition is to compile and evaluate the most recent advances in energy consumption forecasting techniques.</p> <p>&nbsp;</p> <h2>Releases Details</h2> <ul> <li><strong>v1.0</strong>: one year of data from a smart building with readings taken every 5 minutes.</li> <li><strong>v2.0</strong>: 40 days of data from a smart building with readings taken every 5 minutes.</li> <li><strong>v3.x</strong>: a single day of data from a smart building with readings taken every hour. These releases will become available during the first competition period (from 06/01/2025 to 10/01/2025).</li> <li><strong>v4.x</strong>: a single day of data from a smart building with readings taken every hour. These releases will become available during the second competition period (from 14/07/2025 to 18/07/2025).</li> </ul> <p>&nbsp;</p> <h2>Dataset Description</h2> <p>All releases are composed of the following data:</p> <ul> <li>Time: in hours and&nbsp;minutes</li> <li>Power: in Watts</li> <li>Voltage: in Volts</li> <li>Current: in Ampers</li> <li>Generation power: in Watts</li> <li>Temperature: in &ordm;C</li> </ul>

opencc-by-4.0Dec 2024View details →
ClinicalTrials.gov36/100

Evaluation of Outcome Metrics in Alexander Disease

ClinicalTrials.gov study NCT02714764. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
dryad36/100

Data from: A multifaceted suite of metrics for comparative myoelectric prosthesis controller research

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad36/100

Echosounder derived metrics from small-scale coastal surveys

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Data for: How do we measure and increase systems thinking? Comparing self-reported and performative metrics in response to building causal loop models

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad36/100

Testing the utility of alternative metrics of branch support to address the ancient evolutionary radiation of tunas, stromateoids, and allies (Teleostei: Pelagiaria)

Open the record for dataset details and reuse information.

publicApr 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record