Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
552
datasets available to search
ShareScore release 0.7.1
Dataset results
552 results for “Security”
FixMe: An Incremental Lightweight Method for Vulnerability Data Collection for Security Patch Prediction
<div> <div>This repository has the FixMe dataset and the source code for extracting the new dataset. is a lightweight approach for collecting code patches based on analyzing the commits of various version control systems. The practical framework is designed to generate patches across a wide array of programming languages. This open-source tool streamlines the process of gathering vulnerability records from the Common Vulnerabilities and Exposures (CVE) database through an incremental approach. By embracing an incremental methodology, we expedite the acquisition of data, ensuring the inclusion of newly identified vulnerabilities and their corresponding patch pairs. Our methodology involves extracting security issues, obtaining vulnerability-fixing commits, and retrieving relevant source code from various projects. The extracted dataset by the FixMe tool supports for the automated patch prediction, automated program repair, commit classification, vulnerability prediction and more.</div> </div>
An Empirical Evaluation of the Relationship between Technical Debt and Software Security
<p>This dataset contains the static analysis results of 50 open source software applications retrieved from Github. The results were produced by using SonarQube, PMD, CKJM Extended and Findbugs tools. Further details can be found in the relevant publication: </p> <ul> <li>Siavvas, M., Tsoukalas, D., Janković, M., Kehagias, D., Chatzigeorgiou, A., Tzovaras, D., Aničić, N., Gelenbe, E. <em>An Empirical Evaluation of the Relationship between Technical Debt and Software Security</em>. In: Konjović, Z., Zdravković, M., Trajanović, M. (Eds.) ICIST 2019 Proceedings Vol.1, pp.199-203, 2019</li> </ul>
Survey data on people's forest use patterns and perceptions of border security measures in Białowieża Forest region, Poland
<p>A survey was conducted in June-July 2022 to obtain information about people's forest use patterns and opinions and feelings about border security measures (state of emergency, border zone closure, militarization) instituted in northeastern Poland starting in September 2021. Participants were informed that the survey was voluntary and anonymous. Participants were not obliged to respond to all questions and could stop the survey at any time. Survey completion and submission implied consent to participate. Participants had to be at least 18 years of age to take part in the survey. They had to be residents of the Białowieża Forest region. 100 persons participated in the survey. Of these 100 persons, 44 identified as local (born in the region). Data are coded and a key is provided. Some responses are aggregated and only responses to close-ended questions are shared, to prevent disclosure of potentially identifying information. </p>
Corpus of Resolutions: UN Security Council (CR-UNSC)
<h2>Overview</h2> <p>The <strong>Corpus of Resolution: UN Security Council (CR-UNSC)</strong> collects and presents for the first time in human and machine-readable form all resolutions, drafts, and meeting records of the UN Security Council, including detailed metadata, as published by the <a href="https://digitallibrary.un.org/">UN Digital Library</a> and revised by the authors.</p> <p>The United Nations Security Council (UNSC) is the most influential of the principal UN organs. Composed of five permanent and ten non-permanent members, its functioning is constrained by the political context in which it operates. During the Cold War, the complex political relationships between the permanent members and their veto powers significantly affected the capacity of the UNSC to address violations of international peace and security, with only 646 resolutions passed from 1946 to 1989. Since the 1990s, the activity of the UN Security Council has increased dramatically and produced 2721 resolutions up to the end of 2023. The length, complexity and thematic breadth of the resolutions has also increased, prompting calls to redefine it as a quasi-legislative body.</p> <p>Under Articles 24 and 25 of the UN Charter, member states have conferred upon the UNSC the "primary responsibility for the maintenance of international peace and security" and have agreed "to accept and carry out" its decisions. The discharge of this function is carried out through the powers bestowed upon it under Chapter VI of the UN Charter, "Pacific Settlement of Disputes", Chapter VII, "Action with Respect to Threats to the Peace, Breaches of the Peace, and Acts of Aggression", Chapter VIII, "Regional Arrangements", and Chapter XII, "International Trusteeship System". </p> <p>Under the peace and security mandate, its areas of activity cover disarmament, pacific settlement of disputes, enforcement, and, until 1994, strategic areas in a trusteeship agreement. Its functions also pertain to the correct working of the United Nations, covering issues of membership, the appointment of the Secretary General, the elections of judges of the International Court of Justice (ICJ), the calling of special and emergency sessions of the General Assembly, the amendment of the Charter and of the ICJ Statute.</p> <p><strong>Please refer to the Codebook for a detailed explanation of the dataset and instructions on how to make use of it.</strong></p> <p> </p> <h2>Updates</h2> <p>The CR-UNSC will be updated at least once per year.</p> <p>In case of serious errors an update will be provided at the earliest opportunity and a highlighted advisory issued on the Zenodo page of the current version. Minor errors will be documented in the GitHub issue tracker and fixed with the next scheduled release.</p> <p>The CR-UNSC is versioned according to the day of the last run of the data pipeline, in the ISO format YYYY-MM-DD. Its initial release version is 2024-05-03.</p> <p>Notifications regarding new and updated data sets will be published on my academic website at www.seanfobbe.com or on the Fediverse at @seanfobbe@fediscience.org</p> <p> </p> <h2>Changelog</h2> <ul> <li>New variant: EN_TXT_BEST containing a write-out of the English resolution texts equivalent to the CSV file text variable</li> <li>New diagrams: bar charts of top M49 regions and sub-regions of countries mentioned in resolution texts</li> <li>Fixed naming mix-up of BIBTEX and GRAPHML zip archives</li> <li>Fixed whitespace character detection in citation extraction (adds ca. 10% more citations)</li> <li>Fixed improper merging of weights in citation network</li> <li>Fixed "cannot xtfrm data frames" warning</li> <li>Improve REGEX detection for certain geographic entities</li> <li>Improve Codebook (headings, citation network docs)</li> </ul> <p> </p> <h2>Key Metrics</h2> <p><em>Version:</em> 2024-05-19</p> <p><em>Scope:</em> UNSC Resolutions from 1 (1946) up to and including 2722 (2024)</p> <p><em>Tokens:</em> 3,704,016 (English resolution texts)</p> <p><em>Languages: </em>English, French, Spanish, Arabic, Chinese, Russian</p> <p> </p> <h2>Features</h2> <ul> <li>82 Variables</li> <li>Resolution texts in all six official UN languages (English, French, Spanish, Arabic, Chinese, Russian)</li> <li>Draft texts of resolutions in English</li> <li>Meeting record texts in English</li> <li>URLs to draft texts in all other languages (French, Spanish, Arabic, Chinese, Russian)</li> <li>URLs to meeting record texts in all other languages (French, Spanish, Arabic, Chinese, Russian)</li> <li>Citation data as GraphML (UNSC-to-UNSC resolutions and UNSC-to-UNGA resolutions)</li> <li>Bibliographic database in BibTeX/OSCOLA format for e.g. Zotero, Endnote and Jabref</li> <li>Extensive Codebook to explain the uses of the dataset</li> <li>Compilation Report and Quality Assurance Report explain construction and validation of the data set</li> <li>Publication quality diagrams for teaching, research and all other purposes (PDF for printing, PNG for web)</li> <li>Open and platform independent file formats (CSV, PDF, TXT, GraphML)</li> <li>Software version controlled with Docker</li> <li>Publication of full data set (Open Data)</li> <li><a href="../doi/10.5281/zenodo.7319783">Publication of full source code (Open Source)</a></li> <li>Data published under Public Domain waiver (CC Zero 1.0)</li> <li>Source Code is Free Software published under the GNU General Public License Version 3 (GNU GPL v3)</li> <li>Secure cryptographic signatures for all files in version of record (SHA2-256 and SHA3-512)</li> </ul> <p> </p> <h2>Recommended Variants</h2> <table> <tbody> <tr> <td><strong>Traditional Scholars</strong></td> <td> <p>ALL_PDF_Resolutions</p> <p>EN_TXT_BEST</p> <p>BIBTEX_OSCOLA</p> </td> </tr> <tr> <td><strong>Quantitative Scholars</strong></td> <td> <p>ALL_CSV_FULL</p> <p>EN_TXT_BEST</p> <p>CITATIONS_GRAPHML</p> </td> </tr> </tbody> </table> <p> </p> <p>Please refer to the Codebook regarding for details on each variant. The ZIP archives include texts in all languages, unless noted in the filename.</p> <p>We strongly recommend using the CSV files for quantitative analysis, but if you find CSV hard to use and want to analyze only the text of resolutions, the EN_TXT_BEST variant is a mix of expert-revised OCR and born digital texts equivalent to the "text" variable in the CSV file.</p> <p> </p> <h2>Compilation Report and Quality Assurance Report</h2> <p>With every compilation of the full data set, an extensive Compilation Report and detailed Quality Assurance Report are created and published in PDF format.</p> <p>The Compilation Report includes the source code for the pipeline architecture, comments and explanations of design decisions, relevant computational results, exact timestamps and a table of contents with clickable internal hyperlinks to each section.</p> <p>The Quality Assurance Report contains a count of all hard tests and expectations, additional visualizations and documented test results for all soft tests that require further interpretation</p> <p>The Compilation Report, Quality Assurance Report and Source Code are published under the following DOI: <a href="../doi/10.5281/zenodo.7319783">https://zenodo.org/doi/10.5281/zenodo.7319783</a></p> <p> </p> <h2>Attribution and Copyright</h2> <p>This data is derived from the United Nations Digital Library at <a href="https://digitallibrary.un.org">https://digitallibrary.un.org</a>. Records were accessed and downloaded on 13 and 26 March 2024, with additional work on revisions and corrections up to and including the date given as the version number.</p> <p>Pursuant to <a href="https://en.wikisource.org/wiki/Administrative_Instruction_ST/AI/189/Add.9/Rev.2">UN Administrative Instruction ST/AI/189/Add.9/Rev.2 of 17 September 1987</a> all official records and United Nations Documents (including resolutions, compilations of resolutions, drafts and meeting records) are in the public domain. We wish to honor the letter and spirit of this UN policy. To ensure the widest possible distribution of official UN documents and to promote the international rule of law we waive any copyright that might have accrued by creating the dataset under a <a href="https://creativecommons.org/public-domain/cc0/">Creative Commons CC0 1.0 Universal (CC0 1.0) Public Domain Dedication</a>. </p> <p> </p> <h2>Disclaimer</h2> <p>This data set is an academic initiative and is not associated with or endorsed by the United Nations or any of its constituent organs and organizations.</p> <p> </p> <h2>Author Websites</h2> <p><a href="https://www.seanfobbe.com">Personal Website of Seán Fobbe</a></p> <p><a href="https://www.santannapisa.it/en/lorenzo-gasbarri">Personal Website of Lorenzo Gasbarri</a></p> <p><a href="https://www.kcl.ac.uk/people/niccolo-ridi">Personal Website of Niccolò Ridi</a></p> <p> </p> <h2>Contact</h2> <p>Did you discover any errors? Do you have suggestions on how to improve the data set? You can either post these to the <a href="https://github.com/SeanFobbe/cr-unsc/issues">Issue Tracker on GitHub</a> or contact Seán Fobbe via <a href="https://seanfobbe.com/contact/">https://seanfobbe.com/contact/</a></p>
Dataset for : A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
<p>We present a novel solution combining Large Language Model (LLM) capabilities with Formal Verification strategies to falsify and automatically repair software vulnerabilities. Initially, we employ Bounded Model Checking (BMC) to locate the software vulnerability and derive a counterexample. Relying on mathematical proofs, counterexamples provide evidence that the system behaves incorrectly or contains a vulnerability, thereby preventing the generation of false positive alerts. The counterexample that has been detected, along with the source code, are provided to the LLM engine. Our approach involves establishing a specialized prompt language for conducting code debugging and generation to understand the vulnerability's root cause and repair the code. Finally, we use BMC to verify the corrected version of the code generated by the LLM. As a proof of concept, we create \esbmcai based on the Efficient SMT-based Context-Bounded Model Checker (ESBMC) and a pre-trained Transformer model, specifically gpt-3.5-turbo, to detect and fix errors in C programs. We generated a dataset comprising $1{,}000$ C code samples, each consisting of $20$ to $50$ lines of C code. Experimental results show that our proposed method achieved an impressive success rate of up to $80$\% in repairing vulnerable code, encompassing buffer overflow, arithmetic overflow, and pointer dereference failures. To our knowledge, \esbmcai represents the first proposal for a pioneering initiative to integrate a Large Language Model (LLM) with software model checking. We advocate that this automated approach has the potential to incorporate into the software development lifecycle's continuous integration and deployment (CI/CD) process. </p> <p> </p> <p>The uploaded dataset contains 1000 codes, each comprising 20 to 50 lines of C code generated with gpt-3.5-turbo. The material also consists of a version of ESBMC statically compiled with all dependencies, a classifier script, and the output file.</p> <p> </p> <p> </p>
Repetition of Computer Security Warnings Results in Differential Repetition Suppression Effects as Revealed with Functional MRI
Open the record for dataset details and reuse information.
[Dataset] FP-Redemption: Measuring Browser Fingerprinting Adoption for the Sake of Web Security
<p>Full dataset for the paper "FP-Redemption: Measuring Browser Fingerprinting Adoption for the Sake of Web Security"</p> <p>5 files are provided:</p> <ul> <li>dataset.csv. The raw elements collected when browsing the web. Each entry corresponds to one attribute being accessed with one parameter combination by one script on one webpage. A single attribute with the same parameters can be accessed several times. It is represented with the key <em>nbTimes</em></li> <li>domainTags.csv: For each website, it provides its category and country tag.</li> <li>webpageTags.csv: For each webpage, it provides its type.</li> <li>fingerprinters.zip/<filenumber>.js: Our fingerprinters. Out of the 199 we requested, 7 are missing, leading in 192 js files.</li> <li>mapping.csv. 3 columns CSV file: <ul> <li>The first one lists the 199 fingerprinters detected by our algorithm.</li> <li>The second one gives the <filenumber> used to link a fingerprinter and its file in the directory.</li> <li>The third one gives the groups the fingerprinters belongs to. By default, each fingerprinter belongs to his own group. However, several fingerprinters are belonging to the same group as we evaluate there were duplicates. Thus, the number of distinct groups corresponds to the distinct fingerprinters we measured in our dataset: 169.</li> </ul> </li> </ul>
Data and Code for "Does Organic Farming Jeopardize Food Security of Farm Households in Benin?"
<p>This data and code archive provides all the data and code for replicating the empirical analysis that is presented in the journal article "<a href="https://doi.org/10.1016/j.foodpol.2024.102622" target="_blank" rel="noopener">Does Organic Farming Jeopardize Food Security of Farm Households in Benin?</a>" authored by Ghislain B.D. Aïhounton and Arne Henningsen and published in the journal Food Policy (Volume 124, April 2024, 102622, DOI: 10.1016/j.foodpol.2024.102622).</p> <p>We conducted the empirical analysis with the "R" statistical software (version 4.3.3) using the add-on packages "AER" (version 1.2.12), "DescTools" (version 0.99.54), "lmtest" (version 0.9.40), "moments" (version 0.14.1), "sandwich" (version 3.1.0), "stargazer" (version 5.2.3), and "xtable" (version 1.8.4) that are all available at CRAN.</p> <p>This replication package contains the following files:</p> <p>* README<br>This file.</p> <p>* R/dataBenin.csv<br>A CSV file that contains the (unprepared) data set. The variables in this file are described in file R/Variables.csv. This CSV file is imported by R script PrepareDataFoodNutrition.R.</p> <p>* R/Variables.csv<br>A CSV file that describes the variables in the (unprepared) data set (file R/dataBenin.csv).</p> <p>* R/PrepareData.R<br>An R script that imports the (unprepared) data set (file R/dataBenin.csv), calculates additional variables and add theses variables to the data set, removes observations that should not be used in the empirical analysis, and saves the prepared data set as CSV file (R/dataFoodNutrition.csv).</p> <p>* R/dataPrepared.csv<br>A CSV file that contains the (prepared) data set used in the empirical analysis. This CSV file is created by the R script R/PrepareDataFoodNutrition.R. It is imported by the R scripts R/DescriptiveTab.R, FoodNutritionImpact.R, and GridSearchFoodSecurity.R.</p> <p>* R/DescriptiveTab.R<br>An R script that imports the prepared data set (file R/dataFoodNutrition.R) and creates Table 1 of the paper ("Descriptive statistics", file paper/tables/DescriptiveStat.tex) as LaTeX file.</p> <p>* R/Estimations.R<br>An R script that imports the prepared data set (file R/dataFoodNutrition.R), conducts all the analyses presented in the paper, creates Tables 2 and 3 of the paper ("OLS and IV regression results of the conditional associations between organic farming and outcomes" and "OLS and IV regression results of the conditional associations between organic farming and mediating outcomes", LaTeX files paper/tables/estMainReg.tex and paper/tables/estMedReg.tex), creates Figures 1 and 2 of the paper ("Estimated conditional associations of organic farming with outcomes" and "Estimated conditional associations of organic farming with mediating outcomes", 12 PDF files paper/figures/*.pdf), and 45 Tables that are included in the Supplementary Information: 36 tables with detailed regression results (LaTeX files paper/tables/tabels/est*.tex), one table with results of the first-stage probit regression (LaTeX file paper/tables/tabels/estProbit.tex), 6 tables with detailed regression results of estimations for testing the exogeneity of the instrument as suggested by Di Falco et al. (2011) (LaTeX files paper/tables/tabels/estOLS*Falco.tex), and 2 tables with coefficient bounds obtained as suggested by Oster (2019) (LaTeX files paper/tables/tabels/Oster*.tex).</p> <p>* R/GridSearch.R<br>An R script that re-runs our regression analyses with different units of measurement of IHS-transformed variables and calculates various indicators that can can be used to assess the appropriateness of different units of measurement as suggested by Aihounton and Henningsen (2021) and that creates 28 Tables that are included in the Supplementary Information (LaTeX files paper/tables/tabels/grid*.tex).</p> <p>* R/functions/calcOsterBounds.R<br>An R script that defines the R function calcOsterBounds() that calculates coefficient bounds using the method suggested by Oster (2019). This function is used by the R script R/FoodNutritionImpact.R.</p> <p>* R/functions/calcSemiElaOrg.R<br>An R script that defines the R function calcSemiElaOrg() that calculates the semi-elasticity of various log-transformed or IHS-transformed variables with respect to the dummy variable for organic farming. This function is used by the R scripts R/FoodNutritionImpact.R and R/GridSearchFoodSecurity.R.</p> <p>* R/functions/createFormula.R<br>An R script that defines the R function createFormula() that creates the regression formulas for the various empirical analyses that are presented in the paper. This function is used by the R scripts R/FoodNutritionImpact.R and R/GridSearchFoodSecurity.R.</p> <p>* R/functions/functionsTables.R<br>An R script that defines various R functions that are used to create tables in LaTeX format. These functions are used by the R scripts R/FoodNutritionImpact.R and R/GridSearchFoodSecurity.R.</p> <p>* R/functions/predR2.R<br>An R script that defines the R function predR2() that calculates the predictive R-squared value. This R script has been obtained from the replication package of the article:<br>Aïhounton, G. B. D. and Henningsen, A. (2021). Units of measurement and the inverse hyperbolic sine transformation. The Econometrics Journal, 24(2):334–351. https://doi.org/10.1093/ectj/utaa032<br>The function consists of a slightly modified version of the code that is available at: https://tomhopper.me/2014/05/16/can-we-do-better-than-r-squared/ This function is used by the R script R/GridSearchFoodSecurity.R.</p> <p>* paper/figures/*.pdf<br>12 LaTeX files that are the (sub)figures in Figures 1 and 2 of the paper ("Estimated conditional associations of organic farming with outcomes" and "Estimated conditional associations of organic farming with mediating outcomes"). These 12 files are created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/DescriptiveStat.tex<br>A LaTeX file that creates Table 1 of the paper ("Descriptive statistics"). This file is created by the R script R/DescriptiveTab.R.</p> <p>* paper/tables/estMainReg.tex<br>A LaTeX file that creates Table 2 of the paper ("OLS and IV regression results of the conditional associations between organic farming and outcomes"). This file is created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/estMedReg.tex<br>A LaTeX file that creates Table 3 of the paper ("OLS and IV regression results of the conditional associations between organic farming and mediating outcomes"). This file is created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/tabels/est*.tex<br>36 LaTeX files that create 36 tables that are included in the Supplementary Information and present detailed regression results. These 36 files are created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/tabels/estProbit.tex<br>A LaTeX files that creates a table that is included in the Supplementary Information and presents the results of the first-stage probit regression. This file is created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/tabels/estOLS*Falco.tex<br>6 LaTeX files that create 6 tables that are included in the Supplementary Information and present detailed regression results for testing the exogeneity of the instrument as suggested by Di Falco et al. (2011). These 6 files are created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/tabels/Oster*.tex<br>2 LaTeX files that create 2 tables that are included in the Supplementary Information and present coefficient bounds obtined as suggested by Oster (2019). These 2 files are created by the R script R/FoodNutritionImpact.R.</p> <p>* paper/tables/tabels/grid*.tex<br>28 LaTeX files that create 28 tables that are included in the Supplementary Information and present various indicators for assessing the appropriateness of different units of measurement of IHS-transformed variables as suggested by Aihounton and Henningsen (2021). These 28 files are created by the R script R/GridSearchFoodSecurity.R</p>
Analyzing and Mitigating (with LLMs) the Security Misconfigurations of Helm Charts from Artifact Hub
<p>In the corresponding scientific paper, we proposed a pipeline to mine Helm charts from Artifact Hub, a popular centralized repository, and analyze them using state-of-the-art open-source tools like Checkov and KICS. First, such a pipeline runs several chart analyzers and identifies the common and unique misconfigurations reported by each tool. Secondly, it uses LLMs to suggest mitigation for each misconfiguration. Finally, the chart refactoring previously generated is analyzed again by the same tools to see whether it satisfies the tool's policies.</p> <p>In this dataset, you can find all the Helm chart templates downloaded from Artifact Hub (available in June 2024), all the outputs of the tools analyzing such templates, the CSV result files with all LLM queries and answers, and the snippets selected for the manual analysis.</p>
Data from: Saltmarsh vegetation and secured woody debris facilitate mangrove re-colonization
<p>Does the presence of saltmarsh vegetation affect the long-term regeneration of the pioneer mangrove species <em>Avicennia germinans</em> in a degraded dwarf forest? Does immobilized coarse woody debris (CWD) affect regeneration similarly? Do larger trees suppress or facilitate intraspecific saplings? The study was conducted in a dwarf mangrove forest in the high intertidal zone on Bragança peninsula in northern Brazil. The spatial patterns of <em>A. germinans</em>, the herbaceous halophyte <em>Sesuvium portulacastrum</em>, and CWD were mapped in three sample plots (each 400 m<sup>2</sup>) during two consecutive vegetation surveys, conducted in 2011 and 2014. Inhomogeneous Poisson and Thomas point-process models were used to assess the distribution of <em>A. germinans</em> life-history stages (seedlings, saplings, and adult dwarf trees), conditioned on the presence of <em>S. portulacastrum</em> and CWD. In addition, intraspecific interactions between trees and regeneration were assessed based on crown projection mapping. Bivariate point pattern analyses were used to assess the dependence of advance regeneration on dwarf <em>A. germinans</em> trees and <em>S. portulacastrum</em>. <em>A. germinans</em> saplings and trees were positively associated with <em>S. portulacastrum</em> and CWD, whereas seedlings were located around tree crowns. The density of fruit-bearing trees was positively associated with sapling density, indicating that regeneration relied on locally dispersed propagules. Herbaceous vegetation and CWD have an important ecological function in degraded mangroves by retaining tidally dispersed propagules. Here, we show that herbaceous vegetation does not suppress the growth of seedlings but facilitates mangrove recolonization. Due to limited tidal dispersal, regeneration relies on local propagule supply. In addition to hydrological restoration, the observed vegetation patterns suggest that, in the absence of propagule-retaining vegetation, restoration of high-intertidal mangroves can be facilitated by establishing nuclei of planted trees and installing secured logs.</p>
Are marriage-related taxes and Social Security benefits holding back female labor supply?
<p>Margherita Borella, Mariacristina De Nardi, and Fang Yang, "Are marriage-related taxes and Social Security benefits holding back female labor supply?". The Review of Economic Studies, Forthcoming. This package contains all the code and detailed instructions on how to replicate the data analysis and numerical results in the paper.</p>
Defeating Adversarial Attacks Againt Adversarial attacks in Network Security
<p>We investigate if the feature randomization approach to improve the robustness of forensic detectors to targeted attacks in network security, can be extended to detectors based on deep learning features. In particular, we study the transferability of adversarial examples targeting an original CNN image manipulation detector to other detectors that rely on a random subset of the features extracted from the flatten layer of the original network. The results we got by considering, two original network architectures and different classes of attacks, show that feature randomization helps to hinder attack transferability, even if, in some cases, simply changing the architecture of the detector, or even retraining the detector is enough to prevent the transferability of the attacks.</p>
Assessing ambitious nature conservation strategies in a below 2-degree and food-secure world – supplementary spatial data
<p><strong>Assessing ambitious nature conservation strategies in a below 2-degree and food-secure world – supplementary spatial data</strong></p><p><strong>Authors: </strong>Marcel Kok, Johan Meijer, Willem-Jan van Zeist, Jelle Hilbers, Marco Immovilli, Jan Janse, Elke Stehfest, Michel Bakkenes, Andrzej Tabeau, Aafke Schipper, Rob Alkemade</p><p><strong>Point of contact:</strong> <a href="mailto:Marcel.Kok@pbl.nl">Marcel.Kok@pbl.nl</a></p><p><strong>Research paper summary:</strong> Global biodiversity is projected to further decline under a wide range of future socio-economic development pathways, even in sustainability-oriented scenarios. This raises the question how biodiversity can be put on a path to recovery, the core challenge for the implementation of the CBD Kunming-Montreal Global Biodiversity Framework. We designed two ambitious global conservation strategies, 'Half Earth' (HE) and 'Sharing the Planet' (SP), and evaluated their ability to restore terrestrial and freshwater biodiversity and to provide nature's contributions to people (NCP), while also limiting global warming below 2 degrees and ensuring food security. We applied the integrated assessment framework IMAGE with the GLOBIO biodiversity model, using the 'Middle of the Road' Shared Socio-economic Pathway (SSP2) with its projected human population growth as baseline. We found that the HE strategy performs generally better for terrestrial biodiversity (biodiversity intactness (MSA), Area of Habitat, Living Planet Index, Red List Index) in currently still natural regions. The SP strategy yields more improvements for biodiversity in human-used areas, for freshwater biodiversity and for regulating NCP (pest control, pollination, erosion control, water quality). However, both strategies were insufficient to restore biodiversity and corresponded with considerable increases in food security risks and global temperature. Only when we combined the conservation strategies with a portfolio of 'integrated sustainability measures', including climate change mitigation and reductions of food waste and animal product consumption, our scenarios resulted in a restoration of biodiversity and NCP while keeping global warming below two degrees and food security risks below the baseline projection.</p><p><strong>Contents:</strong> This repository contains the supplementary spatial data describing the specific prioritization of conservation areas under the Half Earth (HE) and Sharing the Planet (SP) scenarios, and the resulting scenario land use and MSA data sets for the year 2050, including also a baseline (BL) scenario. All spatial data is in geotiff format at a 10 arcsecond resolution in WGS84 coordinate system. Detailed description of the methodology is provided in the paper listed under "related identifiers".</p><p><strong>Keywords:</strong> Nature conservation, Half Earth, Sharing the Planet, Climate Change, Food Security, Solution-oriented scenarios, Biodiversity, Nature's Contribution to People, NCP</p>
Hybrid Deep Learning Techniques for Securing Bioluminescent Interfaces in Internet of Bio Nano Things
<p>The data-set presents normal and anomalous values of twelve traffic parameters, generated by <strong>Bioluminescent bio-cyber Interfacing </strong>(BBI) in the I<strong>nternet of Bio Nano Things </strong>(IoBNT) based systems.</p> <p>The traffic parameters included in the data-set represent bio-electric and electro-bio transduction unit operation of BBI incorporating normal, as well as abnormal data to train and test machine/deep learning classifiers in discriminating attack scenarios.</p> <p>The parameters considered include the following: <strong>Cumulative concentration of released molecules, Elimination rate, Michaelis-Menten constant, Kinetic constant, Forward rate constant, Catalytic reaction constant, Ligand-receptor binding constant, Concentration of ATP, Concentration of information molecules, Release rate Reverse kinetic constant,</strong> and <strong>Reverse forward rate constant.</strong></p> <p>The data set is divided into training and testing data for simplified analysis, and application.</p>
Taxonomy of Security Weaknesses in Java and Kotlin Android Apps
<p>This is the replication package for the paper "Taxonomy of security weaknesses in Java and Kotlin Android apps" accepted for inclusion in The Journal of Systems & Software</p>
Supplementary Materials for "Food security in Roman Palmyra (Syria) in light of paleoclimatological evidence and its historical implications"
<p>Contained here are the SI files for the article "Food security in Roman Palmyra (Syria) in light of paleoclimatological evidence and its historical implications". With all the materials contained here, as well as the openly accessible datasets cited in S1_File, every step of the study can be reproduced. Detailed instructions are contained within. Includes code for Data Analysis.</p> <p>Article DOI: [forthcoming]</p>
Result data related to Tröndle et al (2024): Rebuilding Ukraine's energy supply in a secure, economic, and decarbonised way
<p>This dataset contains the result data of all the scenarios ran in the scientific article "Rebuilding Ukraine’s energy supply in a secure, economic, and decarbonised way".</p> <p>The results of the main five scenarios of the study are available as PyPSA result files:</p> <ul> <li>nuclear-and-renewables-high.nc: A scenario with nuclear in the mix and high economic growth assumption.</li> <li>nuclear-and-renewables-low.nc: A scenario with nuclear in the mix and low economic growth assumption.</li> <li>only-renewables-high-low-bio.nc: A scenario with only renewables, high economic growth assumption, and only 10% of assumed biomass potential.</li> <li>only-renewables-high.nc: A scenario with only renewables and high economic growth assumption.</li> <li>only-renewables-low.nc: A scenario with only renewables and low economic growth assumption.</li> </ul> <p>See the PyPSA documentation for more information: <a href="https://pypsa.readthedocs.io" target="_blank" rel="noopener">https://pypsa.readthedocs.io</a>.</p> <p>The results of the 330 global sensitivity analysis runs are available as summary files in CSV format:</p> <ul> <li>gsa-capacities-energy-gwh.csv: The installed energy storage capacities for each scenario.</li> <li>gsa-capacities-power-gw.csv: The installed generation capacities for each scenario.</li> <li>gsa-lcoe.csv: The levelised cost of electricity for each scenario.</li> </ul>
Dataset from "Natural capital accounting reveals ecosystems' role in water and energy security in Colombia's Sinú Basin"
<p>The data archived here are associated with the publication titled "Natural capital accounting reveals ecosystems' role in water and energy security in Colombia's Sinú Basin", available at: <a href="https://doi.org/10.1038/s43247-025-02254-9">https://doi.org/10.1038/s43247-025-02254-9</a>. The files within "Sinu_SDR_inputs.zip" and "Sinu_SWY_inputs.zip" were prepared and run in <a href="http://releases.naturalcapitalproject.org/?prefix=invest/3.12.0/">InVEST version 3.12.0</a>. "SDR" refers to the InVEST Sedimnet Delivery Ratio (SDR) model and "SWY" refers to the InVEST Seasonal Water Yield (SWY) model. Results of these model runs are found within "Sinu_SDR_results.zip" and "Sinu_SWY_results.zip" for the SDR and SWY models, respectively. These models were calibrated using observed data on average monthly water flows (from 1959 to 1992) and average annual sediment loads (from 1972 to 1992) from gauge stations on Colombia's Sinú River. Those observed data were obtained from Colombia's Institute of Hydrology, Meteorology, and Environmental Studies (IDEAM) hydrometeorological monitoring network <a href="http://dhime.ideam.gov.co/atencionciudadano/">webportal</a> and are summarized in the files included here, "MeanMonthlyObservedFlowsXgaugeStation.csv" for monthly water flows and "annualObservedSedimentXgaugeStation.csv" for annual sediment loads. "EcosystemTypeTable.xlsx" is the table of ecosystem values. "Cuenta_Sinu_SankeyData_v2_paper.xlsx" contains the Sankey and accounts tables.</p>
Security Bug Conversations
<p>This dataset will be released as part of the following publication.</p> <ul> <li>Benjamin S. Meyers, Nuthan Munaiah, Andrew Meneely, and Emily Prud'hommeaux. <strong>Pragmatic Characteristics of Security Conversation: An Exploratory Linguistic Analysis. </strong><em>Forthcoming.</em><strong> </strong>Proceedings of the 12th International Workshop on Cooperative and Human Aspects of Software Engineering (CHASE 2019). Montréal, QC, Canada.</li> </ul> <p><strong>Files:</strong></p> <pre><code>security_bug_conversations.csv</code></pre> <p>The full dataset containing over 2.1 million comments posted by developers discussing bugs in the Chromium project. The dataset also includes the values we calculated for the five pragmatic features (described in Section 3 of the paper cited above).</p> <p><strong>CSV Fields:</strong></p> <ul> <li><strong>Organizational:</strong> <ul> <li><em>Bug ID:</em> Unique identifier of a bug discussion in the Chromium project. The URL https://bugs.chromium.org/p/chromium/issues/detail?id=<Bug ID> may be used to access the bug online</li> <li><em>Comment ID:</em> Unique identifier of a comment in a bug discussion</li> </ul> </li> <li><strong>Classification:</strong> <ul> <li><em>Is Security:</em> Binary indicator of whether or not a comment is part of a bug that is about security</li> </ul> </li> <li><strong>Natural Language:</strong> <ul> <li><em>Comment Text:</em> The raw natural language text of the bug comment</li> </ul> </li> <li><strong>Linguistic Metrics:</strong> <ul> <li><em>Min. Formality:</em> Minimum of the formality of sentences in the bug comment</li> <li><em>Max. Formality:</em> Maximum of the formality of sentences in the bug comment</li> <li><em>Max. Informativeness:</em> Maximum of the informativeness of sentences in the bug comment</li> <li><em>Max. Implicature:</em> Maximum of the implicature of sentences in the bug comment</li> <li><em>Min. Politeness:</em> Minimum of the politeness of sentences in the bug comment</li> <li><em>Max. Politeness:</em> Maximum of the politeness of sentences in the bug comment</li> <li>Number of Tokens</li> <li>Number of Sentences</li> <li><em>Has Doxastic Uncertainty:</em> Binary indicator of presence of a sentence with doxastic uncertainty in the bug comment</li> <li><em>Has Epistemic Uncertainty:</em> Binary indicator of presence of a sentence with epistemic uncertainty in the bug comment</li> <li><em>Has Conditional Uncertainty:</em> Binary indicator of presence of a sentence with conditional uncertainty in the bug comment</li> <li><em>Has Investigational Uncertainty:</em> Binary indicator of presence of a sentence with investigational uncertainty in the bug comment</li> <li><em>Has Uncertainty:</em> Binary indicator of presence of a sentence with any uncertainty in the bug comment</li> </ul> </li> </ul>
Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis
<p>This repository contain datasets and results for the paper:</p> <p><strong>Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis</strong></p> <p> </p> <p><strong>Github repository for the code: </strong></p> <p><a href="https://github.com/siebeniris/QuantifyingLanguageConfusion/tree/main">Quantifying Language Confusion GitHub repo</a></p> <p> </p> <p><strong>DATA</strong> include the following datasets:</p> <p>i) raw language graphs and</p> <p>ii) the calculated language similarities from the language graphs,</p> <p>iii) <strong>MTEI</strong>: the files from the <a href="https://github.com/siebeniris/vec2text_exp/tree/aaai">experimental results of multilingual inversion attacks</a>, and calculated language confusion entropy from the data;</p> <p>iv) <strong>LCB</strong>: the files from the <a href="https://github.com/for-ai/language-confusion?tab=Apache-2.0-1-ov-file#readme">language confusion benchmark</a> and calculated language confusion entropy from the data </p> <p> </p> <p><strong>Results</strong> include aggregated results for further analysis:</p> <p>i) <strong>inversion_language_confusion</strong>: results from MTEI</p> <p>ii) <strong>prompting_language_confusion</strong>: results from LCB</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.