Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
119
datasets available to search
ShareScore release 0.7.1
Dataset results
119 results for βPrivacyβ
Privacy Policies Paragraph containing Personal Data
<p>The data consists in crawled privacy policies from European privacy policies. They were split into paragraphs and annotated as containing or not personal data.</p> <p>The question that was asked to annotators was "Does this paragraph contain the explicit mention of specific personal data (e.g. name, phone number, social security, …) being collected?".</p> <p>A full description of the dataset can be found in D3.4 of the SMOOTH project</p>
CVoiceFake (crafted by SafeEar: Content Privacy-Preserving Audio Deepfake Detection)
<h1><strong>Introduction:</strong></h1> <p>CVoiceFake (small) is a dataset that features a random selection of 10% of samples from the entire collection. This dataset encompasses <strong>five common languages (English, Chinese, German, French, and Italian)</strong> and utilizes <strong>multi-advanced and classical voice cloning techniques</strong> (Parallel WaveGAN, Multi-band MelGAN, Style MelGAN, Griffin-Lim, WORLD, and DiffWave) to produce audio samples that bear a high resemblance to authentic audio.</p> <ol> <li><strong>Parallel WaveGAN</strong>: As a non-autoregressive vocoder-based model, Parallel WaveGAN produces high-fidelity audio rapidly, ideal for efficient and quality deepfake generation.</li> <li><strong>Multi-band MelGAN</strong>: Multi-band MelGAN is a variant of MelGAN that divides the frequency spectrum into sub-bands for faster and more stable multi-lingual vocoder training, enhancing the robustness and scalability of the dataset.</li> <li><strong>Style MelGAN</strong>: Style MelGAN is designed to capture fine prosodic and stylistic nuances of speech, making it particularly compelling for deepfake applications that require high levels of expressivity and variation in speech synthesis.</li> <li><strong>Griffin-Lim</strong>: This algorithm reconstructs waveforms from spectrograms using an iterative phase estimation method. Though less high-fidelity than neural vocoders, it serves as a traditional baseline for comparing deepfake generation.</li> <li><strong>WORLD</strong>: WORLD is a statistical parameter-based voice synthesis system that offers fine control over the spectral and prosodic features of the synthesized audio. Its fine manipulation is useful for crafting the nuanced variations needed in deepfake datasets.</li> <li>We have also built the SOTA diffusion-based deepfake audio (DiffWave); please contact the author at <code>xinfengli@zju.edu.cn</code> if you are interested in the dataset, particularly the DiffWave portion. Furthermore, any additional discussions are welcomed.<br><strong>DiffWave</strong>: DiffWave is a diffusion probability model for waveform generation. It converts the white noise signal into structured waveform through a Markov chain, capable of both conditional and unconditional generation tasks. DiffWave represents the advanced synthesis method for its fast synthesis speed and high synthesis quality.</li> </ol> <h1><strong>π₯</strong><strong>News:</strong></h1> <p>Please note that we recently released our DiffWave subset in Version 2 in comparison to Version 1, which is available on <a href="../records/14062964" target="_blank" rel="noopener">CVoiceFake Full</a>. You can download the file named CVoiceFake_Large_diffwave_update.tar.gz.xx, and after unzipping it, you will find it retains the same file structure as before.<br> <strong>| CVoiceFake_Large_diffwave_update.tar.gz.00 |<br> | CVoiceFake_Large_diffwave_update.tar.gz.01 |</strong></p> <p> </p> <h1><strong>Full Dataset & Project Page:</strong></h1> <p>The whole dataset is available on <a href="../records/14062964" target="_blank" rel="noopener">CVoiceFake Full</a> as well. Please kindly also refer to the project page: <a title="SafeEar Website" href="https://safeearweb.github.io/Project/" target="_blank" rel="noopener">SafeEar Website</a>.</p> <p> </p> <h1><strong>Citation:</strong></h1> <p>If you find our paper/code/benchmark helpful, please kindly consider citing this work with the following reference:</p> <pre><code>@inproceedings{li2024safeear,<br> author = {Li, Xinfeng and Li, Kai and Zheng, Yifan and Yan, Chen and Ji, Xiaoyu, and Xu, Wenyuan},<br> title = {{SafeEar: Content Privacy-Preserving Audio Deepfake Detection}},<br> booktitle = {Proceedings of the 2024 {ACM} {SIGSAC} Conference on Computer and Communications Security (CCS)}<br> year = {2024},<br>} </code></pre> <div> <div> </div> </div>
Cross-language corpora of privacy policies
<p>The dataset consists of three different privacy policy corpora (in English and Italian) composed of 81 unique privacy policy texts spanning the period 2018-2021. This dataset makes available an example of three corpora of privacy policies. The first corpus is the English-language corpus, the original used in the study by Tang et al. [2]. The other two are cross-language corpora built (one, the source corpus, in English, and the other, the replication corpus, in Italian, which is the language of a potential replication study) from the first corpus.</p> <p>The policies were collected from:</p> <ol> <li>the Alexa top 10 Italy and U.S. websites rank;</li> <li>the Play Store apps rank in the "most profitable games" category of the Play Store for Italy and the U.S.</li> </ol> <p>We manually analyzed the Alexa top 10 Italy websites as of November 2021. Analogously, we analyzed selected apps that, in the same period, had ranked better in the "most profitable games" category of the Play Store for Italy.</p> <p>All the privacy policies are ANSI-encoded text files and have been manually read and verified.<br> The dataset is helpful as a starting point for building comparable cross-language privacy policies corpora. The availability of these comparable cross-language privacy policies corpora helps replicate studies in different languages. <br> Details on the methodology can be found in the accompanying paper.</p> <p>The available files are as follows:</p> <ul> <li><strong>policies-texts.zip</strong> --> contains a directory of text files with the policy texts. File names are the SHA1 hashes of the policy text.</li> <li><strong>policy-metadata.csv</strong> --> Contains a CSV file with the metadata for each privacy policy.</li> </ul> <p>This dataset is the original dataset used in the publication [1]. The original English U.S. corpus is described in the publication [2].</p> <p>[1] F. Ciclosi, S. Vidor and F. Massacci. "Building cross-language corpora for human understanding of privacy policies." Workshop on Digital Sovereignty in Cyber Security: New Challenges in Future Vision. Communications in Computer and Information Science. Springer International Publishing, 2023, In press.</p> <p>[2] J. Tang, H. Shoemaker, A. Lerner, and E. Birrell. Defining Privacy: How Users Interpret Technical Terms in Privacy Policies. Proceedings on Privacy Enhancing Technologies, 3:70–94, 2021.</p>
Supplementary Materials for "Exploration of User Privacy in 802.11 Probe Requests with MAC Address Randomization Using Temporal Pattern Analysis"
<p>Supplementary Materials for "Exploration of User Privacy in 802.11 Probe Requests with MAC Address Randomization Using Temporal Pattern Analysis"</p> <p>This package contains an anonymized packets of 802.11 probe requests captured in in December 2021 at Universitat Jaume I . The packet capture file is in the standardized *.pcap binary format and can be opened with any packet analysis tool such as Wireshark or scapy (Python packet analysis and manipulation package).</p>
Accepted Artifact for Privacy-Respecting Type Error Telemetry at Scale
<p>This artifact packages the data for the paper: <em>Privacy-Respecting Type Error Telemetry at Scale</em></p> <p>There are two files on Zenodo:</p> <ul> <li>data.tar.gz has the original Luau telemetry data</li> <li>artifact.tar.gz has a result PDF, intermediate data, and scripts for processing the data</li> </ul> <p>The artifact code and the source for the paper are also on GitHub:</p> <ul> <li><a href="https://github.com/bennn/luau-telemetry">https://github.com/bennn/luau-telemetry</a></li> </ul> <p>This artifact is primarily a **dataset**. It shows how we reached the conclusions in the paper.</p> <p>The scripts in this artifact are provided as-is for completeness. They may have bugs. They may not work as advertised.</p>
Greek privacy policies dataset from PCI 2023 paper: "A privacy policies dataset in Greek in the GDPR era"
<p>A dataset of privacy policies in the Greek language, with policies coming from top visited websites in Greece with a privacy policy in the Greek language.</p> <p>The dataset, as well as results of its analysis are included.</p> <p>if you want to use this dataset, please cite the relevant conference publication:</p> <p>Georgia M. Kapitsaki and Maria Papoutsoglou, "A privacy policies dataset in Greek in the GDPR era, in Proceedings of the 27th Pan-Hellenic Conference on Informatics, PCI 2023.</p> <p> </p>
Supplementary materials for the paper "Users' Privacy Concerns and Attitudes towards Usage-Based Insurance: an empirical approach"
<p>These are materials necessary to replicate the study discussed in <em>Users' Privacy Concerns and Attitudes towards Usage-Based Insurance: an empirical approach</em>, accepted for publication at VEHITS 2022.</p>
Improving students' privacy awareness β Analysis of a pilot survey to design a VR environment for self-paced learning
<p>In this research, we measured the knowledge of students at the University of Debrecen in the field of data privacy awareness, online and password security.</p> <p><strong>Description</strong></p> <ul> <li>In the questionnaire, green-highlighted answer signs the correct answer to each question.</li> <li>Total data set contains the answers to each question and the respondent's age.</li> <li>Correct/incorrect data set contains information if the answer is correct to each question, and it also contains the respondent's age. One means the answer was correct, and zero means the answer was incorrect.</li> </ul>
Privacy-by-Design Maturity Model: literature review, coding, model creation and evaluation
<p>Results from two multivocal literature reviews (MLRs) and subsequent coding, formulation of capabilities and dependencies, creation of maturity matrix and evaluation results. Used in the creation of a PbD domain model and extraction of core activities for PbD in Information Systems design. Part of the <a href="https://www.privacymaturity.org/" target="_blank" rel="noopener">Privacy-by-Design Maturity</a> research project by the <a href="https://www.uu.nl/en/research/ai-labs/ai-lab-for-the-public-services" target="_blank" rel="noopener">AI Lab for Public Services</a>.</p>
Mobile Application Privacy Risk Assessments from User-authored Scenarios
<p>Mobile applications (apps) provide users valuable benefits at the risk of exposing users to privacy harms. Improving privacy in mobile apps faces several challenges, in particular, that many apps are developed by low resourced software development teams, such as end-user programmers or in startups. In addition, privacy risks are primarily known to users, which can make it difficult for developers to prioritize privacy for sensitive data. In this paper, we introduce a novel, lightweight method that allows app developers to elicit scenarios and privacy risk scores from users directly using only an app screenshot. The technique relies on named entity recognition (NER) to identify information types in user-authored scenarios, which are then fed in real-time to a privacy risk survey that users complete. The best-performing NER model predicts information types with a weighted average precision of 0.70 and recall of 0.72, after post-processing to remove false positives. The model was trained on a labeled 300-scenario corpus, and evaluated in an end-to-end evaluation using an additional 203 scenarios yielding 2,338 user-provided privacy risk scores. Finally, we discuss how developers can use the risk scores to prioritize, select and apply privacy design strategies in<br> the context of four user-authored scenarios.</p>
Database for article: "Privacy Perceptions in Digital Games: A Study with Information Technology (IT) Undergraduates"
<p>This database is an addendum to the article "<strong>Privacy Perceptions in Digital Games: A Study with Information Technology (IT) Undergraduates</strong>" to provide information regarding the anonymously collected data.</p><p><strong>Abstract of the article</strong></p><p>This study explores the perceptions and practices of undergraduates in Information Technology (IT) regarding privacy issues in digital games. This topic becomes relevant in the current scenario where artificial intelligence (AI) is increasingly integrated into digital games, providing an enhanced experience for players. However, this integration poses security and privacy challenges, the understanding of which is crucial for both players and developers.<br>The primary objective of this research is to comprehend the participants' perceptions and understandings of privacy in digital games. We employed a qualitative and quantitative methodology to address our research inquiries. Through an online form of data collection, we obtained 61 responses. Among the obtained information, we observed that 40\% of the students are interested in pursuing a career in game development, and 49.18% would consider this possibility. Noteworthy among the identified issues is the necessity for companies to devise more effective means of communicating their privacy policies to players/users, adapting the language to their target audience. Participants reported attacks related to online multiplayer games and expressed concerns about the security of personal data.</p>
Designing a Training Journey for Privacy and Information Security Practitioners in the Federal Public Administration
<p> Context : The Ministry of Management and Innovation in Public Services (MGI) leads the formulation and coordination of the Digital Government Strategy (EGD). DEPSI, under the Secretariat of Digital Government (SGD), is responsible for the Privacy and Information Security Program (PPSI), which aims at data privacy, compliance, and institutional resilience. Problem: The culture of privacy and information security in the Federal Public Administration faces development challenges. Despite the PPSI, there is a lack of awareness initiatives, training, a clear strategy, best practices, and performance indicators. Proposed Solution: The proposal aims to develop a training journey for Practitioners working in roles related to privacy and information security, with the goal of identifying and promoting the best practices, skills, and competencies required for these roles. IS Theory: This study aligns with Organizational Information Processing Theory, providing mechanisms to help organizations adapt to regulatory uncertainties in privacy and security. Method: We employed a mixed approach, combining document analysis, a literature review, and a survey. Guidelines and standards were analyzed to map competencies and responsibilities, while the survey gathered practitioners' perceptions of the proposed training journey. Summary of Results: We identified the key profiles and their corresponding responsibilities, and proposed a personalized training journey. Survey results indicated that the journey meets Practitioners’ expectations, being well-evaluated in terms of criteria and assigned weights. Contributions and Impact in IS: This work contributes by presenting a proposal for a Training Journey to assess the knowledge of Federal Public Administration employees and guide them on the best paths for professional development. </p>
Contact tracing solutions for COVID-19: applications, data privacy and security :: Suplementary Material
<p>Supplementary Material for the paper "Contact tracing solutions for COVID-19: applications, data privacy and security"</p>
Data for research article "Privacy Explanations β A Means to End-User Trust"
<p>Research data for article "<strong>Privacy Explanations – A Means to End-User Trust</strong>". This package includes the survey and its results.</p>
Source Location Privacy Aware Routing Protocols Selection Results
<p>This is the dataset used to generate the results for the journal paper " A Decision Theoretic Framework for Selecting Source Location Privacy Aware Routing Protocols in Wireless Sensor Networks" at Future Generation Computer Systems (FGCS) 2018.</p> <p> </p> <p> </p>
FedscGen: privacy-aware federated batch effect correction of single-cell RNA sequencing data -- Preprocessed datasets
<div> <div> <div> <div> <p>This dataset accompanies the publication "FedscGen: Privacy-Aware Federated Batch Effect Correction of Single-Cell RNA Sequencing Data" and includes eight single-cell RNA sequencing (scRNA-seq) datasets used to benchmark the FedscGen and scGen methods. The datasets are provided in <code>.h5ad</code> format and include comprehensive metadata necessary for replication and further analysis.</p> <h3>Datasets</h3> <p>We analyze various datasets to compare FedscGen against scGen (centralized) in terms of batch correction. For simplicity, we refer to the dataset by abbreviations:</p> <ol> <li> <p><strong>Cell Line (CL)</strong>:</p> <ul> <li>Derived from the 293t_jurkat experiment with three batches: Zheng et al., 2017.</li> </ul> </li> <li> <p><strong>Human Dendritic Cells (HDC)</strong>:</p> <ul> <li>scRNA-seq data of human dendritic cells across two batches: Villani et al., 2017.</li> </ul> </li> <li> <p><strong>Human Pancreas (HP)</strong>:</p> <ul> <li>Consolidated data from five sources with 14,767 cells each: Baron et al., 2016; Muraro et al., 2016; Segerstolpe et al., 2016; Wang et al., 2016; Xin et al., 2016.</li> </ul> </li> <li> <p><strong>Mouse Brain (MB)</strong>:</p> <ul> <li>Merged datasets with 691,600 and 141,606 cells: Saunders et al., 2018; Rosenberg et al., 2018.</li> </ul> </li> <li> <p><strong>Mouse Cell Atlas (MCA)</strong>:</p> <ul> <li>Data focusing on 11 cell types from various organs: Han et al., 2018; The Tabula Muris Consortium, 2018.</li> </ul> </li> <li> <p><strong>Mouse Hematopoietic Stem and Progenitor Cells (MHSPC)</strong>:</p> <ul> <li>Data from SMART-seq2 and MARS-seq protocols: Nestorowa et al., 2016; Paul et al., 2015.</li> </ul> </li> <li> <p><strong>Mouse Retina (MR)</strong>:</p> <ul> <li>Data from two unassociated laboratories with 26,830 and 44,808 cells: Macosko et al., 2015; Shekhar et al., 2016.</li> </ul> </li> <li> <p><strong>PBMC (human Peripheral Blood Mononuclear Cell)</strong>:</p> <ul> <li>scRNA-seq data with two batches: Zheng et al., 2017.</li> </ul> </li> </ol> <p><strong>Usage Notes</strong>: Each dataset is provided in <code>.h5ad</code> format, compatible with common single-cell analysis tools such as Scanpy. Detailed metadata is included within each file.</p> <p><strong>Keywords</strong>: Single-cell RNA sequencing, scRNA-seq, Batch effect correction, Privacy-aware, Federated learning, scGen, FedscGen, Clinical multi-center studies, Genomics, Bioinformatics</p> <p><strong>Contact</strong>: For questions or further information, please contact Mohammad Bakhtiari at <a href="mailto:mohammad.bakhtiari@uni-hamburg.de.">mohammad.bakhtiari@uni-hamburg.de.</a></p> <p><strong>License</strong>: Creative Commons Attribution 4.0 International (CC BY 4.0)</p> </div> </div> </div> </div> <div> <div> <div> </div> </div> </div>
Disclosure Privacy Information Social Media on TikTok
<p><em><span>With the rapid development of technology, the use of social media by the public, especially among young people, is increasing. One of the social media platforms currently used by young people is the TikTok application. It is a video-based TikTok feature accompanied by music, writing, and pictures that are considered attractive, so teenagers like it to show their existence and self-disclosure. Therefore, this study aims to examine the intention of users to disclose their privacy. As the basis of the theory, this study deployed the privacy calculus theory, where the perceived benefits and perceived risks play crucial roles in the intention to disclose their privacy</span></em></p>
Supplemental Materials to Accompany: Abowd and Schmutte "An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices"
<p>Materials to supplement "<a href="https://www.aeaweb.org/articles?id=10.1257/aer.20170627">An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices</a>" by John Abowd and Ian Schmutte.</p>
A National Forum on Web Privacy and Web Analytics β Participant Survey Instrument
<p>This survey instrument was administered to participants of the <em>National Forum on Web Privacy and Web Analytics</em>. Results informed the Forum event and Forum deliverables.</p> <p>The <em>National Forum on Web Privacy and Web Analytics </em>was held September 2018 in Bozeman, Montana, where 40 librarians, technologists, and privacy researchers collaborated in producing a practical roadmap for enhancing our analytics practice in support of privacy.</p> <p>More information is available on our project site: <a href="https://osf.io/gnfpu/">https://osf.io/gnfpu/</a>. </p> <p>This project is made possible in part by the Institute of Museum and Library Services, through grant <a href="https://www.imls.gov/grants/awarded/lg-73-18-0100-18"># LG-73-18-0100-18</a>.</p>
Exploring Augmented Reality Privacy Icons for Smart Home Devices and their Effect on Users' Privacy Awareness
<p><strong>Exploring Augmented Reality Privacy Icons for Smart Home Devices and their Effect on Users' Privacy Awareness</strong></p> <p><strong>Authors</strong></p> <p>Kathrin Knutzen, Florian Weidner, Wolfgang Broll</p> <p> </p> <p><strong>About</strong></p> <p>This data represents the supplementary material for the conference paper with above title submitted at ISMAR 2021.</p> <p> </p> <p><strong>Contents</strong></p> <p>The supplementary material contains five files:</p> <ol> <li>The abstraction of paraphrases and transcripts after each condition respectively.<br> According to qualitative content analysis procedure, the conducted interviews were transcribed, paraphrased and subsequently abstracted to generate a category system. Every category is described with a definition and some exemplary quotes. Statements of participants are condensed and abstracted. Number of participants who made statements regarding a category, and most prominent valence are taken as basis to generate tree maps in Figure 5 and 6.<br> Please note that the prevalences represent the views or opinions of the participants on the single categories. Also, the mentioned categories have several subcategories and only the most important regarding privacy awareness are mentioned in the article.<br> <br> Transcripts, audio files and paraphrases are available upon request.</li> <li>The experimental task description. It served as exposition for the task that the participants had to complete.</li> <li>The interview guideline. Please note that this study was part of a larger project that also focused on topics such as usability and immersion, however, the article reports only on privacy awareness.</li> <li>The R script file to generate the tree maps in Figures 5 and 6. The dataset is created using data from the abstraction Excel sheet.</li> <li>A demonstration video of the experimental setup.</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.