Skip to main content
zenodorestricted

A dataset of Data Subject Access Request Packages

<h3>Overview</h3> <p>This dataset is a minimal example of Data Subject Access Request Packages (SARPs), as they can be retrieved under data protection laws, specifically the GDPR. It includes data from two data subjects, each with accounts for five major sevices, namely Amazon, Apple, Facebook, Google, and Linkedin.</p> <p>&nbsp;</p> <h3>Purpose and Usage</h3> <p>This dataset is meant to be an initial dataset that allows for manual exploration of structures and contents found in SARPs. Hence, the number of controllers and user profiles should be minimal but sufficient to allow cross-subject and cross-controller analysis. This dataset can be used to explore structures, formats and data types found in real-world SARPs. Thereby, the planning of future SARP-based research projects and studies shall be facilitated.<br><br>We invite other researchers to use this dataset to explore the structure of SARPs. The envisioned primary usage includes the development of user-centric privacy interfaces and other technical contributions in the area of data access rights. Moreover, these packages can also be used for examplified data analyses, although no substantive research questions can be answered using this data. In particular, this data does not reflect how data subjects behave in real world. However, it is representative enough to give a first impression on the types of data analysis possible when using real world data.&nbsp;</p> <h3>&nbsp;</h3> <h3>Data Generation&nbsp;</h3> <p>In order to allow cross-subject analysis, while keeping the re-identification risk minimal, we used research-only accounts for the data generation. A detailed explanation of the data generation method can be found in the paper corresponding to the dataset, accepted for the Annual Privacy Forum 2024.</p> <p>In short, two user profiles were designed and corresponding accounts were created for each of the five services. Then, those accounts were used for two to four month. During the usage period, we minimized the amount of identifying data and also avoided interactions with data subjects not part of this research. Afterwards, we performed a data access request via the controller's web interface. Finally, the data was cleansed as described in detail in the acconpanying paper and in brief within the following section.</p> <h3>&nbsp;</h3> <h3>Data Cleansing</h3> <p>Before publication, both possibly identifying information and security relevant attributes need to be obfuscated or deleted. Moreover, multi-party data (especially messages with external entities) must be deleted. If data is obfuscated, we made sure to substitute multiple occurances of the same information with the same replacement.<br>We provide a list of deleted and obfuscated items, the obfuscation scheme and, if applicable, the replacement.</p> <p>The list of obfuscated items looks like the following example:</p> <table> <tbody> <tr> <td>path</td> <td>filetype</td> <td>filename</td> <td>attribute</td> <td>scheme</td> <td>replacement</td> </tr> <tr> <td>linkedin\Linkedin_Basic</td> <td>csv</td> <td>messages.csv</td> <td>TO</td> <td>semantic description</td> <td>Firstname Lastname</td> </tr> <tr> <td>gooogle\Meine Aktivit&auml;ten\Datenexport</td> <td>html</td> <td>MeineAktivit&auml;ten.html</td> <td>IP Address</td> <td>loopback</td> <td>127.142.201.194</td> </tr> <tr> <td>facebook\personal_information</td> <td>json</td> <td>profile_information.json</td> <td>emails</td> <td>semantic description</td> <td>firstname.lastname@gmail.com</td> </tr> </tbody> </table> <h3>&nbsp;</h3> <h3>Data Characterization</h3> <p>To give you an overview of the dataset, we publicly provide some meta-data about the usage time and SARP characteristics of exports from subject A/ subject B.</p> <table> <tbody> <tr> <td>provider</td> <td>usage time<br>(in month)</td> <td>export options</td> <td>file types</td> <td># subfolders</td> <td># files</td> <td>export size</td> </tr> <tr> <td>Amazon</td> <td>2/4</td> <td>all categories</td> <td>CSV (32/49)<br>EML (2/5)<br>JPEG (1/2)<br>JSON (3/3)<br>PDF (9/10)<br>TXT (4/4)</td> <td>41/49</td> <td>51/73</td> <td>1.2 MB / 1.4 MB</td> </tr> <tr> <td>Apple</td> <td>2/4</td> <td>all data<br>max. 1 GB/ max. 4 GB</td> <td>CSV (8/3)</td> <td>20/1</td> <td>8/3</td> <td>71.8 KB / 294.8 KB</td> </tr> <tr> <td>Facebook</td> <td>2/4</td> <td> <p>all data</p> <p>JSON/HTML</p> <p>on my computer</p> </td> <td>JSON (39/0)<br>HTML (0/63)<br>TXT (29/28)<br>JPG (0/4)<br>PNG (1/15)<br>GIF (7/7)</td> <td>45/76</td> <td>76/117</td> <td>12.3 MB / 13.5 MB</td> </tr> <tr> <td>Google</td> <td>2/4</td> <td> <p>all data</p> <p>frequency once</p> <p>ZIP</p> <p>max. 4 GB</p> </td> <td>HTML (8/11)<br>CSV (10/13)<br>JSON (27/28)<br>TXT (14/14)<br>PDF (1/1)<br>MBOX (1/1)<br>VCF (1/0)<br>ICS (1/0)<br>README (1/1)<br>JPG (0/2)</td> <td>44/51</td> <td>64/71</td> <td>1.54 MB /1.2 MB</td> </tr> <tr> <td>LinkedIn</td> <td>2/4</td> <td>all data</td> <td>CSV (18/21)</td> <td>0/0 (part 1/2)<br>0/0 (part 1/2)</td> <td>13/18<br>19/21</td> <td> <p>3.9 KB / 6.0 KB</p> <p>6.2 KB / 9.2 KB</p> </td> </tr> </tbody> </table> <h3><br>Authors</h3> <p>This data collection was performed by Daniela P&ouml;hn (Universit&auml;t der Bundeswehr M&uuml;nchen, Germany), Frank Pallas and Nicola Leschke (Paris Lodron Universit&auml;t Salzburg, Austria). For questions, please contact nicola.leschke@plus.ac.at.</p> <h3>Accompanying Paper</h3> <p>The dataset was collected according to the method presented in:<br>Leschke, P&ouml;hn, and Pallas (2024). "How to Drill Into Silos: Creating a Free-to-Use Dataset of Data Subject Access Packages". Accepted for Annual Privacy Forum 2024.</p>

ShareScore

16/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
8
Reuse readiness
0
Engagement
0