AOL Dataset for Browsing History and Topics of Interest
<p><strong>AOL Dataset for Browsing History and Topics of Interest</strong></p>
<p>This record provides the datasets of the paper <em>The Privacy-Utility Trade-off in the Topics API</em> (DOI: <a href="https://doi.org/10.1145/3658644.3670368" target="_blank" rel="noopener">10.1145/3658644.3670368</a>; arXiv: <a href="https://arxiv.org/abs/2406.15309" target="_blank" rel="noopener">2406.15309</a>).</p>
<p>The datasets generating code and the experimental results can be found in <a href="https://doi.org/10.5281/zenodo.11229402" target="_blank" rel="noopener">10.5281/zenodo.11229402</a> (<a href="https://github.com/nunesgh/topics-api-analysis" target="_blank" rel="noopener">github.com/nunesgh/topics-api-analysis</a>).</p>
<p><strong>Files</strong></p>
<ol>
<li><code>AOL-treated.csv</code>: This dataset can be used for analyses of browsing history vulnerability and utility, as enabled by third-party cookies. It contains singletons (individuals with only one domain in their browsing histories) and one outlier (one user with 150.802 domain visits in three months) that are dropped in some analyses.</li>
<li><code>AOL-treated-unique-domains.csv</code>: Auxiliary dataset containing all the unique domains from <code>AOL-treated.csv</code>.</li>
<li><code>Citizen-Lab-Classification.csv</code>: Auxiliary dataset containing the Citizen Lab Classification data, as of commit <a href="https://github.com/citizenlab/test-lists/tree/ebd0ee8d41977b381972b2f6c471af5437d8d015/lists" target="_blank" rel="noopener">ebd0ee8</a>, treated for inconsistencies and filtered according to Mozilla's Public Suffix List, as of commit <a href="https://github.com/publicsuffix/list/tree/5e6ac3a082505ac4cf08858bdb38382d9a912833" target="_blank" rel="noopener">5e6ac3a</a>, extended by the discontinued TLDs: .bg.ac.yu, .ac.yu, .cg.yu, .co.yu, .edu.yu, .gov.yu, .net.yu, .org.yu, .yu, .or.tp, .tp, and .an.</li>
<li><code>AOL-treated-Citizen-Lab-Classification-domain-match.csv</code>: Auxiliary dataset containing domains matched from <code>AOL-treated-unique-domains.csv</code> with domains and respective topics from <code>Citizen-Lab-Classification.csv</code>.</li>
<li><code>Google-Topics-Classification-v1.txt</code>: Auxiliary dataset containing the Google Topics API taxonomy v1 data as provided by Google with the Chrome browser.</li>
<li><code>AOL-treated-Google-Topics-Classification-v1-domain-match.csv</code>: Auxiliary dataset containing domains matched from <code>AOL-treated-unique-domains.csv</code> with domains and respective topics from <code>Google-Topics-Classification-v1.txt</code>.</li>
<li><code>AOL-reduced-Citizen-Lab-Classification.csv</code>: This dataset can be used for analyses of browsing history vulnerability and utility, as enabled by third-party cookies, and for analyses of topics of interest vulnerability and utility, as enabled by the Topics API. It contains singletons and the outlier that are dropped in some analyses.<br>This dataset can be used for analyses including the (data-dependent) randomness of trimming-down or filling-up the top-s sets of topics for each individual so each set has s topics. Privacy results for Generalization and utility results for Generalization, Bounded Noise, and Differential Privacy are expected to slightly vary with each run of the analyses over this dataset.</li>
<li><code>AOL-reduced-Google-Topics-Classification-v1.csv</code>: This dataset can be used for analyses of browsing history vulnerability and utility, as enabled by third-party cookies, and for analyses of topics of interest vulnerability and utility, as enabled by the Topics API. It contains singletons and the outlier that are dropped in some analyses.<br>This dataset can be used for analyses including the (data-dependent) randomness of trimming-down or filling-up the top-s sets of topics for each individual so each set has s topics. Privacy results for Generalization and utility results for Generalization, Bounded Noise, and Differential Privacy are expected to slightly vary with each run of the analyses over this dataset.</li>
<li><code>AOL-experimental.csv</code>: This dataset can be used to empirically verify code correctness for <a href="https://doi.org/10.5281/zenodo.11229402" target="_blank" rel="noopener">10.5281/zenodo.11229402</a>. All privacy and utility results are expected to remain the same with each run of the analyses over this dataset.</li>
<li><code>AOL-experimental-Citizen-Lab-Classification.csv</code>: This dataset can be used to empirically verify code correctness for <a href="https://doi.org/10.5281/zenodo.11229402" target="_blank" rel="noopener">10.5281/zenodo.11229402</a>. All privacy and utility results are expected to remain the same with each run of the analyses over this dataset.</li>
<li><code>AOL-experimental-Google-Topics-Classification-v1.csv</code>: This dataset can be used to empirically verify code correctness for <a href="https://doi.org/10.5281/zenodo.11229402" target="_blank" rel="noopener">10.5281/zenodo.11229402</a>. All privacy and utility results are expected to remain the same with each run of the analyses over this dataset.</li>
</ol>
<p><strong>License</strong></p>
<p><a href="https://creativecommons.org/licenses/by-nc-sa/4.0/" target="_blank" rel="noopener">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International</a>.</p>
openJun 2024View details →