Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

209

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

209 results for “Anonymity”

Learn how ShareScore rates datasets ↗
zenodo40/100

"Do you want to know who you are?" The rise of genetic ancestry testing and the search for genealogies: an anonymized survey from Sweden

<p>Full, anonymized survey data on genetic genealogy, ancestry and identity conducted by the Swedish Genealogical Society&nbsp;as part of a research project funded by the HERA joint research program &quot;Uses of the Past&quot;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Guaranteed Publisher/Subscriber Anonymity with Interest Profiles for Smart Cities Applications: Trends and Gaps

<p>Studies carried out in the context of anonymous Publisher/Subscriber (MQTT protocol) and with a profile of interest, are topics that the literature and the IoT (Internet of Things) industry face to implement models, architecture and architectural patterns, mainly related to anonymity, and the profile of interest there is nothing in literature or industry. In this way, the creation of an architecture based on these themes can facilitate the communication of IoT devices in smart cities. Therefore, this article aims to identify in the literature/industry how to guarantee the anonymity of Publisher/Subscriber with interest profiles using the MQTT protocol, and as future works to create an architecture based on the theme so that it can be used by academia and industry. IoT. The search string returned 291 (Two hundred and Ninety-One) works, of which 4 (Four) were selected according to the criteria of the Systematic Literature Review.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Crocosmia x crocosmiiflora (Lemoine ex Anonymous) N.E.Br. (BR0000012560592)

Belgium Herbarium image of <a href="https://www.plantentuinmeise.be">Meise Botanic Garden</a>.

opencc-by-sa-4.0May 2019View details →
zenodo40/100

Crocosmia x crocosmiiflora (Lemoine ex Anonymous) N.E.Br. (BR0000012406395)

Belgium Herbarium image of <a href="https://www.plantentuinmeise.be">Meise Botanic Garden</a>.

opencc-by-sa-4.0May 2019View details →
zenodo40/100

Crocosmia x crocosmiiflora (Lemoine ex Anonymous) N.E.Br. (BR0000025021974)

Belgium Herbarium image of <a href="https://www.plantentuinmeise.be">Meise Botanic Garden</a>.

opencc-by-sa-4.0May 2019View details →
zenodo40/100

Crocosmia x crocosmiiflora (Lemoine ex Anonymous) N.E.Br. (BR0000005159079)

Belgium Herbarium image of <a href="https://www.plantentuinmeise.be">Meise Botanic Garden</a>.

opencc-by-sa-4.0May 2019View details →
zenodo40/100

Crocosmia x crocosmiiflora (Lemoine ex Anonymous) N.E.Br. (BR0000014459825)

Belgium Herbarium image of <a href="https://www.plantentuinmeise.be">Meise Botanic Garden</a>.

opencc-by-sa-4.0May 2019View details →
zenodo40/100

UFRRJ Students Dataset - Anonymized

<p>Anonymized Dataset collected from 2000 until 2013 from the Federal Rural University of Rio de Janeiro students, containing 13 attributes.</p> <p>This dataset also contains mean GPA grades of each course.</p> <p>This dataset was used to build the experiments contained at :</p> <p>Ferreira, Raul S., and Carlos E. Mello. "W-SAGE: Ferramenta Web para Análise de Dados Geoespaciais." III Escola Regional de Sistemas de Informação (III ERSI-RJ)</p>

opencc-by-4.0Aug 2017View details →
zenodo40/100

Anonymized survey responses from two Inception Workshops

<p>To evaluate the effectiveness of the Inception Workshop approach we conducted a survey, in two experimental workshops. The survey was done in two stages, with pre- and post- workshop questionnaires. Both questionnaires included the same set of questions, in which the respondents were asked to evaluate their familiarity with environmental domain concepts and procedures, and IT tools and technologies. Anonymized responses are published here. The first workshop responses are in file `iw.q1.txt` and the second workshop in `iw.q2.txt`. Each row includes the question category (pre or post workshop), the question number (Qu1-Qu20), and the respondents answers in a 5-level Liker scale (from 1 to 5). A full list of questions are in the corresponding paper.</p>

opencc-by-sa-4.0Dec 2016View details →
zenodo40/100

Urban Agriculture and Health: Insights from Anonymized Expert Interviews on Spatial Planning in Greater Lomé, Togo

<p>Transcripts of anonymized interviews with 11 urban planning experts in Greater Lom&eacute; on the subject of urban agriculture, health and spatial planning.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Anonymized Responses to the Student Expectation of Learning Analytics Questionnaire (SELAQ) in 2022 and 2023

<p>Contains responses of 566 bachelor students to the Student Expectation of Learning Analytics Questionnaire (SELAQ) conducted at Bochum University of Applied Sciences in Germany in fall 2022 and summer 2023. The Questionnaire consists of 12 statements (in the dataset: enumerated from 1-12) which are evaluated by students regarding both their desire and their expectation (in the dataset: D and E). This results in 24 items (in the dataset: 1D, 1E, 2D, 2E, ..., 12D, 12E). Each item is evaluated by the students on a Likert-scale from 1 (strongly disagree) to 7 (strongly agree). Additionally, the students' study track (in the dataset: Study_Track) and their current semester (in the dataset: Semester) were also recorded. The questionnaires were carried out in paper form at the beginning of lectures. Statements 1, 2, 3, 5 and 6 deal with ethical and privacy expectations, while statements 4, 7, 8, 9, 10, 11 and 12 deal with service feature expectations regarding learning analytics. For a text version of the items, please refer to the paper related to this dataset (DOI 10.1145/3636555.3636923), the additional descriptions below or the attached description file.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Dataset: ABIVAX Société Anonyme (ABVX) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: ABIVAX Société Anonyme (ABVX) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Multilingual Speaker Anonymization Trials for CommonVoice and Multilingual LibriSpeech

<p>This dataset contains the speaker verification trial files of the evaluation data splits for Multilingual LibriSpeech (MLS) and CommonVoice (CV) that we propose in our paper "Probing the Feasibility of Multilingual Speaker Anonymization". <strong>The actual audio files are not included and have to be obtained separately</strong>, following the licenses of the respective corpus creators. The files in this dataset only contain the audio file IDs that can be used to prepare the data for the evaluation.&nbsp;</p> <h2>Data</h2> <p>All files contain the utterance IDs of the original MLS and CV corpora which makes it possible to align them to the audio files as provided by the dataset creators. If you use these datasets, you need to cite the original sources. We do not claim any rights to the audios.</p> <p>The dataset does not include all IDs of the corpora. For MLS, "dev" and "test" correspond to the dev and test splits as provided in the MLS corpus. For CV,&nbsp; the corpus was divided randomly into dev and test while ensuring no speaker overlap between both splits. We use the CV 16.1 version of the corpus and restrict it to validated audios where the user specified their gender as either female or male. Further information about the data restrictions can be found in the paper.</p> <p>Please note that we use the "client ID" of CV to distinguish between speakers. We acknowledge that this can lead to having the same speaker under different name multiple times in the corpus if they were assigned different client IDs. Based on our results for the original (non-anonymized) data, we believe this to have only little effect on our evaluation data.</p> <p>The following languages are included in this dataset: <strong>English (en), German (de), Dutch (nl), French (fr), Spanish (es), Italian (it), Portuguese (pt), Polish (pl), and Russian (ru).</strong></p> <h2>Structure</h2> <p>The directory contains separate folders for MLS and CV, with subfolders for each language. Each language subfolder contains 7 files:</p> <ul> <li>dev_enrolls</li> <li>dev_trials_f</li> <li>dev_trials_m</li> <li>test_enrolls</li> <li>test_trials_f</li> <li>test_trials_m</li> <li>utt2spk</li> </ul> <p>The file structures follow the evaluation data of the Voice Privacy Challenges (<a href="https://www.voiceprivacychallenge.org" rel="nofollow">https://www.voiceprivacychallenge.org</a>). The "enrolls" file contain the list of utterance IDs (i.e., audio files) that are used for enrollment of the speaker verification model. "trials_f" and "trials_m" correspond to the trial files for female and male speakers, respectively. Each line in a trial file consists of three constituents, separated by space: "enrollment speaker" "trial utterance" "target/nontarget". The last constituent signals whether the trial utterance was originally (i.e., before the anonymization) spoken by the enrollment speaker (target) or not (nontarget). The utt2spk file contains the true mapping between utterance and original speaker. This file is especially important for the CV corpus where we created new speaker names to replace the long client IDs, and where the speaker assignment is not visible in the file name.</p> <h2>Creation Process</h2> <p>For the preparation of the data into enrollment and trial subset, we tried to follow the dev and test files of the Voice Privacy Challenge 2022. This results in far more nontarget than target trials, and additional speakers in the trial set that are not contained in the enrollment set.</p> <h3>MLS</h3> <p>The MLS corpus (<a href="https://www.openslr.org/94/" rel="nofollow">available here</a>) already comes with a split into train / dev / test, which we reuse here. The data is further divided into 8 languages: en, de, nl, fr, es, it, pt and pl. We only take the dev and test sets for 6 languages (de, nl, fr, es, it, and pt). For en, we already have an alternative from the Voice Privacy Challenges based on the monolingual English LibriSpeech, so there is no need for another part from MLS. For pl, the MLS part only contains 2 speakers per gender and dev / test split, which is too small for effective speaker verification. <em>(Sidenote: Dutch is not significantly bigger with only 3 speakers per gender and split, but we decided to keep it anyway).</em> We do not further restrict the number of utterances or speaker in each language and dev / test set, which results in a large inbalance across languages. However, as given in the original MLS splits, the languages itself are balanced in terms of gender.</p> <h3>CV</h3> <p>Mozilla's CommonVoice data collection is significantly bigger than MLS, so we could select more speakers per language and make sure that the datasets per language were more or less balanced. We take the data from CV Version 16.1 (<a href="https://commonvoice.mozilla.org/en/datasets" rel="nofollow">available here</a>). Please note that users who <em>donated</em> their voice to the data collection can opt out of being included in the data at any point. <strong>This means that some speakers or utterances contained in our trial data might be missing in future downloads of the CV corpus.</strong> We further want to mention that we had to use the field <code>client ID</code> in the CV corpus for speaker assignment which is not fully accurate. The same speaker might end up with several client IDs if they are recording the utterances in different sessions. However, for the purpose of speaker anonymization, this issue is not as relevant as for pure speaker recognition.</p> <p>We use CV for all of our 9 languages. CV does not come with a division into train / dev / test splits, so we randomly sample speakers and utterances from it. For this sampling, we consider only speakers that have gender annotated as either female or male, and have recorded at least 50 validated audios. We further make sure that we have the same number of female as male speakers which leads for several languages to a large reduction in size, with most languages having far more male speakers in CV than female speakers. This leads especially to a smaller dataset for pl, for which only 14 speakers per gender and dev / test split are available. We randomly select at most 20 speakers per gender for each dev and test in each language, and randomly choose up to 70 utterances per speaker. As our lower bound for speaker selection was originally 50 utterances per speaker, this results in 50-70 utterances per speaker.</p> <h3>Separation into Enrollment and Trials</h3> <p>In MLS, we use all speakers for the enrollment and trial. The only exception is de, for which more speakers are available. In the MLS-de data, we reserve 5 speakers per gender and split for trial only, which creates some unseen distraction speakers in the trial data. In CV, we use 15 speakers per gender and split for enrollment (except for the smaller pl part, for which it is only 10), and also reserve up to 5 speakers per gender and split only for trial. Naturally, all enrollment speakers are also used in trial.</p> <p>15% of all utterances of a speaker (at least 5 utterances) are used as enrollment utterances, the rest for trial. All trial utterances are paired with each enrollment speaker of the respective gender. If an enrollment speaker is the actual speaker of that utterance, this is denoted as <em>target</em>, otherwise as <em>nontarget</em>.</p> <p>During the trials, the enrollment speaker is modeled as an average of the speaker embeddings of all enrollment utterances of that speaker.</p> <h2>Statistics</h2> <p>The following section displays the statistics for each dataset, language and dev / test split. In these statistics, female and male speakers are not distinguished, but the numbers are balanced for each subset.</p> <p>The following information is given for each dataset and language:</p> <ul> <li># speakers: number of speakers used for both enrollment and trial (50% female / 50% male)</li> <li># add.trial speakers: number of speakers additionally used only in trial</li> <li># enroll utts: total number of utterances used in enrollment (across all speakers)</li> <li># trial utts: total number of utterances used in trials (across all speakers)</li> <li># target trials: number of target trials (enrollment speaker == trial speaker)</li> <li># nontarget trials: number of nontarget trials (enrollment speaker != trial speaker)</li> <li># words: total number of words across all trial utterances (the WER is computed based on them)</li> <li># avg. utt length: average length of all utterances in the dataset, in seconds</li> </ul> <div> <h3>Development Data</h3> <p><strong>Total dataset statistics:</strong></p> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># speakers</th> <th># add. trial speakers</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> <th>avg. utt length</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>20</td> <td>10</td> <td>333</td> <td>3,136</td> <td>1,936</td> <td>29,424</td> <td>111,245</td> <td>14.90</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>18</td> <td>0</td> <td>354</td> <td>2,062</td> <td>2,062</td> <td>16,496</td> <td>73,007</td> <td>15.03</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>10</td> <td>0</td> <td>183</td> <td>1,065</td> <td>1,065</td> <td>4,260</td> <td>34,636</td> <td>14.88</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>20</td> <td>0</td> <td>349</td> <td>2,059</td> <td>2,059</td> <td>18,531</td> <td>74,782</td> <td>14.95</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>10</td> <td>0</td> <td>119</td> <td>707</td> <td>707</td> <td>2,828</td> <td>24,733</td> <td>15.90</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>6</td> <td>0</td> <td>461</td> <td>2,634</td> <td>2,634</td> <td>5,268</td> <td>11,0384</td> <td>14.83</td> </tr> <tr> <td>CV</td> <td>en</td> <td>30</td> <td>10</td> <td>279</td> <td>2,306</td> <td>1,691</td> <td>32,899</td> <td>22,394</td> <td>4.96</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>30</td> <td>9</td> <td>289</td> <td>2,396</td> <td>1,738</td> <td>34,202</td> <td>21,421</td> <td>5.16</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>30</td> <td>10</td> <td>293</td> <td>2,432</td> <td>1,761</td> <td>34,719</td> <td>23,240</td> <td>4.93</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>30</td> <td>10</td> <td>286</td> <td>2,421</td> <td>1,722</td> <td>34,593</td> <td>23,745</td> <td>5.69</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>30</td> <td>10</td> <td>290</td> <td>2,401</td> <td>1,735</td> <td>34,280</td> <td>22,569</td> <td>5.23</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>30</td> <td>10</td> <td>284</td> <td>2,398</td> <td>1,708</td> <td>34,262</td> <td>16,411</td> <td>3.97</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>30</td> <td>10</td> <td>288</td> <td>2,401</td> <td>1,725</td> <td>34,290</td> <td>22,108</td> <td>4.51</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>20</td> <td>8</td> <td>199</td> <td>1,739</td> <td>1,196</td> <td>16,194</td> <td>12,876</td> <td>4.39</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>30</td> <td>10</td> <td>289</td> <td>2,429</td> <td>1,737</td> <td>34,698</td> <td>20,599</td> <td>5.24</td> </tr> </tbody> </table> <div> <p><strong>Dataset statistics per speaker (average)</strong></p> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>16.6</td> <td>104.5</td> <td>96.8</td> <td>1,471.2</td> <td>3,553.0</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>19.7</td> <td>114.6</td> <td>114.6</td> <td>916.4</td> <td>4,350.0</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>18.3</td> <td>106.5</td> <td>106.5</td> <td>426.0</td> <td>4,405.5</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>17.4</td> <td>103.0</td> <td>103.0</td> <td>926.6</td> <td>4,300.5</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>11.9</td> <td>70.7</td> <td>70.7</td> <td>282.8</td> <td>2,084.0</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>76.8</td> <td>439.0</td> <td>439.0</td> <td>878.0</td> <td>45,515.0</td> </tr> <tr> <td>CV</td> <td>en</td> <td>9.3</td> <td>59.1</td> <td>56.4</td> <td>1,096.6</td> <td>627.0</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>9.6</td> <td>59.9</td> <td>57.9</td> <td>1,140.1</td> <td>492.5</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>9.8</td> <td>60.8</td> <td>58.7</td> <td>1,157.3</td> <td>606.0</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>9.5</td> <td>60.5</td> <td>57.4</td> <td>1,153.1</td> <td>594.5</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>9.7</td> <td>60.0</td> <td>57.8</td> <td>1,142.7</td> <td>494.0</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>9.5</td> <td>60.0</td> <td>56.9</td> <td>1,142.1</td> <td>308.0</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>9.6</td> <td>60.0</td> <td>57.5</td> <td>1,143.0</td> <td>520.5</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>10.0</td> <td>62.1</td> <td>59.8</td> <td>809.7</td> <td>514.5</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>9.6</td> <td>60.7</td> <td>57.9</td> <td>1,156.6</td> <td>502.0</td> </tr> </tbody> </table> <h3>Test Data</h3> <p><strong>Total dataset statistics:</strong></p> <p>&nbsp;</p> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># speakers</th> <th># add. trial speakers</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> <th>avg. utt length</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>30</td> <td>0</td> <td>329</td> <td>3,065</td> <td>1,906</td> <td>28,744</td> <td>110,202</td> <td>15.18</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>18</td> <td>0</td> <td>357</td> <td>2,069</td> <td>2,069</td> <td>16,552</td> <td>79,524</td> <td>14.94</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>10</td> <td>0</td> <td>185</td> <td>1,077</td> <td>1,077</td> <td>4,308</td> <td>34,796</td> <td>15.07</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>20</td> <td>0</td> <td>348</td> <td>2,037</td> <td>2,037</td> <td>18,333</td> <td>75,536</td> <td>15.11</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>10</td> <td>0</td> <td>125</td> <td>746</td> <td>746</td> <td>2,984</td> <td>26,769</td> <td>15.47</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>6</td> <td>0</td> <td>458</td> <td>2,617</td> <td>2,617</td> <td>5,234</td> <td>108,489</td> <td>14.96</td> </tr> <tr> <td>CV</td> <td>en</td> <td>30</td> <td>9</td> <td>289</td> <td>2,344</td> <td>1,733</td> <td>33,427</td> <td>22,560</td> <td>5.09</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>30</td> <td>10</td> <td>284</td> <td>2,377</td> <td>1,713</td> <td>33,942</td> <td>21,492</td> <td>5.24</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>30</td> <td>10</td> <td>291</td> <td>2,408</td> <td>1,745</td> <td>34,275</td> <td>22,887</td> <td>4.74</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>30</td> <td>10</td> <td>298</td> <td>2,444</td> <td>1,786</td> <td>34,874</td> <td>24,091</td> <td>5.37</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>30</td> <td>10</td> <td>284</td> <td>2,377</td> <td>1,703</td> <td>33,952</td> <td>22,655</td> <td>5.25</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>30</td> <td>7</td> <td>290</td> <td>2,213</td> <td>1,743</td> <td>31,452</td> <td>15,727</td> <td>4.23</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>30</td> <td>10</td> <td>283</td> <td>2,254</td> <td>1,704</td> <td>32,106</td> <td>19,913</td> <td>4.32</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>20</td> <td>8</td> <td>196</td> <td>1,712</td> <td>1,181</td> <td>15,939</td> <td>13,809</td> <td>4.82</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>30</td> <td>10</td> <td>294</td> <td>2,435</td> <td>1,759</td> <td>34,766</td> <td>20,509</td> <td>5.13</td> </tr> </tbody> </table> <div> <h5>Dataset statistics per speaker (average)</h5> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>16.4</td> <td>102.2</td> <td>95.3</td> <td>1,437.2</td> <td>4,047.0</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>19.8</td> <td>114.9</td> <td>114.9</td> <td>919.6</td> <td>3,945.5</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>18.5</td> <td>107.7</td> <td>107.7</td> <td>430.8</td> <td>3,658.0</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>17.4</td> <td>101.8</td> <td>101.8</td> <td>916.6</td> <td>4,870.5</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>12.5</td> <td>74.6</td> <td>74.6</td> <td>298.4</td> <td>3,141.5</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>76.3</td> <td>436.2</td> <td>436.2</td> <td>872.3</td> <td>24,892.5</td> </tr> <tr> <td>CV</td> <td>en</td> <td>9.6</td> <td>60.1</td> <td>57.8</td> <td>1,114.2</td> <td>551.5</td> </tr> <tr> <td>&nbsp;</td> <td>de</td> <td>9.5</td> <td>59.4</td> <td>57.1</td> <td>1,131.4</td> <td>568.5</td> </tr> <tr> <td>&nbsp;</td> <td>fr</td> <td>9.7</td> <td>60.2</td> <td>58.2</td> <td>1,145.8</td> <td>618.0</td> </tr> <tr> <td>&nbsp;</td> <td>it</td> <td>9.9</td> <td>61.1</td> <td>59.5</td> <td>1,162.5</td> <td>612.0</td> </tr> <tr> <td>&nbsp;</td> <td>es</td> <td>9.5</td> <td>59.4</td> <td>56.8</td> <td>1,131.7</td> <td>610.0</td> </tr> <tr> <td>&nbsp;</td> <td>pt</td> <td>9.7</td> <td>59.7</td> <td>58.1</td> <td>1,048.4</td> <td>409.5</td> </tr> <tr> <td>&nbsp;</td> <td>nl</td> <td>9.4</td> <td>59.3</td> <td>56.8</td> <td>1,070.2</td> <td>564.0</td> </tr> <tr> <td>&nbsp;</td> <td>pl</td> <td>9.8</td> <td>61.1</td> <td>59.0</td> <td>797.0</td> <td>515.0</td> </tr> <tr> <td>&nbsp;</td> <td>ru</td> <td>9.8</td> <td>60.9</td> <td>58.6</td> <td>1,158.9</td> <td>392.5</td> </tr> </tbody> </table> <h2>More Information</h2> <h3>Paper</h3> <p>The paper in which this dataset is proposed will be published at Interspeech 2024.</p> <p><a href="https://arxiv.org/abs/2407.02937">The preprint is available on arXiv</a>: Meyer, Sarina, Florian Lux, and Ngoc Thang Vu. "Probing the Feasibility of Multilingual Speaker Anonymization." <em>arXiv preprint arXiv:2407.02937</em> (2024).&nbsp;</p> <h3>Code</h3> <p>All code related to this data, including the data and descriptions, as well as preparation scripts to use this data for speaker anonymization, can be found in our Github repository:<a href="https://github.com/DigitalPhonetics/speaker-anonymization/tree/multilingual"> https://github.com/DigitalPhonetics/speaker-anonymization/tree/multilingual</a></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Negev Walking Focusing Interviews: Original Transcripts (Anonymous)

<p>The attached files are 30 interview transcripts for the article: &quot;In Search for the Authentic Nature Experience: Walking-Focusing Interviews as a Tool for Evaluating Cultural Ecosystem Services&quot;, written for People and Nature.&nbsp;</p> <p>The material is available for download as individual files (one for each interview) or as a ZIP file (.rar) if you wish to download all.</p> <p>The interviews took place in the Negev Desert, Israel, on a trail named &quot;Bor Hemet&quot;, which is part of an official protected area, in October-December 2018 by the authors. Their analysis&nbsp;intends to provide an indication to&nbsp;whether this form of interviews - Walking-Focusing interviews - can provide meaningful qualitative and holistic information pertaining to Cultural Ecosystem Services that other methodologies have difficulties providing.&nbsp;</p> <p>&nbsp;</p> <p>For more details on the study, see the following poster:&nbsp;&nbsp;</p> <p>https://www.researchgate.net/publication/325924203_Walking_Focusing_and_Evaluating_the_Cultural_Ecosystem_Services_of_Drylands</p> <p>&nbsp;</p> <p>And the People and Nature website (open access journal):</p> <p>&nbsp;</p> <p>https://www.britishecologicalsociety.org/publications/journals/people-and-nature/</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

What makes research software sustainable? Anonymized interview transcripts

<p>Anonymized transcripts of&nbsp;a series of interviews with the developers of research software.</p> <p>We also include the Participant Information Sheet provided to all participants before the interview commences.</p>

opencc-by-sa-4.0Feb 2019View details →
zenodo40/100

The PInSoRo dataset -- anonymized version

<p>The <strong>PInSoRo Dataset</strong> (also called <em>Freeplay Sandbox dataset</em>) is a large (120 children, 45h+ of RGB-D video recordings), open-data dataset of child-child and child-robot social interactions.</p> <p>These interactions are recorded during little-constrained <strong>free play</strong> episodes. They encompass a rich and diverse set of social behaviours.</p> <p>More on the dataset website: <a href="https://freeplay-sandbox.github.io/">https://freeplay-sandbox.github.io/</a></p> <p>Documentation of the dataset structure: <a href="https://github.com/freeplay-sandbox/dataset/blob/master/data/README.md#pinsoro-dataset---data-structure">online</a> or check data/README.md inside the dataset zip file.</p> <p>You can also check the <a href="https://github.com/freeplay-sandbox/dataset/blob/master/CHANGELOG.md">CHANGELOG</a>.</p>

opencc-by-sa-4.0Nov 2017View details →
zenodo40/100

FT40 Anonymous (11) 4-key fagottino: measurements, photos, endoscopic video

<p>Dataset of FT 40 Anonymous (11) 4-key fagottino&nbsp;containing detailed external and internal measurements, photos, and an endoscopic video.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

FT46 Anonymous (12) 5-key tenoroon: measurements, photos, endoscopic video

<p>Dataset of FT46 Anonymous (12) 5-key tenoroon containing&nbsp;detailed external and internal measurements, photos, and an endoscopic video.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [Eng: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (anonymized version - first part)

<p><a title="GENEActiv" href="https://activinsights.com/technology/geneactiv/" target="_blank" rel="noopener">GENEActiv</a> accelerometer .csv files converted with a 1 second epoch from raw GENEActiv .bin files recorded during the project entitled "<strong>Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXI&egrave;me si&egrave;cle: exemple de la Nouvelle-Cal&eacute;donie</strong>" [en: "<strong>Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia</strong>"]. Devices are 60-Hz triaxial accelerometers.</p> <p>This dataset also contains <strong>participantCharacteristics.csv</strong> that povides basic information about participants and <strong>read_a_binFile_share.R</strong> that is a short R code aiming at converting and saving accelerometer data from .bin files in 1 second epoch .csv files (consider the Methods section).</p> <p>Participant characteristics: 10 to 16 years old students and some parents.</p> <p>Number of participants: 231 (206 adolescents + 25 adults).</p> <p>Year of the study: 2018 - 2019.</p> <p>Place of the study: New Caledonia.</p> <p>The accelerometer .csv files with a 1 second epoch and extracted from raw .bin files are available in open datasets:</p> <ul> <li><a title="Open dataset - first part" href="https://doi.org/10.5281/zenodo.12615468" target="_blank" rel="noopener">anonymized version - first part</a></li> <li><a title="Open dataset - second part" href="https://doi.org/10.5281/zenodo.12638746" target="_blank" rel="noopener">anonymized version - second part</a></li> <li><a title="Open dataset - third part" href="https://doi.org/10.5281/zenodo.12682660" target="_blank" rel="noopener">anonymized version - third part</a></li> </ul> <p>The accelerometer raw .bin files are available in&nbsp;<strong>restricted datasets</strong>:</p> <ul> <li><a title="Restricted dataset - first part" href="https://doi.org/10.5281/zenodo.11594645" target="_blank" rel="noopener">non-anonymized version - first part</a></li> <li><a title="Restricted dataset - second part" href="https://doi.org/10.5281/zenodo.12638965" target="_blank" rel="noopener">non-anonymized version - second part</a></li> <li><a title="Restricted dataset - third part" href="https://doi.org/10.5281/zenodo.12661429" target="_blank" rel="noopener">non-anonymized version - third part</a></li> </ul> <p>Other participant characteristics (age, place of living, cultural community and socio-economic status) are available in a <a title="Information associated with GENEActiv accelerometer files collected during the project entitled &quot;Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXI&egrave;me si&egrave;cle: exemple de la Nouvelle-Cal&eacute;donie&quot; [en: &quot;Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia&quot;] (non-anonymized information version)" href="https://doi.org/10.5281/zenodo.12195186" target="_blank" rel="noopener">restricted non-anonymized dataset</a>.</p> <p>When using this dataset, please cite the following reference:<br><a title="Wattelez et al. 2025" href="https://doi.org/10.1016/j.dib.2024.111228" target="_blank" rel="noopener">G. Wattelez, S. Frayon, O. Galy, Assessing physical activity/behavior of adolescents living in the Pacific with accelerometer data: 231 GENEActiv records in New Caledonia, Data in Brief 58 (2025) 111228, doi: 10.1016/j.dib.2024.111228</a></p>

embargoedcc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record