Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
209
datasets available to search
ShareScore release 0.9.0
Dataset results
209 results for “anonymization”
WP2 Task 2.2 results of desktop study of TNs (anonymized)
<p>WP2 Task 2.2 results of desktop study of TNs (personal info removed)</p>
CERN Anonymized Mattermost Data
<p>The information generated in large organizations has been a great source of information for researchers, facilitating innovation and advances in various areas. Promoting an open and transparent approach has provided diverse benefits and competitive advantages for organizations. This dataset contains anonymized data from the CERN Mattermost service. The selected entities are Channel Information, Channel Member, Channel Member History, Teams, Team Members, and User.</p> <p>Document "<strong>CERN_Mattermost_OpenData_Tables.pdf</strong>" contains the data descriptions of data entities available from Mattermost before risk assessment and anonymization.</p> <p>Document "<strong>CERN_Mattermost_OpenData_Tables-Anonymus.pdf</strong>", contains data descriptions of data entities published within this dataset found in the "matermost.json" file.</p> <p>The dataset is released under the "CC BY-NC Creative Commons Attribution Non-Commercial Licence", to prevent the commercial application of the dataset, since the intent is to use the dataset for research purposes.</p> <p>The table below provides a brief overview of the data contained in the dataset:</p> <table> <thead> <tr> <th scope="col">Description</th> <th scope="col">Value</th> </tr> </thead> <tbody> <tr> <td>Number of Users </td> <td>21231</td> </tr> <tr> <td>Number of Teams</td> <td>2367</td> </tr> <tr> <td>Number of Buildings</td> <td>151</td> </tr> <tr> <td>Number of Organisational Units</td> <td>163</td> </tr> <tr> <td>Number of Channels</td> <td>12773</td> </tr> </tbody> </table> <p>A small framework to work with the CERN Anonymized Mattermost Data Set was created and can be reached via the following link:</p> <ul> <li><a href="https://github.com/mpobaschnig/cdhf">CERN Data Handling Framework</a> (DOI: 10.5281/zenodo.6572538)</li> </ul> <p><strong>Frequently Asked Questions</strong></p> <p><strong><em>What does it mean if a user does not have an Organisational Id and/or Building Id?</em></strong></p> <ul> <li>This means that the User is an external user, for example a Guest Lecturer which has access to CERN Mattermost but is not part of an Organisational Unit or Building.</li> </ul> <p> </p>
anonymized author dataset from the publication "Impact of the COVID-19 pandemic on publishing in astronomy in the initial two years"
<p>We provide the anonymized author dataset that was prepared for the research presented in <a href="https://arxiv.org/pdf/2203.15621.pdf">https://arxiv.org/pdf/2203.15621.pdf</a> . The data is provided as two pickled pandas data frames. The first data frame contains the number of publications each author has written in each year per author position (as 1st author etc.) and their assigned gender. The second file contains the corresponding country of affiliation for each author and year. The two data frames can be matched by author_id.</p> <p>Columns in "matched_author_dataset_publication_counts_zenodo.pkl":</p> <pre>['author_id', 'gender', 'P(gender)', 'pub1_tot', 'first_auth_pub_year', 'pub2_tot', 'auth_2_pub_year', 'pub3_tot', 'auth_3_pub_year', 'pub4_tot', 'auth_4_pub_year', 'pub5_tot', 'auth_5_pub_year', 'pub6_tot', 'auth_6_pub_year', 'pub7_tot', 'auth_7_pub_year', 'pub8_tot', 'auth_8_pub_year', 'pub9_tot', 'auth_9_pub_year', 'pub10_tot', 'auth_10_pub_year', 'pub11_tot', 'auth_11_pub_year', 'pub12_tot', 'auth_12_pub_year', 'pub13_tot', 'auth_13_pub_year', 'pub14_tot', 'auth_14_pub_year', 'pub15_tot', 'auth_15_pub_year', 'pub16_tot', 'auth_16_pub_year', 'tot_pub', 'tot_pub_year', 'last_pub', 'first_pub']</pre> <p>'author_id' is a unique id assigned to the author. 'gender' contains the author's most likely gender and 'P('gender')' the likelihood of correct assignment. 'pubX_tot' contains the total number of publications the author has written as Xth author, 'auth_X_pub_year' contains a list with one entry per year from 1950 to 2022 counting the number of papers the author has published as Xth author in that year. 'tot_pub' counts the total number of publications (summed over all years) and 'tot_pub_year' for each year/ 'last_pub' and 'first pub' contain the list indices of the years when the author last/first published.</p> <p>Columns in author_dataframe_country.pkl:</p> <pre>['author_id', 'aff_year', 'aff_country_author']</pre> <p>'aff_year' is a list of years for which country information could be inferred for this author. 'aff_country_author' is a list of countries for each entry in 'aff_year'.</p> <p> </p>
Head of the Buddha, Anonymous, Rijksmuseum
" h 24.0cm × w 15.0cm × d 11.8cm. More details The protuberance or topknot on the head (ushnisha) and the raised spot on the forehead (urna) are standard features of the Buddha. The wavy hair combed back from the face is characteristic of Buddha figures from Gandhara. It was only during the early decades of the 1st century AD that the first Buddha images began to appear." - https://www.rijksmuseum.nl/en/search/objecten?q=head+buddha&p=1&ps=12&ii=1#/AK-MAK-1231,1 Source: Objaverse 1.0 / Sketchfab
Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES) Dataset - Anonymized
<p>Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES) and this corresponding dataset aim to provide tools for measuring user enjoyment from an external perspective to supplement self-reported user enjoyment responses in human-robot interaction research, with future potential application for autonomous detection of user enjoyment in real-time in robots and agents for adapting conversations contingently to provide enjoyable and long-lasting interactions.</p> <p>The dataset consists of 25 older adults' (12 men, 13 women) open-domain dialogue with an autonomous companion robot with an integrated large language model (GPT-3.5, text-davinci-003) from participatory design workshops conducted in March 2023. The conversations are annotated for user enjoyment based on HRI CUES by 3 expert annotators, as described in the paper (arXiv:2405.01354). Robot architecture and participatory design workshops are described in DOI: 10.21203/rs.3.rs-2884789/v1.</p> <p><strong>Exchanges</strong> file contains the participant ID, the number of the turn (conversation exchange by Robot-Participant response), the start and end of the turn, the anonymized transcript for the turn, and three annotator scores for the user enjoyment in the exchange. </p> <p><strong>Overall </strong>file contains the participant ID, self-reported user perception scores from the questionnaire ("I was satisfied with my conversation with the robot", "It was fun talking to the robot", "The conversation with the robot was interesting", "It felt strange talking to the robot") and three annotator scores for the user enjoyment in the overall interaction.</p> <p>The conversations are in Swedish. Participants' mean age is 74.6 (SD=5.8). 20 participants had no prior interaction with a robot, and only one had previously talked with a robot. The average interaction duration is 7.4 min (SD=1.5) with 12 to 29 turns. Each turn lasts 5 to 61 seconds (M=17.7, SD=7.2). The total duration of the interactions is 174 min, corresponding to 590 turns. </p> <p><em>Videos of the interactions are available upon request, contingent upon a signed agreement to maintain data confidentiality in accordance with GDPR regulations.</em></p> <p>Anonymization macros:</p> <p>[P_NAME]: Participant's name (may include surname). The robot always uses the first name even when the surname is given.</p> <p>[NAME_REMOVED]: A name of another person mentioned by the participant.</p> <p>[LOCATION_REMOVED]: Small town/village/area where the participant lives or lived.</p> <p>[MEDICAL_INFO_REMOVED]: Medical information shared by the participant.</p> <p>[AGE_REMOVED]: Participant's or other person's age.</p> <p>[INFORMATION_REMOVED]: Sensitive information shared by the participant.</p> <p>[MISTAKEN_NAME]: Speech recognition error resulted in the name being misunderstood.</p>
Serum aminoacid levels in patients with spinocelular cancer of head and neck: anonymized dataset
<p><strong>Dataset containing analysis of free serum amino acid concentrations in patients with head and neck spinocellular tumors.</strong></p> <p>The study was conducted in accord with the Helsinki Declaration of 1964 and all subsequent revisions thereof. It was approved by the ethical committee of St. Anne’s Faculty Hospital, Brno, and by ethical committee of University Hospital Motol, Prague, Czech Republic.</p> <p>All blood samples were obtained from HNSCC patients who developed histologically verified primary HNSCC after they signed the informed consent. Blood samples were obtained by venipuncture. The blood samples were centrifuged at 3000 rpm at 4°C for 10 min within 60 min after collection. Serum was aliquoted and stored at −80°C until analysis.</p> <p>The inclusion criteria were: age 40-95 years; no prior chemotherapy; no endocrinologic or metabolic disorders; no uncontrolled hypertension or infections; normal liver, heart and kidney function; and adequate bone marrow reserve.</p> <p>Amino acid profiles were examined using ion-exchange liquid chromatography (AAA-400, Ingos, Prague, Czech Republic) with post-column derivatization by ninhydrin and absorbance detector in visible light range (IEC-Vis).</p>
Information associated with GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [Eng: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (non-anonymized information)
<p>Additional restricted participant characteristics of the following datasets:</p> <ul> <li><a title="GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" (non-anonymized version - first part)" href="https://doi.org/10.5281/zenodo.11594645" target="_blank" rel="noopener">GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [en: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (non-anonymized version - first part)</a></li> <li><a title="GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" (non-anonymized version - second part)" href="https://doi.org/10.5281/zenodo.12638965" target="_blank" rel="noopener">GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [en: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (non-anonymized version - second part)</a></li> <li><a title="GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" (non-anonymized version - third part)" href="https://doi.org/10.5281/zenodo.12661429" target="_blank" rel="noopener">GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [en: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (non-anonymized version - third part)</a></li> <li><a title="GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" (anonymized version - first part)" href="https://doi.org/10.5281/zenodo.12615468" target="_blank" rel="noopener">GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [en: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (anonymized version - first part)</a></li> <li><a title="GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" (anonymized version - second part)" href="https://doi.org/10.5281/zenodo.12638746" target="_blank" rel="noopener">GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [en: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (anonymized version - second part)</a></li> <li><a title="GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" (anonymized version - third part)" href="https://doi.org/10.5281/zenodo.12682660" target="_blank" rel="noopener">GENEActiv accelerometer files collected during the project entitled "Cultures et comportements alimentaires de la jeunesse dans les pays francophones du Pacifique au XXIème siècle: exemple de la Nouvelle-Calédonie" [en: "Eating cultures and behaviors of young people in French-speaking Pacific countries in the 21st century: the example of New Caledonia"] (anonymized version - third part)</a></li> </ul> <p>Participant characteristics: 10 to 16 years old students (n = 206) and parents (n = 25).</p> <p>Number of participants: 231.</p> <p>Year of the study: 2018 - 2019.</p> <p>Place of the study: New Caledonia.</p> <p>When using this dataset, please cite the following reference:<br><a title="Wattelez et al. 2025" href="https://doi.org/10.1016/j.dib.2024.111228" target="_blank" rel="noopener">G. Wattelez, S. Frayon, O. Galy, Assessing physical activity/behavior of adolescents living in the Pacific with accelerometer data: 231 GENEActiv records in New Caledonia, Data in Brief 58 (2025) 111228, doi: 10.1016/j.dib.2024.111228</a></p>
Anonymous Submission
<p>The dataset provided in this repository encompasses extensive code quality metrics and description completeness measures for each pull request (PR) analyzed in the study.</p> <p><strong>Code Quality Metrics:</strong> The dataset includes 15 code quality metrics for each PR, obtained using SonarQube scans of the changed source code files within those PRs. These metrics provide insights into various aspects of code quality including but not limited to cognitive complexity, maintainability, and technical debt.</p> <p><strong>Description Completeness:</strong> Each PR's description completeness was assessed using the Llama model to identify and evaluate the presence of key components within each description. This process involved automated analysis to determine how well each PR description was crafted, focusing on the inclusion of essential information that aids in understanding and evaluating the PR's purpose and impact.</p>
GENEActiv accelerometer files collected in Vanuatu during FALAH project (anonymized version - first part)
<p><a title="GENEActiv" href="https://activinsights.com/technology/geneactiv/" target="_blank" rel="noopener">GENEActiv</a> accelerometer .csv and .RData files converted with a 1 second epoch from raw GENEActiv .bin files recorded in Vanuatu during <a title="FALAH website" href="https://falah.unc.nc/" target="_blank" rel="noopener">FALAH</a> project. Devices are 60-Hz triaxial accelerometers.</p> <p>This dataset also contains <strong>participantCharacteristics.csv</strong> that povides basic information about participants and <strong>read_a_binFile_share.R</strong> that is a short R code aiming at converting and saving accelerometer data from .bin files in 1 second epoch .csv files (consider the Methods section).</p> <p>Participant characteristics: 13 to 17 years old students.</p> <p>Number of participants: 72.</p> <p>Year of the study: 2023.</p> <p>Place of the study: Vanuatu.</p> <p>The accelerometer .csv and .RData files with a 1 second epoch and extracted from raw .bin files are available in the restricted datasets:</p> <ul> <li><a title="Anonymized dataset (first part)" href="https://doi.org/10.5281/zenodo.14043332" target="_blank" rel="noopener">anonymized version (first part)</a></li> <li><a title="Anonymized dataset (second part)" href="https://doi.org/10.5281/zenodo.14089478" target="_blank" rel="noopener">anonymized version (second part)</a></li> </ul> <p>The accelerometer raw .bin files are available in the restricted datasets:</p> <ul> <li><a title="Non-anonymized dataset (first part)" href="https://doi.org/10.5281/zenodo.14043547" target="_blank" rel="noopener">non-anonymized version (first part)</a></li> <li><a title="Non-anonymized dataset (second part)" href="https://doi.org/10.5281/zenodo.14089527">non-anonymized version (second part)</a></li> </ul> <p>Other participant characteristics (age, place of living, ...) and responses to questionnaires are available in <a title="Information and questionnaire associated with GENEActiv accelerometer files collected in Vanuatu during FALAH project (non-anonymized information)" href="https://doi.org/10.5281/zenodo.14189884" target="_blank" rel="noopener">a restricted non-anonymized dataset</a>.</p>
anonymous submission
<p>pending</p>
Anonymized Survey Responses
<p>This file contains the survey questions, 101 participants' responses to the survey questions and information on coding.</p>
FT4 Anonymous (4) 4-key Fagottino
<p>Dataset of FT4 Anonymous (4) 4-key fagottino, containing basic measurements and photos.</p>
Anonymized data for "The Impact of Argument Arrangement on Persuasiveness in Online Discussions"
<p>Anonymous upload of the data for the paper "The Impact of Argument Arrangement on Persuasiveness in Online Discussions" for blind review.</p>
Anonymous submission
<p>## Representations</p> <p>llmcomp_data_{humaneval,winogrande}.zip files contain the representations of the LLMs used in our work. Unzipped they will take approximately 73 GB of storage.</p> <p>## Similarity Scores</p> <p>The pre-computed similarity scores are contained in .parquet files.</p>
Anonymized source data files for figures in: Recurrent processes support a cascade of hierarchical decisions
Open the record for dataset details and reuse information.
Colorado Ongoing Basin Emissions Study (COBE) anonymized final data set of emissions measurements
Open the record for dataset details and reuse information.
Caleg 2019 Biodata Datasets For K-Anonymity Research
<p>No description provided.</p>
Data from: Stronger transferability but lower variability in transcriptomic- than in anonymous microsatellites: evidence from Hylid frogs.
A simple way to quickly optimize microsatellites in non-model organisms is to re-use loci available in closely related taxa; however, this approach can be limited by the stochastic and low cross-amplification success experienced in some groups (e.g. amphibians). An efficient alternative is to develop loci from transcriptome sequences. Transcriptomic microsatellites have been found to vary in their levels of cross-species amplification and variability, but this has to date never been tested in amphibians. Here, we compare the patterns of cross-amplification and levels of polymorphism of 18 published anonymous microsatellites isolated from genomic DNA versus 17 loci derived from a transcriptome, across nine species of tree frogs (Hyla arborea and Hyla cinerea group). We established a clear negative relationship between divergence time and amplification success, which was much steeper for anonymous than transcriptomic markers, with half-lives (time at which 50% of the markers still amplify) of 1.1 and 37 My respectively. Transcriptomic markers are significantly less polymorphic than anonymous loci, but remain variable across diverged taxa. We conclude that the exploitation of amphibian transcriptomes for developing microsatellites is an optimal approach for multi-species surveys (e.g. analyses of hybrid zones, comparative linkage mapping), while anonymous microsatellites may be more informative for fine-scale analyses of intraspecific variation. Moreover, our results confirm the pattern that microsatellite cross-amplification is greatly variable among amphibians, and should be assessed independently within target lineages. Finally, we provide a bank of microsatellites for Palearctic tree frogs (so far only available for H. arborea), which will be useful for conservation and evolutionary studies in this radiation.
Data from: Comparative population genetic analysis of bocaccio rockfish Sebastes paucispinis using anonymous and gene-associated simple sequence repeat loci
Comparative population genetic analyses of traditional and emergent molecular markers aid in determining appropriate use of new technologies. The bocaccio rockfish Sebastes paucispinis is a high-gene-flow marine species off the west coast of North America that experienced strong population decline over the past three decades. We used 18 anonymous and 13 gene associated simple sequence repeat loci (EST-SSRs) to characterize range-wide population structure with temporal replicates. No FST-outliers were detected using the LOSITAN program, suggesting that neither balancing nor divergent selection affected the loci surveyed. Consistent hierarchical structuring of populations by geography or year class was not detected regardless of marker class. The EST-SSRs were less variable than the anonymous SSRs, but no correlation between FST and variation or marker class was observed. General Linear Model analysis showed that low EST-SSR variation was attributable to low mean repeat number. Comparative genomic analysis with Gasterosteus aculeatus, Takifugu rubripes, and Oryzias latipes showed consistently lower repeat number in EST-SSRs than SSR loci that were not in ESTs. Purifying selection likely imposed functional constraints on EST-SSRs resulting in low repeat numbers that affected diversity estimates, but did not affect the observed pattern of population structure.
Social circles from Facebook (anonymized)
<p>Taken from the SNAP Datasets, we used this dataset to inspire our random dataset generating machinery to evaluate several tools part of the DICE methodology.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.