Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data set of Reversing anterior insular cortex neuronal hypoexcitability attenuates compulsive behavior in juvenile rats
<p>Development of self-regulatory competencies during adolescence is partially dependent on normative brain maturation. Here we report that adolescent rats as compared to adults exhibit impulsive and compulsive-like behavioral traits, the latter being associated with lower expression of mRNA levels of the immediate early gene zif268 in the anterior insula cortex (AIC). This suggests that underdeveloped AIC function in adolescent rats could contribute to an immature pattern of interoceptive cue integration in decision-making and a compulsive phenotype. In support of this, we report that layer 5 pyramidal neurons in the adolescent rat AIC are hypoexcitable and receive fewer glutamatergic synaptic inputs compared to adults. Chemogenetic activation of the AIC attenuated compulsive traits in adolescent rats supporting the idea that in early stages of AIC maturity there exists a suboptimal integration of sensory and cognitive information that contributes to inflexible behaviors in specific conditions of reward availability.</p>
Uruguayan candombe drumming - beat and downbeat tracking data set
<p>Uruguayan Candombe drumming - beat and downbeat data set</p> <p><br>===========<br>Description<br>===========</p> <p>This dataset includes more than 2 hours of Candombe recordings, with annotated beats and downbeats. It features 35 complete performances by renowned players, playing in groups of three to five drums. A total of 26 tambor players from various generations took part, representing the three key traditional Candombe styles. These recordings were created in a studio setting over a period of more than two decades as part of musicological research. The sessions from 1992 and 1995 were produced by Luis Jure, and the session from 2014 was produced by Luis Jure and Martín Rocamora.</p> <p>The audio files are in stereo with a sampling rate of 44.1 kHz and 16-bit precision. An expert annotated the location of beats and downbeats, totaling more than 4700 downbeats. The audio is provided in flac format and the annotations are in .csv files. The values in the first column of the CSV file represent the time instants of the beats. The numbers in the second column indicate both the bar number and the beat number within the bar. For example, 1.1, 1.2, 1.3, and 1.4 represent the four beats of the first bar. Therefore, each label ending in .1 indicates a downbeat. Another set of annotations is provided as .beats files in which the bar numbers are removed.</p> <p><br>========<br>Citation<br>========</p> <p>This dataset was released with the publication of the following paper. So if you find the dataset useful and want to reference it in your publications, please cite it.</p> <p>"Beat and Downbeat Tracking Based on Rhythmic Patterns Applied to the Uruguayan Candombe Drumming”. Leonardo Nunes, Martín Rocamora, Luis Jure, Luiz W. P. Biscainho. Proceedings of the 16th International Society for Music Information Retrieval Conference (ISMIR 2015), pages 264-270, Málaga, Spain, 26-30 October, 2015.</p> <p>@inproceedings{Nunes2015,<br> author = {Leonardo Nunes and Martín Rocamora and Luis Jure and Luiz W. P. Biscainho},<br> title = {{Beat and Downbeat Tracking Based on Rhythmic Patterns Applied to the Uruguayan Candombe Drumming}},<br> booktitle = {Proceedings of the 16th International Society for Music Information Retrieval Conference (ISMIR 2015)},<br> month = {Oct.},<br> address = {Málaga, Spain},<br> pages = {264--270},<br> year = {2015}<br>}</p> <p><br>========<br>Candombe<br>========</p> <p>Candombe is a vital part of Uruguayan popular culture, with thousands of practitioners and its rhythm influencing various genres of popular music. In 2009, UNESCO recognized it as part of the Intangible Cultural Heritage of Humanity. While it originated in Uruguay, Candombe has its roots in the culture brought by African slaves in the 18th century. Over time, it has evolved to incorporate the descendants of European immigrants and has become a part of the entire society. Candombe drumming, with its unique rhythm, is the essential element of this tradition, which also includes dancing, symbolic characters, and costumes.</p> <p>The drum used in Candombe is called tambor, which is Spanish for "drum." There are three different sizes: chico (small), repique (medium), and piano (big). Each size has its own unique sound, ranging from high to low frequency, and its own specific rhythmic patterns. All three drums are played with a stick in the dominant hand and the other one hitting the skin directly. The stick is also used to hit the shell when playing the clave or madera pattern. A minimal ensemble of drums (cuerda de tambores) must have at least one of each of the three drums. During a llamada de tambores, the ensemble usually consists of around 20 to 60 drums. When marching, the players walk forward with short steps synchronized with the beat, which is important for embodying the rhythm, even though it is not audible.</p> <p>===============<br>Acknowledgments<br>===============</p> <p>This work was partially supported by the funding agency Comisión Sectorial de Investigación Científica, Universidad de la República, Uruguay, and by the Department of Culture of the Municipality of Montevideo.</p> <p>This is the complete list of performers, in alphabetical order: Mariano Barroso, Eduardo 'Cacho' Giménez, Eduardo 'Malumba' Giménez, Francisco Giménez, José Luis Giménez, Jorge 'Foqué' Gómez, José Pedro 'Perico' Gularte, Luis 'Pocholo' Maciel, Julio Magariños, Raúl 'Neno' Magariños, Marcelo Magariños, Javier 'Cerdo' Martirena, Wilson Martirena, Eduardo 'Tierra' Nilo, Sergio Ortuño, Fernando 'Lobo' Núñez, Edinson 'Palo' Oviedo, Gustavo Oviedo, Egdardo Pintos, Luis 'Mocambo' Quiroz, Rodolfo 'Pelado' Rodríguez, Fernando 'Hurón' Silva, Juan Silva, Raúl Silva, Waldemar 'Cachila' Silva, and Héctor Manuel Suárez.</p>
MItosis DOmain Generalization Challenge 2022 (MICCAI MIDOG 2022), Training data set (PNG version)
<p>This is the training dataset of the MItosis DOmain Generalization (MIDOG) challenge 2022, held in conjunction with MICCAI 2022. Please find the structured challenge description at 10.5281/zenodo.6362337.</p> <p>The training set consists of 405 tumor cases in total across six tumor types:</p> <ul> <li>Canine Lung Cancer (44 cases, scanned with 3DHistech Pannoramic Scan II)</li> <li>Human Breast Cancer (150 cases, scanned using three scanners, part of MIDOG2021 dataset)</li> <li>Canine Lymphoma (55 cases, scanned with 3DHistech Pannoramic Scan II)</li> <li>Human neuroendocrine tumor (55 cases, scanned with Hamamatsu NanoZoomer XR)</li> <li>Canine Cutaneous Mast Cell Tumor (50 cases, scanned with Aperio ScanScope CS2)</li> <li>Human melanoma (51 cases, scanned with Hamamatsu NanoZoomer XR) (no labels provided)</li> </ul> <p>From each WSI, a trained pathologist selected an area of 2mm² corresponding to approximately 10 high power fields, according to the grading scheme of Elston and Ellis. We cropped this area and provide it as PNG files in this data set due to restrictions in data set size on zenodo. Each file includes the resolution (in dots per inch, DPI) of the original scanned images.</p> <p>The training set contains 9501 mitotic figures (MF) and 11051 hard examples (non-mitotic figures). All annotations are provided in MS COCO JSON format and as SQLITE database (SlideRunner format).</p>
Code and data sets for "MS²Rescore: Data-driven rescoring dramatically boosts immunopeptide identification rates"
<p>Code used to prepare data sets, train and evaluate new MS²PIP models, evaluate MS²Rescore for immunopeptidomics, and generate figures. See README.md for more information on how to use these files and reproduce the results reported in the manuscript titled "MS²Rescore: Data-driven rescoring dramatically boosts immunopeptide identification rates".</p>
Crop Classification Data-set
<p>The dataset consists of crop type training and testing data of more than 10 classes, collected using Ground Truth Surveys, in Harichand region of Khyber Pakhtoonkhwa, Pakistan. </p> <p>The dataset also contains 2 tiff files having Planet-Scope and Sentinel-2 raster data. </p> <p>https://drive.google.com/drive/folders/1SweabTezj78btq9wd3PRZYWrR_4gUuC4</p>
Inhabiting Extraterrestrial Space - Data set of 477 images
<p><strong>Data set of 477 images on Space Habitat</strong></p> <p>Abstracted from the research project « Habiter l’espace extraterrestre »</p> <p>HEAD-Genève<br> Partenaire scientifique : L’Observatoire de l’Espace, le laboratoire culturel du CNES (Paris)<br> Projet financé par le Fonds national suisse de la recherche scientifique (FNS)</p> <p>website : habitat-extraterretre.ch</p> <p> </p> <p><em><strong>Abstract of the project</strong></em></p> <p>Objects conceived and realized to inhabit extraterrestrial space are arousing a strong resurgence of interest with regards to the ecological questions as well as the economic issues they cover. Thus, "Leaving earth" has become a contemporary reality often mentioned but whose understanding is, however, rarely based on scientific research dedicated to space or extraterrestrial habitats and remains most often far from a specific basis, whether historical, material or cultural.</p> <p>The two historical and cultural lines followed by this research start from the same point: the implementation of a corpus of space research images made up of documents produced within the framework of authenticated works, supported by state institutions at an international level (Space Agencies) belonging to an official voted project and specifying one or more elements of an inhabited space object. At the margins of this framework, we have opened this corpus of documents to "pioneer" engineers who worked before the advent of the space age in 1957 and whose influence will be lasting on its future developments. </p> <p>The corpus of images from the fields of space communication, architecture, cinema and visual arts, are based on proven links between the objects and images they gather and those of space research. They are also opened to other images and objects, whose connections stem from these first connections.</p> <p>The relationships of these images and objects can be queried through descriptors applying to all the fields of the database. They ensure a given research to have transversal or multidisciplinary dimensions. Each corpus generated by these researches is a scientific material of knowledge related to the cultural history of space habitat. This research main goal is to enable researchers from the human sciences as well as artists to seize these corpuses and to contribute in the writing of this history.</p> <p> </p> <p><em><strong>Research team</strong></em></p> <p>Christophe Kihm, Associate Professor, HEAD - Geneva, HES-SO (principal applicant)</p> <p>Floriane Germain, Doctor, PHD in museum studies, mediation, heritage, space archives expert (scientific assistant),</p> <p>Jill Gasparina, Assistant Professor, HEAD - Geneva, HES-SO (scientific collaborator)</p> <p>Anne-Lyse Renon, Senior lecturer at the Laboratory of practices and theory of contemporary art, University of Rennes 2 (scientific assistant)</p> <p>With the cllaboration of Gérard Azoulay, director of the Observatoire de l'Espace, the cultural laboratory of the CNES (Paris).</p> <p> </p> <p><em><strong>General information </strong></em></p> <p>Images in the database are classified according to five fields:</p> <p>- Space research</p> <p>- Space communication</p> <p>- Architecture</p> <p>- Cinema</p> <p>- Visual arts</p> <p>The following dataset gives access to the 477 image files of the database. Each image file includes descriptors and metadata linking the image to other images in the database such as:</p> <p>Field, Subfield, Title, Author(s), Date, Sponsor(s), Country of Origin, Type of document, Dimensions, Creation Techniques , Medium of the Original, Conservation Place of the Original, Diffusion Medium of the Document, Inhabitants, Type of Habitats, Situation in Space, Dedicated Activitie(s), Localisation (Associated Space), Related Material Object(s), Situation in the Research Process, Type of Collaboration, Collaborator(s), Related Theme(s), Comments, Related Project, Related Object.</p>
GouDa - Generation of universal Data Sets
<p>GouDa is a tool for the generation of universal data sets to evaluate and compare existing data preparation tools and new research approaches. It supports diverse error types and arbitrary error rates. Ground truth is provided as well. It thus permits better analysis and evaluation of data preparation pipelines and simplifies the reproducibility of results.</p> <p>Publication: <em>V. Restat, G. Boerner, A. Conrad, and U. Störl. GouDa - Generation of universal Data Sets. In Proceedings of Data Management for End-to-End Machine Learning (DEEM’22), Philadelphia, USA, 2022. https://doi.org/10.1145/3533028.3533311</em></p>
Data set for the journal article Structural Analysis of Metal Coordination Sites in Single-Atom Catalysts Based on Carbon Nitrides
<p>The data is organized according to the figure in the manuscript. </p>
Data set of antibiotic study-1
<p>This data set includes the information related to appropriate antibiotic use (labeled as "appropriate_use_antibiotics") among the urban people of Bangladesh. The data set is included sociodemographic information such as age, sex, marital status, and educational status. The attitudes towards antibiotic use are included the information of within 2 months of antibiotic taking frequency (labeled as "took_antibiotics"), awareness of use (labeled as "attitude_awareness_use"), abuse of antibiotics (labeled as "attitude_abuse_antibiotic"), antibiotic resistance (labeled as "attitude_resistance"), and effect of resistance (labeled as "attitude_effect_resistance"). The knowledge of antibiotics treatable diseases such as COVID-19 (labeled as "knowledge_COVID-19"), dengue (labeled as "knowledge_dengue"), diabetics (labeled as "knowledge_diabetics"), pneumonia (labeled as "knowledge_pneumonia"), and tuberculosis (labeled as "knowledge_tuberculosis") are included. The knowledge of types of disease specification are also included bacterial (labeled as "knowledge_bacterial"), viral (labeled as "knowledge_viral"), parasitic (labeled as "knowledge_parasitic"), fungal (labeled as "knowledge_fungal"), and helminthic (labeled as "knowledge_helminthic"). Finally, knowledge of antimicrobials drugs specifications is included Penicillin (labeled as "knowledge_Penicillin"), Amoxicillin (labeled as knowledge_Amoxicillin), Cefixime (labeled as "knowledge_Cefixime"), Azithromycin (labeled as "knowledge_Azithromycin"), Remdisivir (labeled as "knowledge_Remdisivir"), and Albendazole (labeled as "knowledge_Albendazole").</p> <p>Note: The response "Yes_cor" or "No_cor", indicated that the responses were correct to the respective items. </p>
Galaxy Training Data for "Evaluating and ranking a set of pathways based on multiple metrics"
<pre>This dataset provides the inputs needed for the Galaxy Pathway Analysis workflow training tutorial (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). This workflow asseses the performance of predicted pathways by computing 4 criteria (target product flux, thermodynamic feasibility, pathway length, and enzyme availability). A score inform the user about the best candidate pathways to produce a compound of interest. The generated output is a collection of scored and ranked heterologous pathways. The content of the dataset is as follows: - A set of pathways provided in the SBML format (Systems Biology Markup Language) to be ranked, modeling heterologous pathways such as those outputted by the RetroSynthesis workflow (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). - The GEM (Genome-scale metabolic models) which is a formalized representation of the metabolism of the host organism (the model is E. coli iML1515), provided in the SBML format.</pre>
Data set of manuscript entiteld Sigma Oscillations Protect or Reinstate Motor Memory Depending on their Temporal Coordination with Slow Waves
<p>Data set of manuscript entiteld Sigma Oscillations Protect or Reinstate Motor Memory Depending on their Temporal Coordination with Slow Waves</p> <p>EEG files : Brainvision</p> <p>*.dat; *.edf : eeg files of the experimental nap </p> <p>*.vhdr: corresponding header files</p> <p>*.vmrk: corresponding marker files time starting from recording sample</p> <p>*._Pauselog.txt: markers with condition in the stimulation computer time</p> <p>Sleep Scores: </p> <p>Scores by the sleep technician (one file per participants). </p> <p>Surveys:</p> <p>General survey (first screening) and complete survey (for inclusion). Code correspondance is provided in ID_correspondance.txt</p> <p>BehavData:</p> <p>For each participant, a text file with all the cues and responses logged for each tasks. *_raw.txt are present when an issue arised during experiment and a task needed to be re-started. </p> <p> </p>
The input data set includes 729 objects (patients) and 39 variables (clinical qualitative and quantitative descriptors).
<p>For reliable data treatment and interpretation qualitative descriptors were omitted and only numerical clinical indicators were included in the data matrix. Finally, the data set dimension was [729 x 18].</p> <p> The data were treated by hierarchical cluster analysis and factor analysis. The major goal of the data mining was to reach statistically significant partitioning of the objects and variables into similarity patterns (clusters) which helps to better understand the data structure, to assess the meaning of the partitioning achieved, thus promoting the evaluation of the health status of the patients and the role of specific descriptors for the formation of the partitioning patterns.</p> <p>3D classification Python tool.</p>
Groundwater level change data set and SAR analysis data set from the end of 2016 to the end of 2020 in the Osaka Plain and Kyoto Basin, Japan
<p>The .dat file contains data on groundwater level changes. Groundwater level data is hourly, and each line contains data for one day (24 hours). The name of the .dat files consists of the name of the groundwater level station and the observation period. Datasets that include raw in the name contain the respective date in the first column.</p> <p>Source: Water information System, Ministry of Land, Infrastructure, Transport and Tourism, 2017-2020. Groundwater level search results, http://www1.river.go.jp/ (accessed on 20 October 2021) [Translated from Japanese.] (In Japanese).</p> <p>SLC data for InSAR analysis can be obtained by running the .py file in python. python files are provided to obtain SLC data by two orbits, Ascending and Descending, respectively.</p> <p>Source: European Space Agency (ESA), https://search.asf.alaska.edu/#/</p>
Data set for the fauna material excavated at the Liang Abu site (East Kalimantan, Indonesia)
<p>Data set for the fauna material excavated at the Liang Abu site (East Kalimanta, Indonesia): count by subphylum, class, order, family, genus, and species for the squares 11Ec, 11Ed, and12Eb.</p>
Data set for the lithic material excavated at the Liang Abu site (East Kalimantan, Indonesia)
<p>Three tables describing the lithic material excavated at the Liang Abu site (East Kalimantan, Indonesia):</p> <ul> <li>abu-lithic-class.csv: count by class by stratigraphic layer</li> <li>abu-lithic-spatial-distribution.csv: count by square and stratigraphic layer</li> <li>abu-lithic-flakes.csv: morphometric values of the flakes</li> </ul>
Applied Statistics data-sets
<p>This collection of data-sets is a companion to the <em><strong>Applied Statistics eBook</strong></em>, which is available to download at <a href="https://zenodo.org/record/6783846#.Yr3EQ3bMKUk">https://zenodo.org/record/6783846#.Yr3EQ3bMKUk</a> . The datasets have been selected for illustrative purposes and should not be relied upon as a basis for substantive research. </p>
Login Data Set for Risk-Based Authentication
<p><strong>Login Data Set for Risk-Based Authentication</strong></p> <blockquote> <p>Synthesized login feature data of >33M login attempts and >3.3M users on a large-scale online service in Norway. Original data collected between February 2020 and February 2021.</p> </blockquote> <p>This data sets aims to foster research and development for <a href="https://riskbasedauthentication.org">Risk-Based Authentication (RBA)</a> systems. The data was synthesized from the real-world login behavior of more than 3.3M users at a large-scale single sign-on (SSO) online service in Norway.</p> <p>The users used this SSO to access sensitive data provided by the online service, e.g., a cloud storage and billing information. We used this data set to study how the <a href="https://doi.org/10.14722/ndss.2016.23240">Freeman et al. (2016)</a> RBA model behaves on a large-scale online service in the real world (see <a href="#publication">Publication</a>). The synthesized data set can reproduce these results made on the original data set (see <a href="#study-reproduction">Study Reproduction</a>). Beyond that, you can use this data set to evaluate and improve RBA algorithms under real-world conditions.</p> <p><strong>WARNING:</strong> The feature values are plausible, but still <strong>totally</strong> <strong>artificial</strong>. Therefore, you should NOT use this data set in productive systems, e.g., intrusion detection systems.</p> <p><strong>Overview</strong></p> <p>The data set contains the following features related to each login attempt on the SSO:</p> <table> <thead> <tr> <th>Feature</th> <th>Data Type</th> <th>Description</th> <th>Range or Example</th> </tr> </thead> <tbody> <tr> <td>IP Address</td> <td>String</td> <td>IP address belonging to the login attempt</td> <td>0.0.0.0 - 255.255.255.255</td> </tr> <tr> <td>Country</td> <td>String</td> <td>Country derived from the IP address</td> <td>US</td> </tr> <tr> <td>Region</td> <td>String</td> <td>Region derived from the IP address</td> <td>New York</td> </tr> <tr> <td>City</td> <td>String</td> <td>City derived from the IP address</td> <td>Rochester</td> </tr> <tr> <td>ASN</td> <td>Integer</td> <td>Autonomous system number derived from the IP address</td> <td>0 - 600000</td> </tr> <tr> <td>User Agent String</td> <td>String</td> <td>User agent string submitted by the client</td> <td>Mozilla/5.0 (Windows NT 10.0; Win64; ...</td> </tr> <tr> <td>OS Name and Version</td> <td>String</td> <td>Operating system name and version derived from the user agent string</td> <td>Windows 10</td> </tr> <tr> <td>Browser Name and Version</td> <td>String</td> <td>Browser name and version derived from the user agent string</td> <td>Chrome 70.0.3538</td> </tr> <tr> <td>Device Type</td> <td>String</td> <td>Device type derived from the user agent string</td> <td>(<code>mobile</code>, <code>desktop</code>, <code>tablet</code>, <code>bot</code>, <code>unknown</code>)<a href="#fn1"><sup>1</sup></a></td> </tr> <tr> <td>User ID</td> <td>Integer</td> <td>Idenfication number related to the affected user account</td> <td>[Random pseudonym]</td> </tr> <tr> <td>Login Timestamp</td> <td>Integer</td> <td>Timestamp related to the login attempt</td> <td>[64 Bit timestamp]</td> </tr> <tr> <td>Round-Trip Time (RTT) [ms]</td> <td>Integer</td> <td>Server-side measured latency between client and server</td> <td>1 - 8600000</td> </tr> <tr> <td>Login Successful</td> <td>Boolean</td> <td><code>True</code>: Login was successful, <code>False</code>: Login failed</td> <td>(<code>true</code>, <code>false</code>)</td> </tr> <tr> <td>Is Attack IP</td> <td>Boolean</td> <td>IP address was found in known attacker data set</td> <td>(<code>true</code>, <code>false</code>)</td> </tr> <tr> <td>Is Account Takeover</td> <td>Boolean</td> <td>Login attempt was identified as account takeover by incident response team of the online service</td> <td>(<code>true</code>, <code>false</code>)</td> </tr> </tbody> </table> <p><strong>Data Creation</strong></p> <p>As the data set targets RBA systems, especially the <a href="https://doi.org/10.14722/ndss.2016.23240">Freeman et al. (2016)</a> model, the statistical feature probabilities between all users, globally and locally, are identical for the categorical data. All the other data was randomly generated while maintaining logical relations and timely order between the features.</p> <p>The timestamps, however, are not identical and contain randomness. The feature values related to IP address and user agent string were randomly generated by publicly available data, so they were very likely not present in the real data set. The RTTs resemble real values but were randomly assigned among users per geolocation. Therefore, the RTT entries were probably in other positions in the original data set.</p> <ul> <li> <p>The country was randomly assigned per unique feature value. Based on that, we randomly assigned an ASN related to the country, and generated the IP addresses for this ASN. The cities and regions were derived from the generated IP addresses for privacy reasons and do not reflect the real logical relations from the original data set.</p> </li> <li> <p>The device types are identical to the real data set. Based on that, we randomly assigned the OS, and based on the OS the browser information. From this information, we randomly generated the user agent string. Therefore, all the logical relations regarding the user agent are identical as in the real data set.</p> </li> <li> <p>The RTT was randomly drawn from the login success status and synthesized geolocation data. We did this to ensure that the RTTs are realistic ones.</p> </li> </ul> <p><strong>Regarding the Data Values</strong></p> <p>Due to unresolvable conflicts during the data creation, we had to assign some unrealistic IP addresses and ASNs that are not present in the real world. Nevertheless, these do not have any effects on the risk scores generated by the <a href="https://doi.org/10.14722/ndss.2016.23240">Freeman et al. (2016)</a> model.</p> <p>You can recognize them by the following values:</p> <ul> <li> <p>ASNs with values >= 500.000</p> </li> <li> <p>IP addresses in the range 10.0.0.0 - 10.255.255.255 (10.0.0.0/8 CIDR range)</p> </li> </ul> <p><strong>Study Reproduction</strong></p> <p>Based on our evaluation, this data set can reproduce our study results regarding the RBA behavior of an RBA model using the IP address (IP address, country, and ASN) and user agent string (Full string, OS name and version, browser name and version, device type) as features.</p> <p>The calculated RTT significances for countries and regions inside Norway are not identical using this data set, but have similar tendencies. The same is true for the Median RTTs per country. This is due to the fact that the available number of entries per country, region, and city changed with the data creation procedure. However, the RTTs still reflect the real-world distributions of different geolocations by city.</p> <p>See <a href="RESULTS.md">RESULTS.md</a> for more details.</p> <p><strong>Ethics</strong></p> <p>By using the SSO service, the users agreed in the data collection and evaluation for research purposes. For study reproduction and fostering RBA research, we agreed with the data owner to create a synthesized data set that does not allow re-identification of customers.</p> <p>The synthesized data set does not contain any sensitive data values, as the IP addresses, browser identifiers, login timestamps, and RTTs were randomly generated and assigned.</p> <p><strong>Publication</strong></p> <p>You can find more details on our conducted study in the following journal article:</p> <p><a href="https://doi.org/10.1145/3546069">Pump Up Password Security! Evaluating and Enhancing Risk-Based Authentication on a Real-World Large-Scale Online Service</a> (2022)<br> <em>Stephan Wiefling, Paul René Jørgensen, Sigurd Thunem, and Luigi Lo Iacono</em>.<br> <em>ACM Transactions on Privacy and Security</em></p> <p><strong>Bibtex</strong></p> <pre>@article{Wiefling_Pump_2022, author = {Wiefling, Stephan and Jørgensen, Paul René and Thunem, Sigurd and Lo Iacono, Luigi}, title = {Pump {Up} {Password} {Security}! {Evaluating} and {Enhancing} {Risk}-{Based} {Authentication} on a {Real}-{World} {Large}-{Scale} {Online} {Service}}, journal = {{ACM} {Transactions} on {Privacy} and {Security}}, doi = {10.1145/3546069}, publisher = {ACM}, year = {2022} }</pre> <p><strong>License</strong></p> <p>This data set and the contents of this repository are licensed under the <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0)</a> license. See the <a href="LICENSE">LICENSE</a> file for details. If the data set is used within a publication, the following journal article has to be cited as the source of the data set:</p> <p>Stephan Wiefling, Paul René Jørgensen, Sigurd Thunem, and Luigi Lo Iacono: Pump Up Password Security! Evaluating and Enhancing Risk-Based Authentication on a Real-World Large-Scale Online Service. In: ACM Transactions on Privacy and Security (2022). doi: <a href="https://doi.org/10.1145/3546069">10.1145/3546069</a></p> <ol> <li> <p>Few (invalid) user agents strings from the original data set could not be parsed, so their device type is empty. Perhaps this parse error is useful information for your studies, so we kept these 1526 entries.<a href="#fnref1">↩︎</a></p> </li> </ol>
Lagrangian particles in turbulence: An experimental data set
<p>This data set contains the coordinates of Lagrangian trajectories in a quasi-homogeneous isotropic turbulent flow. The trajectories were measured in a water tank experiment using the 3D-PTV method through the MyPTV open-source software (https://github.com/ronshnapp/MyPTV). The total duration of this data set corresponds to 20 seconds of recording at a rate of 500 Hz, and it holds roughly 575,000 trajectories.</p> <p> <br> The data is stored in two text files in a tab-separated format, where each file coresponds to 10 seconds of recording. Each row in the files outlines a single measurement point from the experiment. The various columns correspond to the following information:<br> 1 - trajectory_id <br> 2 - x [mm] <br> 3 - y [mm]<br> 4 - z [mm]<br> 5 - vx [mm/frame]<br> 6 - vy [mm/frame]<br> 7 - vz [mm/frame]<br> 8 - ax [mm/frame^2]<br> 9 - ay [mm/frame^2]<br> 10 - az [mm/frame^2]<br> 11 - time [frame]<br> where trajectory_id uniquely marks samples that correspond to the same physical trajectory; x, y, and z are the position components in the three orthogonal space directions; vx, vy, and vz correspond to the velocity components; ax, ay, and az that correspond to the acceleration components.</p> <p><br> The flow in the experiment was forced using a system of 8 propellers, powered by DC motors that were positioned at the corners of the cylindrical, octagonally shaped, water tank. The propellers were changing their direction of rotation at random time intervals with an average interval duration of 0.1 seconds. The root mean square of the turbulent flow fluctuations is about 100 millimeters per second. There is a time-averaged secondary circulation with a magnitude of roughly 66% of the root mean squared fluctuation strength. The Taylor microscale Reynolds number is estimated as about 188. </p>
An experimental data set for analysis of the thermophysical behavior of a single-story mechanically ventilated double-skin façade (DSF) in fixed boundary conditions corresponding to winter/mid-season and summer cases
<p>Double-skin facades (DSFs) are dynamic and flexible building envelopes that employ a ventilated cavity to either prevent or reduce the solar-induced cooling load or exploit solar energy for passive solar heating. The mechanical ventilation of the cavity offers higher flexibility and control than natural ventilation, as the latter largely depends on stochastic and unpredictable external conditions. Furthermore, when mechanical ventilation rates are combined with the operation of a shading device, the possibilities for controlling the accumulated heat in the cavity of the DSF increase further. Therefore, this experimental campaign systematically investigates how these two important features interact in controlling the cavity's thermal load and airflow conditions. The measurement collected during the experiments constitutes a dataset that contains the results of a series of experimental runs where the different configurations of DSF, in terms of mechanical ventilation rate and venetian blinds, have been subjected to two representative boundary conditions through a climate simulator facility equipped with a solar simulator device. The full-scale DSF mock-up, which includes venetian blinds installed in a 200 mm ventilated cavity, is operated in this experiment in two modes: outdoor air curtain (OAC) and supply air (SA) mode. Tests were carried out under a steady-state regime with different boundary conditions. For the analysis of the utilization of the excess heat accumulated in the cavity and prevention of DSF overheating, boundary conditions corresponding to g-value calculations were selected. For the analysis of air preheating in the DSF cavity, the boundary conditions corresponding to late winter/mid-season weather (cold outdoor air and low-to-moderate solar irradiance) were chosen. The entire set of experimental data collected during the tests is made publicly available to enable the scientific community to access experimental data to further analyze this problem or for model validation purposes. The data set supplements the open-access paper entitled "<strong>Control of heat transfer in single-story mechanically ventilated facades</strong>" (<a href="https://doi.org/10.1016/j.enbuild.2022.112304">https://doi.org/10.1016/j.enbuild.2022.112304</a>), where additional information about the aims of the experiments, the detailed methods, and other data processing procedures can be found. The database is supported by a guide ("Guide.pdf"), where further explanations about how to read data and schematic drawings of the sensor layout are provided. The collection of experimental tests is divided into two files, according to two considered cases:</p> <ul> <li><strong>DSF operating in outdoor air curtain mode </strong>(24 steady-state measurements). The following factors were changed: mechanical ventilation rate (0, 10, 15, 20, 30, 40, 50, and 100 % of maximum fan power) and venetian blind configuration (closed θ=0 º, semi-open θ=45 º, and raised blinds). The outdoor and indoor temperatures, 30 ℃ and 25 ℃, and solar irradiance of 500 Wm<sup>-2</sup> were replicated. [file name: "Summer.csv"],</li> <li><strong>DSF operating in supply air mode</strong> (27 steady-state measurements). The following factors were changed: mechanical ventilation rate (0, 10, 15, 20, 30, 40, 50, 75, and 100 % of maximum fan power) and venetian blind configuration (closed θ=0 º, semi-open θ=45 º, and raised blinds). The outdoor and indoor temperatures, 10 ℃ and 25 ℃, and solar irradiance of 300 Wm<sup>-2</sup> were replicated. [file name: " Winter_MidSeason.csv"]</li> </ul> <p>Any inquiries about the experimental data can be sent to: <a href="mailto:aleksandar.jankovic@ntnu.no">aleksandar.jankovic@ntnu.no</a></p> <p>The activities presented in this paper were carried out within the research project "REsponsive, INtegrated, VENTilated - REINVENT – windows," supported by the Research Council of Norway through the research grant 262198, and the partners SINTEF, Hydro Extruded Solutions, Politecnico di Torino and Aalto University.</p>
data set of scopus about technology and halal meat supply chain publications
<p>This is the dataset of papers related to the topics of technology and halal meat supply chain collected from Scopus. The duration of year was 2008-2022</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.