Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
Dataset for algorithmic thinking skills assessment: Results from the virtual CAT large-scale study in Swiss compulsory education
<p><strong>Overview</strong><br>This dataset was collected during a main study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.<br>As algorithmic thinking becomes increasingly vital in our digital age, this study bridges the gap between traditional assessments and the needs of today's learners by introducing a digital platform. The virtual CAT, a digital adaptation of an unplugged assessment activity, offers scalable, automated assessments with reduced human intervention.</p> <p><strong>Study Context, Location and Participants</strong><br>To comprehensively investigate algorithmic competencies within compulsory education, exploring their variations and determining the factors influencing them, in Spring 2023 we conducted an experimental study with the virtual CAT's.<br>The sample comprises 129 students (65 girls and 64 boys), selected from nine classes across five public schools in Ticino and Solothurn cantons.</p> <p><strong>Data Collection</strong><br>During the data collection process, session and participant details were manually recorded by the administrator. <br>Each session has been assigned a unique identifier, and specific details, such as the date, canton, school name and type, and the students’ HarmoS grade (HG) level, have been recorded. <br>Student information are limited to sex and date of birth, with birth dates used to calculate ages, a significant factor in our demographic analysis. <br>To protect student privacy, unique identifiers have been assigned to each participant, keeping the data anonymous and secure. <br>The assessment tool automatically tracked all user interaction within the platform.<br>All data collected have been pseudonymised, aligning with prevailing open science practices in Switzerland (SNSF, 2021). <br>Data collection was integrated into a validation module of the app. </p> <p><strong>Data Features</strong><br>The dataset comprises the following files:</p> <ul> <li>STUDENTS_SESSIONS.csv</li> <li>RESULTS.csv</li> <li>LOGS.csv</li> <li>CANTONS.csv</li> <li>ALGORITHMS.csv</li> </ul> <p>These files collectively provide insights into the algorithmic actions of the students, demographic details, session logs, results, and more.</p> <p><strong>Usage & Ethics</strong><br>In the spirit of open science, this dataset is made available to the public after meticulous anonymisation to ensure all participants' privacy and ethical treatment. <br>Initial authorisations were secured from school administrators, teachers, and parents. <br>Detailed communication regarding the study's nature, data handling, and objectives was transparently shared with all stakeholders.</p> <p><strong>REFERENCES</strong></p> <p><strong>[1]</strong> A. Piatti, G. Adorni, L. El-Hamamsy, L. Negrini, D. Assaf, L. Gambardella & F. Mondada. (2022). The CT-cube: A framework for the design and the assessment of computational thinking activities. Computers in Human Behavior Reports, 5, 100166. <a href="https://doi.org/10.1016/j.chbr.2021.100166">https://doi.org/10.1016/j.chbr.2021.100166</a></p> <p><strong>[2]</strong> Adorni, G., & Piatti, S., & Karpenko, V. (2023). virtual CAT: An app for algorithmic thinking assessment within Swiss compulsory education. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10027851">https://doi.org/10.5281/zenodo.10027851</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/</a></p> <p><strong>[3]</strong> Adorni, G., & Karpenko, V. (2023). virtual CAT programming language interpreter. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10016535">https://doi.org/10.5281/zenodo.10016535</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/</a></p> <p><strong>[4]</strong> Adorni, G., & Karpenko, V. (2023). virtual CAT data infrastructure. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10015011">https://doi.org/10.5281/zenodo.10015011</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure</a></p> <p> </p>
Sample data for "Machine learning for large-scale forecasting"
<p>This dataset includes sample data for the Netherlands to run the machine learning baseline as described in the paper titled <em>Machine learning for large-scale crop yield forecasting</em>, accessible at <a href="https://doi.org/10.1016/j.agsy.2020.103016">https://doi.org/10.1016/j.agsy.2020.103016</a>. The software implementation of the machine learning baseline is available at: <a href="https://github.com/BigDataWUR/MLforCropYieldForecasting">https://github.com/BigDataWUR/MLforCropYieldForecasting</a>.</p> <p><strong>Notes:</strong></p> <p>The NUTS classification (Nomenclature of territorial units for statistics) is a hierarchical system for dividing up the economic territory of the EU and the UK (see Eurostat, 2016) for more details).</p> <p>Data</p> <p>The dataset consists of 11 CSV files. They are formatted to work as sample inputs to the machine learning baseline.</p> <ol> <li><strong>Crop Area Fractions </strong>(NUTS2, NUTS1): We aggregated the predictions of the machine learning baseline from NUTS2 to national (NUTS0) level by weighting them on the modeled crop area. Cerrani and López Lozano (2017) have described in detail the algorithm used to model crop areas for different NUTS levels. The data comes from the MARS Crop Yield Forecasting System (MCYFS) of European Commission's Joint Research Centre (JRC) (see Lecerf et al., 2019).</li> <li><strong>Centroids (NUTS2)</strong>: Data includes latitude, longitude and distance to coast of the centroids of NUTS2 regions.</li> <li><strong>Meteo Daily Data and Meteo Dekadal Data </strong>(NUTS2): The data comes from MCYFS (see EC-JRC, 2020). By default, the implementation uses daily data.</li> <li><strong>Remote Sensing Data</strong> (NUTS2, see Copernicus Global Land Service, 2020): Data includes fraction of absorbed photosynthetically active radiation (FAPAR) aggregated to NUTS2.</li> <li><strong>Soil Data</strong>: Data includes soil moisture information that can be used to calculate soil water holding capacity. The data comes from MCYFS (see Lecerf et al., 2019).</li> <li><strong>WOFOST data </strong>(NUTS2): The World Food Studies (WOFOST) crop model (van Diepen et al., 1989; Supit et al., 1994; de Wit et al. 2019) is a simulation model for the quantitative analysis of the growth and production of annual field crops. It is a mechanistic, dynamic model that explains daily crop growth on the basis of the underlying processes, such as photosynthesis, respiration and how these processes are influenced by environmental conditions. The crop simulation is fed by weather, soil and crop data. Observed meteorological data is interpolated on a regular 25 km grid using a method based on the distance, altitude and climatic region similarity between the center of grid cells and weather stations (see Van der Goot, 1998). WOFOST runs on the intersection between the 25 km meteorological grid and soil units based on the European soil map (http://esdac.jrc.ec.europa.eu/). In order to have the output data aggregated to administrative regions such as countries or provinces, simulation units are further intersected with the boundaries of these regions. The outputs at soil unit (STU) level are aggregated to grid level in an area weighted manner. Gridded simulations are aggregated to lowest NUTS level 3 considering the arable land area of each grid, derived from GLOBCOVER and CORINE Land Cover (Cerrani and Lopez Lozano, 2017). From NUTS3 to higher levels, crop area fractions for the current year, retrieved from Eurostat, are used to weight and aggregate the output (Cerrani and Lopez Lozano, 2017).</li> <li><strong>GAES data</strong>: GAES data includes agro-climatic features of regions, such as elevation and slope (from USGS-EROS, 2021), field size (from Lesiv et al., 2019), irrigated (crop) areas (from EC-JRC, 2020) and crop areas (from EC-JRC, 2020).</li> <li><strong>National yield statistics </strong>(NUTS0): These are the official Eurostat national yield statistics (Eurostat, 2020a). We used these yield statistics as reference to compare the machine learning predictions aggregated to NUTS0 and the actual MCYFS forecasts (see van der Velde and Nisini, 2019).</li> <li><strong>Regional yield statistics </strong>(NUTS2): We used NUTS2 yield statistics as labels to train and evaluate machine learning algorithms. We got NUTS2 yield statistics from The Central Bureau of Statistics (CBS) of the Netherlands (NL-CBS, 2020).</li> <li><strong>Past MCYFS Yield Forecasts </strong>(NUTS0): These are actual forecasts made by MCYFS in the past (see van der Velde and Nisini, 2019). We used the official Eurostat national yield statistics (see point 7 above) as the reference to compare the machine learning predictions aggregated to NUTS0 and MCYFS forecasts.</li> </ol> <p><strong>Crop ID and name mapping</strong></p> <p>2 : grain maize</p> <p>6 : sugar beets</p> <p>7 : potatoes</p> <p>90 : soft wheat</p> <p>93 : sunflower</p> <p>95 : spring barley</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We would like to thank S. Niemeyer from the European Commission’s Joint Research Centre (JRC) for the permission to provide open access to the Netherlands data. Similarly, we would like to thank M. van der Velde, L. Nisini and I. Cerrani from JRC for sharing with us past MCYFS forecasts and Eurostat national yield statistics.</p>
Widespread Sampling Biases in Herbaria Revealed from Large-Scale Digitization 1656-2016
Non-random collecting practices may bias conclusions drawn from analyses of herbarium records. Recent efforts to fully digitize and mobilize regional floras offer a timely opportunity to assess commonalities and differences in herbarium sampling biases. We determined spatial, temporal, trait, phylogenetic, and collector biases in ~5 million herbarium records, representing three of the most complete digitized floras of the world: Australia (AU), South Africa (SA), and New England, USA (NE) We identified numerous shared and unique biases among these regions. Shared biases included specimens i) collected close to roads and herbaria; ii) collected more frequently during spring; iii) of threatened species collected less frequently; and iv) of close relatives collected in similar numbers. Regional differences included i) over-representation of graminoids in SA and AU and of annuals in AU; and ii) peak collection during the 1910s in NE, 1980s in SA, and 1990s in AU. Finally, in all regions, a disproportionately large percentage of specimens were collected by a few individuals. These mega-collectors, and their associated preferences and idiosyncrasies, may have shaped patterns of collection bias via ‘founder effects’. Studies using herbarium collections should account for sampling biases and future collecting efforts should avoid compounding these biases.
Large-scale Ridesharing DARP Instances Based on Real Travel Demand
<p>This repository presents a set of large-scale Dial-a-Ride Problem (DARP) instances. The instances were created as a standardized set of ridesharing DARP problems for the purpose of benchmarking and comparing different solution methods.</p><p>The instances are based on real demand and realistic travel time data from 3 different US cities, Chicago, New York City and Washington, DC. The instances consist of real travel requests from the selected period, positions of vehicles with their capacities and realistic shortest travel times between all pairs of locations in each city.</p><p>The instances and results of two solution methods, the Insertion Heuristic, and the optimal Vehicle-group Assignment method, can be found in the dataset. The dataset and methodology used to create it are described in the paper <a href="https://arxiv.org/abs/2305.18859">Large-scale Ridesharing DARP Instances Based on Real Travel Demand</a>.</p>
S1000 corpus, large-scale tagging results and other supplementary files
<p>Data associated with the S1000 corpus</p><p>The tagger software for which the dictionary files in <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/tagger-organisms-dictionary-S1000.tar.gz">tagger-organisms-dictionary-S1000.tar.gz </a>can be used with can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a></p><p>The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/s1000-corpus-annotation-guidelines/</a></p><p>The S1000 corpus split in training, development and test sets in BRAT format can be found in <a href="https://zenodo.org/api/records/10285825/files/S1000-corpus.tar.gz">S1000-corpus.tar.gz</a><a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-corpus.tar.gz?versionId=ac7ce430-c265-49bb-8c8f-9b5f8e271cbe"> </a>and in CoNLL format here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/s1000-conll.tar.gz">s1000-conll.tar.gz</a></p><p>The tagging results of Jensenlab tagger for the S1000 test set are here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-jensenlab-tagger.tar.gz?versionId=d8d9c9f5-ee3b-4738-aefa-a4a95475d25d">S1000-jensenlab-tagger.tar.gz</a></p><p>The result from the large scale run in entire PubMed and PMC Open Access articles for Jensenlab tagger is provided here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz?versionId=48825928-9fc9-423c-8a4c-4f8994e95805">Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz</a></p><p>The model used for the large scale run of the transformer-based method is here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000_Transformer_based_tagger_large_scale_model.tar.gz?versionId=8e974f64-9abc-4449-a377-e3f97e91d612">S1000_Transformer_based_tagger_large_scale_model.tar.gz</a> and the results from the large scale tagging here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip?versionId=dc21a6ba-9763-4130-9f02-0341a885c692">Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip</a></p>
THÖR-MAGNI: A Large-scale Indoor Motion Capture Recording of Human Movement and Interaction
<h1>The THÖR-MAGNI Dataset Tutorials</h1> <p>THÖR-MAGNI datasets is a novel dataset of accurate human and robot navigation and interaction in diverse indoor contexts, building on the previous <a href="https://ieeexplore.ieee.org/abstract/document/8954833/">THÖR dataset protocol</a>. We provide position and head orientation motion capture data, 3D LiDAR scans and gaze tracking. In total, THÖR-MAGNI captures <strong>3.5 hours of motion of 40 participants on 5 recording days</strong>.</p> <p>This data collection is designed around systematic variation of factors in the environment to allow building cue-conditioned models of human motion and verifying hypotheses on factor impact. To that end, THÖR-MAGNI encompasses 5 scenarios, in which some of them have different conditions (i.e., we vary some factor):</p> <ul> <li>Scenario 1 (plus conditions A and B): <ul> <li> Participants move in groups and individually;</li> <li> Robot as static obstacle;</li> <li> Environment with 3 obstacles and lane marking on the floor for <strong>condition B</strong>;</li> </ul> </li> </ul> <ul> <li> Scenario 2: <ul> <li> Participants move in groups, individually and transport objects with variable difficulty (i.e. bucket, boxes and a poster stand);</li> <li> Robot as static obstacle;</li> <li> Environment with 3 obstacles;</li> </ul> </li> </ul> <ul> <li>Scenario 3 (plus conditions A and B): <ul> <li> Participants move in groups, individually and transporting objects with variable difficulty (i.e. bucket, boxes and a poster stand). We denote each role as: <em>Visitors-Alone, Visitors-Group 2, Visitors-Group 3, Carrier-Bucket, Carrier-Box, Carrier-Large Object;</em></li> <li> Teleoperated robot as moving agent: in <strong>condition A</strong>, the robot moves with differential drive; in <strong>condition </strong>B, the robot moves with omni-directional drive;</li> <li> Environment with 2 obstacles;</li> </ul> </li> </ul> <ul> <li>Scenario 4 (plus conditions A and B): <ul> <li> All participants, denoted as <em>Visitors-Alone HRI</em> interacted with the teleoperated mobile robot;</li> <li> Robot interacted in two ways: in <strong>condition A</strong> (Verbal-Only), the Anthropomorphic Robot Mock Driver (ARMoD), a small humanoid NAO robot on top of the mobile platform, only used speech to communicate the next goal point to the participant; in <strong>condition B</strong> the ARMoD used speech, gestures and robotic gaze to convey the same message;</li> <li> Free space environment</li> </ul> </li> </ul> <ul> <li>Scenario 5: <ul> <li> Participants move alone (<em>Visitors-Alone</em>) and one of the participants, denoted as <em>Visitors-Alone HRI</em>, transport objects and interact with the robot;</li> <li> The ARMoD is remotely controlled by an experimenter and proactively offers help;</li> <li> Free space environment;</li> </ul> </li> </ul> <h2>Preliminary steps</h2> <p>Before proceeding, make sure to download the data from ZENODO</p> <h3>1. Directory Structure</h3> <p>├── CLiFF_Maps <- Directory for CLiFF Maps for all files</p> <p> ├── Files <- Directory for the csv files</p> <p> ├── Readme.md</p> <p>├── CSVs_Scenarios <- Directory for aligned data for all scenarios</p> <p> ├── Scenario_1 <- Directory for the csv files for Scenario 1</p> <p> ├── Scenario_2 <- Directory for the csv files for Scenario 2</p> <p> ├── Scenario_3 <- Directory for the csv files for Scenario 3</p> <p> ├── Scenario_4 <- Directory for the csv files for Scenario 4</p> <p> ├── Scenario_5 <- Directory for the csv files for Scenario 5</p> <p>├── docs</p> <p> ├── tutorials.md <- Tutorials document on how to use the data</p> <p>├── Lidar_sample</p> <p> ├── Files <- Directory for sample files</p> <p> ├── 170522_SC3B_1 <- Directory for the pcd files</p> <p> ├── 170522_SC3B_1.csv <- Synchronization file with QTM</p> <p> ├── manual_view_point.json <- json file with manual view point for visualization</p> <p> ├── requirements.txt <- script pip requirements</p> <p> ├── visualize_pcd.py <- script visualize the lidar data</p> <p> ├── Readme.md</p> <p>├── maps <- Directory for maps of the environment (PNG files) and offsets (json file)</p> <p> ├── offsets.json <- Offsets of the map with respect to the global coordinate frame origin</p> <p> ├── {date}_SC{sc_id}_map.png <- Maps for `date` in {1205, 1305, 1705, 1805} and `sc_id` in {1A, 1B, 2, 3}</p> <p> ├── 3009_map.png <- Map for the Scenarios 4A, 4B and 5</p> <p>├── MP4_Videos</p> <p> ├── Files <- Directory for the mp4 files</p> <p> ├── pupil_scene_camera_instrinsics.json <- json file with the intrinsics of pupil camera</p> <p>├── TSVs_RAWET <- Directory for the TSV files for the Raw Eyetracking data for all Scenarios</p> <p> ├── synch_info.csv <- Event markers necessary to align motion capture with eyetracking data</p> <p> ├── Files <- Directory with all the raw eyetracking TSV files</p> <p>├── goals_positions.csv <- File with the goals locations</p> <p> </p> <h3>2. Data Structure and Dataset Files</h3> <p>Withing each Scenario directory, each csv file contains:</p> <p><strong>2.1. Headers</strong></p> <p>The dataset metadata overview contains important information found in the CSV file headers. This reference is designed to help users understand and use the dataset effectively. The headers include details such as FILE_ID, which provides information on the date, scenario, condition, and run associated with each recording. The header of the document includes important quantities such as the number of frames recorded (N_FRAMES_QTM), the count of rigid bodies (N_BODIES), and the total number of markers (N_MARKERS).</p> <p>It also provides information about the order of the contiguous rotation matrix (CONTIGUOUS_ROTATION_MATRIX), modalities measured with units, and specified measurement units. The text presents details on the eyetracking devices used in each recording, including their infrared sensor and scene camera frequencies, as well as an indication of the presence of eyetracking data.</p> <p>The header provides specific information about rigid bodies, including their names (BODY_NAMES), role labels (BODY_ROLES), and the number of markers associated with each rigid body (BODY_NR_MARKERS). Finally, the table lists all marker names used in the file.</p> <p>This metadata provides researchers and practitioners with essential guidance on recording information, data quantities, and specifics about rigid bodies and markers. It is a valuable resource for understanding and effectively using the dataset in the CSV files.</p> <p><strong>2.2. Trajectory Data</strong></p> <p>The remaining portion of the CSV file integrates merged data from the motion capture system and eye tracking devices, organized based on participants' helmet rigid bodies. Columns within the dataset include XYZ coordinates of all markers, spatial centroid coordinates, 6DOF orientation of the object's local coordinate frame, and <em>if available</em> eye tracking data, encompassing 2D/3D gaze coordinates, scene recording frame numbers, eye movement types, and IMU data.</p> <p>Missing data is denoted by "N/A" or an empty cell. Temporal indexing is facilitated by the "Time" or "Frame" column, indicating timestamps or frame numbers. The motion capture system records at 100Hz, Tobii Glasses at 50Hz (Raw); 25 Hz (Camera), and Pupil Glasses at 100Hz (Raw); 30 Hz (Camera). The dataset is structured around motion capture recordings, and for each rigid body, such as "Helmet_1," details per frame include XYZ coordinates of markers, centroid coordinates, and a 9-element rotational matrix describing helmet orientation.</p> <table> <tbody> <tr> <td><strong>Header</strong></td> <td><strong>Explanation</strong></td> </tr> <tr> <td>Helmet_1 - 1 X</td> <td>X-Coordinate of Marker Number 1</td> </tr> <tr> <td>Helmet_1 - 1 Y</td> <td>Y-Coordinate of Marker Number 1</td> </tr> <tr> <td>Helmet_1 - 1 Z</td> <td>Z-Coordinate of Marker Number 1</td> </tr> <tr> <td>Helmet_1 - [...]</td> <td><em>Same for Marker 2 and 3 of Helmet_1</em></td> </tr> <tr> <td>Helmet_1 Centroid_X</td> <td>X-Coordinate of the Centroid</td> </tr> <tr> <td>Helmet_1 Centroid_Y</td> <td>Y-Coordinate of the Centroid</td> </tr> <tr> <td>Helmet_1 Centroid_Z</td> <td>Z-Coordinate of the Centroid</td> </tr> <tr> <td>Helmet_1 R0</td> <td>1st Element of the CONTIGUOUS_ROTATION_MATRIX</td> </tr> <tr> <td>Helmet_1 R[..]</td> <td>Same for R1- R7</td> </tr> <tr> <td>Helmet_1 R8</td> <td>9th Element of the CONTIGUOUS_ROTATION_MATRIX</td> </tr> </tbody> </table> <p> </p> <p><strong>2.3. Eyetracking Data</strong></p> <p>The eye tracking data in the dataset includes 16 participants, providing a comprehensive dataset of over 500 minutes of recorded data across the different activities and scenarios with three different eyetracking devices. Devices are denoted with a special "Tracker_ID" in the dataset, i.e.:</p> <table> <tbody> <tr> <td><strong>Tracker ID</strong></td> <td><strong>Eyetracking Device</strong></td> </tr> <tr> <td>TB2</td> <td>Tobii 2 Glasses</td> </tr> <tr> <td>TB3</td> <td>Tobii 3 Glasses</td> </tr> <tr> <td>PPL</td> <td>Pupil Insivisible Glasses</td> </tr> </tbody> </table> <p>Gaze points are classified into fixations and saccades using the Tobii I-VT Attention filter, which is specifically optimized for dynamic scenarios with a velocity threshold of 100°. Eyetracking devices were systematically repeated after each 4-minute recording to account for natural variations in participants' eye shapes and to improve the gaze estimation algorithms. In addition, gaze estimation adjustments for the pupil invisible glasses were made after each 4-minute recording to mitigate potential drifts. It's worth noting that the scene cameras of the eye tracking glasses had different fields of view. The scene camera of the Pupil Invisible Glasses had a 1088x1080 image with both horizontal (HFOV) and vertical (VFOV) opening angles of 80°, while the Tobii Glasses provided a 1920x1080 image with different opening angles for Tobii Glasses 3 (HFOV: 95°, VFOV: 63°) and Tobii Glasses 2 (HFOV: 82°, VFOV: 52°).</p> <p><strong>NOTE AS OF 2024:</strong> <strong>Videos are NOW part</strong> of the dataset</p> <p>For one participant, wearing the Tobii Glasses 3 and Helmet_6, the data would be denoted as:</p> <table> <tbody> <tr> <td><strong>Header</strong></td> <td><strong>Explanation</strong></td> </tr> <tr> <td><em>Helmet_6 - [...]</em></td> <td><em>*X,Y,Z Coordinates for 5 markers*</em></td> </tr> <tr> <td><em>Helmet_6 [...]</em></td> <td><em>X,Y,Z Coordinates for 1 Centroid* </em></td> </tr> <tr> <td><em>Helmet_6 R[...]</em></td> <td><em>9 Elements of the CONTIGUOUS_ROTATION_MATRIX</em></td> </tr> <tr> <td> <p>Helmet_6 TB3_Accelerometer_[...]</p> </td> <td>Accelerometer data along the X,Y,Z Axis</td> </tr> <tr> <td>Helmet_6 TB3_Gyroscope_[...]</td> <td>Gyroscope data along the X,Y,Z Axis</td> </tr> <tr> <td>Helmet_6 TB3_Magnetometer_[...]</td> <td>Magnetometer data along the X,Y,Z Axis</td> </tr> <tr> <td>Helmet_6 TB3_G2D_[...]</td> <td>2D Eye tracking data (X,Y)</td> </tr> <tr> <td>Helmet_6 TB3_G3D_[...]</td> <td>3D Cyclopic Eye gaze Vector (X,Y,Z)</td> </tr> <tr> <td>Helmet_6 TB3_Movement</td> <td>Eye movement type (N/A, Fixation or Saccade)</td> </tr> <tr> <td>Helmet_6 TB3_SceneFNr</td> <td>Frame number of the scene camera recording </td> </tr> </tbody> </table> <h2>How to use and tools</h2> <p><a href="https://github.com/tmralmeida/magni-dash/tree/dash-public">magni-dash</a></p> <p><a href="https://magni-dash.streamlit.app">This</a> is a dashboard to quickly visualize our data: trajectories, speeds, eye-tracking data and LiDAR visualization (for Scenario 3). If you cannot use the dashboard from the streamlit cloud service, just run it locally by following the <a href="https://github.com/tmralmeida/magni-dash/tree/dash-public">README File</a>.</p> <p><a href="https://github.com/tmralmeida/thor-magni-tools">thor-magni-tools</a></p> <p>To install and use the package, follow the instructions on the <a href="https://github.com/tmralmeida/thor-magni-tools/blob/main/README.md">README file</a> . This package comprises:</p> <ul> <li>3D trajectory restoration: agents in the scene wore an helmet. The helmet is equipped with markers, which are tracked by the Mocap system. 3D trajectory restoration stands for <a href="https://github.com/tmralmeida/thor-magni-tools/blob/main/thor_magni_tools/preprocessing/cfg.yaml#L3">two different ways</a> of aggregating the trackings of the various markers in each helmet: (1) <em>3D-restoration</em> and (2) <em>3D-best marker</em>. The former applies an average over the locations of all visible markers while the latter uses the marker with highest tracking duration.</li> <li>3D pre-processing of restored trajectories: interpolation, downsampling and smoothing. To run the 3D pre-processing, check <a href="https://github.com/tmralmeida/thor-magni-tools?tab=readme-ov-file#preprocessing#preprocessing">this</a>.</li> <li>trajectory analysis: trajectory-related metrics like tracking duration (in seconds), number of 8s <em>tracklets</em>, motion speed, path efficiency score, and minimal distance between people. To run the trajectory analysis, check <a href="https://github.com/tmralmeida/thor-magni-tools?tab=readme-ov-file#preprocessing#analysis">this</a>.</li> </ul>
Nutrient amendment effects on phytoplankton, water chemistry, and cyanotoxins in the 2018 Large-Scale Mesocosm Experiment at the University of Kansas Field Station
This dataset includes water physicochemical parameters, phytoplankton community composition, and cyanobacteria metabolites collected during a 21-day nutrient amendment experiment conducted from 23 July to 13 August 2018 at the University of Kansas Biological Station, Lawrence, KS, United States (39.049674°N, 95.190777°W). The experiment was performed using 18 large-scale, closed-bottom fiberglass tanks (volume: 11,000 L; height: 1.25 m; diameter: 3 m). Three tanks served as ambient controls (CON), while the others received one of the following nutrient treatments: nitrogen only (280 µM) as either ammonium chloride (NH4) or sodium nitrate (NO3); nitrogen (280 µM) plus phosphorus (200 µM) as either ammonium chloride + dipotassium phosphate (NHP) or sodium nitrate + dipotassium phosphate (NOP); and phosphorus only (200 µM) as dipotassium phosphate (P). Each tank received an initial nutrient dose on Day 0.5, followed by weekly additions of 20% of the initial amendment to maintain treatment conditions. All data were quality controlled to correct basic errors and to remove measurements outside the manufacturer’s standard operational ranges.
Large-Scale Dataset for Radio Frequency based Device-Free Crowd Estimation
<p>This dataset serves to estimate the status, in particular the size, of a crowd given the impact on radio frequency communication links within a wireless sensor network. To quantify this relation, signal strengths across sub-GHz communication links are collected at the premises of the Tomorrowland music festival. The communication links are formed between the network nodes of wireless sensor networks deployed in three of the festival's stage environments. </p> <p>The table below lists the eighteen dataset files. They are collected at the music festival's 2017 and 2018 editions. There are three environments, labeled: ‘Freedom Stage 2017’, ‘Freedom Stage 2018’, and ‘Main Comfort 2018’. Each environment has both 433 MHz and 868 MHz data. The measurements at each environment were collected over a period of three festival days. The dataset files are formatted as Comma-Separated Values (CSV).</p> <pre><code class="language-markdown">| Dataset file | Reference file | Number of messages | |-------------------- |------------------------- |-------------------- | | free17_433_fri.csv | None | 393 852 | | free17_868_fri.csv | None | 472 202 | | free17_433_sat.csv | free17_transactions.csv | 996 033 | | free17_868_sat.csv | free17_transactions.csv | 1 023 059 | | free17_433_sun.csv | free17_transactions.csv | 1 007 066 | | free17_868_sun.csv | free17_transactions.csv | 1 036 456 | | free18_433_fri.csv | None | 765 024 | | free18_868_fri.csv | None | 757 657 | | free18_433_sat.csv | free18_transactions.csv | 711 438 | | free18_868_sat.csv | free18_transactions.csv | 714 390 | | free18_433_sun.csv | free18_transactions.csv | 648 329 | | free18_868_sun.csv | free18_transactions.csv | 656 290 | | main18_433_fri.csv | None | 791 462 | | main18_868_fri.csv | None | 908 407 | | main18_433_sat.csv | main18_counts.csv | 863 666 | | main18_868_sat.csv | main18_counts.csv | 884 682 | | main18_433_sun.csv | main18_counts.csv | 903 862 | | main18_868_sun.csv | main18_counts.csv | 894 496 |</code></pre> <p>In addition to the datasets and reference files, a software example is provided to illustrate the data use and visualise the initial findings and relation between crowd size and network signal strength impact.</p> <p>In order to use the software, please retain the following file structure: </p> <pre><code class="language-markdown">. ├── data ├── data_reference ├── graphs └── software</code></pre> <p>The peer-reviewed data descriptor for this dataset has now been published in MDPI Data - an open access journal aiming at enhancing data transparency and reusability, and can be accessed here: <a href="https://doi.org/10.3390/data5020052">https://doi.org/10.3390/data5020052</a>.<br> Please cite this when using the dataset.</p>
Data from paper: "Large-scale variations in the dynamics of Amazon forest canopy gaps from airborne lidar data and opportunities for tree mortality estimates"
<p>Data from the paper:</p> <p>Dalagnol, R. <em>et al.</em> Large-scale variations in the dynamics of Amazon forest canopy gaps from airborne lidar data and opportunities for tree mortality estimates. <em>Sci Rep</em> <strong>11, </strong>1388 (2021). https://doi.org/10.1038/s41598-020-80809-w</p> <p>Link: https://www.nature.com/articles/s41598-020-80809-w</p> <p> </p> <p>This repository contains:</p> <p>1) Data frame with data from static and dynamic gaps used in Figure 2 (Dalagnol_2020_Data_Multitemporal_gaps.csv). Each row is the aggregated measurement at 5-km resolution. The site component referes to the five site studied with multitemporal data. Site order from 1 to 5 is DUC, TAP, FN1, BON and TAL.</p> <p>2) Data frame with data from static gaps and environmental factors used in Table 1, Figure 3, 4, 5 (Dalagnol_2020_Data_Singledate_gaps_Modeling.csv). Each row is the aggregated measurement of one site observed by airborne lidar data.</p> <p>3) Raster file at 5-km resolution with dynamic gap fraction estimates presented in Figure 5 (dynamic_gap_fraction_amazon.tif).</p> <p> </p> <p>If you need anything else, please contact the corresponding author: Ricardo Dalagnol (ricds@hotmail.com).</p>
Exploring Large-Scale Entanglement in Quantum Simulation
<p>Here we provide data for the manuscript " <a href="https://arxiv.org/abs/2306.00057">Exploring Large-Scale Entanglement in Quantum Simulation</a> " with arXiv id <a href="https://arxiv.org/abs/2306.00057">"arXiv:2306.00057</a>". The data set contains both raw and analyzed data saved as ".mat files" Please see the uploaded readme file to understand the data structure. The peer-reviewed article will appear in the future. Please check the published article for recent figures. </p>
Dataset for the paper "Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset"
<p>We present a large-scale anomaly detection dataset collected from IBM Cloud's Console over approximately 4.5 months. This high-dimensional dataset captures telemetry data from multiple data centers, specifically designed to aid researchers in developing and benchmarking anomaly detection methods in large-scale cloud environments. It contains 39,365 entries, each representing a 5-minute interval, with 117,448 features/attributes, as interval_start is used as the index. The dataset includes detailed information on request counts, HTTP response codes, and various aggregated statistics. The dataset also includes labeled anomaly events identified through IBM's internal monitoring tools, providing a comprehensive resource for real-world anomaly detection research and evaluation.</p> <p><strong>File Descriptions</strong></p> <ul> <li><code>location_downtime.csv</code> - Details planned and unplanned downtimes for IBM Cloud data centers, including start and end times in ISO 8601 format.</li> <li><code>unpivoted_data.parquet</code> - Contains raw telemetry data with 413 million+ rows, covering details like location, HTTP status codes, request types, and aggregated statistics (min, max, median response times).</li> <li><code>anomaly_windows.csv</code> - Ground truth for anomalies, listing start and end times of recorded anomalies, categorized by source (Issue Tracker, Instant Messenger, Test Log).</li> <li><code>pivoted_data_all.parquet</code> - Pivoted version of the telemetry dataset with 39,365 rows and 117,449 columns, including aggregated statistics across multiple metrics and intervals.</li> <li><code>demo/demo.[ipynb|html]</code>: This demo file provides examples of how to access data in the Parquet files, available in Jupyter Notebook (<code>.ipynb</code>) and HTML (<code>.html</code>) formats, respectively.</li> </ul> <p>Further details of the dataset can be found in <strong>Appendix B: Dataset Characteristics</strong> of the <a href="https://arxiv.org/abs/2411.09047">paper</a> titled <strong><em>"Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset."</em></strong> Sample code for training anomaly detectors using this data is provided in <a href="https://doi.org/10.5281/zenodo.14598119" target="_blank" rel="noopener">this package</a>.</p> <p> </p> <p>When using the dataset, please cite it as follows:</p> <pre><code>@misc{islam2024anomaly,</code><br><code> title={Anomaly Detection in Large-Scale Cloud Systems: An Industry Case and Dataset}, </code><br><code> author={Mohammad Saiful Islam and Mohamed Sami Rakha and William Pourmajidi and Janakan Sivaloganathan and John Steinbacher and Andriy Miranskyy},</code><br><code> year={2024},</code><br><code> eprint={2411.09047},</code><br><code> archivePrefix={arXiv},</code><br><code> url={https://arxiv.org/abs/2411.09047}</code><br><code>}</code></pre> <p> </p>
Massive IoT for Large-Scale Public Events in the 5GENESIS Surrey Platform
<p>This dataset contains the results of the trials conducted within the context of the main IoT use case of the 5GENESIS Surrey Platform.</p>
Supplementary data: Accurate large-scale simulations of siliceous zeolites by neural network potentials
<p><strong>Content</strong></p> <p><em>1. Zeolite databases</em></p> <ul> <li>Deem database containing 331170 hypothetical zeolite frameworks [Deem09, Pophale11] geometrically optimized at the NNPscan level (note, the first row of the database is alpha-quartz): "DEEM_NNPscan.db"</li> <li>Database of 236 exiting zeolite frameworks of the <a href="http://www.iza-structure.org/databases/">International Zeolite Association (IZA) </a>optimized at the NNPscan level: "IZA_NNPscan.db"</li> <li>Both databases are <a href="https://wiki.fysik.dtu.dk/ase/ase/db/db.html">ASE SQLite database files</a> of the <a href="https://wiki.fysik.dtu.dk/ase/index.html">Atomic Simulation Environment</a> containing the ASE <a href="https://wiki.fysik.dtu.dk/ase/ase/atoms.html">Atoms objects</a> with energies and forces (NNPscan level); readable with ASE's <a href="https://wiki.fysik.dtu.dk/ase/ase/io/io.html">I/O module</a></li> <li>Additionally, relevant quantities can be extracted with, e.g., the following queries (further information: ase db --help):</li> </ul> <pre><code class="language-bash">ase db DEEM_NNPscan.db -c id,formula,natoms,volume,mass,density,energy_per_tsite,n_tsites,relative_energy # Output id|formula|natoms| volume| mass|density|energy_per_tsite|n_tsites|relative_energy 1|O6Si3 | 9|111.161|180.249| 26.988| -31.796| 3| 0.000 2|O16Si8 | 24|433.858|480.664| 18.439| -31.638| 8| 15.265 3|O16Si8 | 24|421.114|480.664| 18.997| -31.596| 8| 19.359 4|O16Si8 | 24|426.557|480.664| 18.755| -31.614| 8| 17.613 5|O16Si8 | 24|412.410|480.664| 19.398| -31.613| 8| 17.677 6|O16Si8 | 24|393.544|480.664| 20.328| -31.594| 8| 19.546 7|O16Si8 | 24|422.400|480.664| 18.939| -31.657| 8| 13.476 8|O16Si8 | 24|394.405|480.664| 20.284| -31.581| 8| 20.797 9|O12Si6 | 18|265.201|360.498| 22.624| -31.611| 6| 17.868 10|O16Si8 | 24|357.047|480.664| 22.406| -31.581| 8| 20.785 11|O16Si8 | 24|434.894|480.664| 18.395| -31.621| 8| 16.911 12|O16Si8 | 24|384.158|480.664| 20.825| -31.657| 8| 13.448 13|O12Si6 | 18|258.977|360.498| 23.168| -31.679| 6| 11.278 14|O16Si8 | 24|466.429|480.664| 17.152| -31.593| 8| 19.588 15|O16Si8 | 24|423.469|480.664| 18.892| -31.639| 8| 15.179 16|O16Si8 | 24|450.716|480.664| 17.750| -31.628| 8| 16.219 17|O16Si8 | 24|331.528|480.664| 24.131| -31.642| 8| 14.857 18|O16Si8 | 24|458.573|480.664| 17.445| -31.635| 8| 15.572 19|O16Si8 | 24|359.298|480.664| 22.266| -31.655| 8| 13.636 20|O16Si8 | 24|464.264|480.664| 17.232| -31.612| 8| 17.750 Rows: 331171 (showing first 20) Keys: density, energy_per_tsite, n_tsites, relative_energy ase db IZA_NNPscan.db -c id,formula,natoms,volume,mass,density,energy_per_tsite,n_tsites,relative_energy,iza_code # Output id|formula |natoms| volume| mass|density|energy_per_tsite|n_tsites|relative_energy|iza_code 1|O16Si8 | 24| 435.488| 480.664| 18.370| -31.676| 8| 11.594|ABW 2|O32Si16 | 48| 961.419| 961.328| 16.642| -31.645| 16| 14.612|ACO 3|O96Si48 | 144|3154.579|2883.984| 15.216| -31.664| 48| 12.810|AEI 4|O80Si40 | 120|2102.921|2403.320| 19.021| -31.703| 40| 9.021|AEL 5|O96Si48 | 144|2417.286|2883.984| 19.857| -31.666| 48| 12.586|AEN 6|O144Si72| 216|4075.300|4325.976| 17.667| -31.674| 72| 11.831|AET 7|O96Si48 | 144|2786.810|2883.984| 17.224| -31.675| 48| 11.716|AFG 8|O48Si24 | 72|1400.247|1441.992| 17.140| -31.690| 24| 10.268|AFI 9|O64Si32 | 96|1764.823|1922.656| 18.132| -31.653| 32| 13.809|AFN 10|O80Si40 | 120|2080.330|2403.320| 19.228| -31.707| 40| 8.632|AFO 11|O64Si32 | 96|2097.384|1922.656| 15.257| -31.655| 32| 13.622|AFR 12|O112Si56| 168|3820.116|3364.648| 14.659| -31.650| 56| 14.150|AFS 13|O144Si72| 216|4732.720|4325.976| 15.213| -31.664| 72| 12.793|AFT 14|O60Si30 | 90|1897.074|1802.490| 15.814| -31.659| 30| 13.268|AFV 15|O96Si48 | 144|3154.885|2883.984| 15.214| -31.664| 48| 12.776|AFX 16|O32Si16 | 48|1137.335| 961.328| 14.068| -31.591| 16| 19.790|AFY 17|O48Si24 | 72|1283.812|1441.992| 18.694| -31.620| 24| 17.034|AHT 18|O96Si48 | 144|2479.287|2883.984| 19.360| -31.681| 48| 11.155|ANA 19|O64Si32 | 96|1797.086|1922.656| 17.807| -31.662| 32| 12.924|APC 20|O64Si32 | 96|1751.393|1922.656| 18.271| -31.678| 32| 11.422|APD Rows: 236 (showing first 20) Keys: density, energy_per_tsite, iza_code, n_tsites, relative_energy # Filtering of the database, e.g., for structures with relative energies < 10 kJ/(mol Si) ase db IZA_NNPscan.db relative_energy\<10 -c density,energy_per_tsite,n_tsites,relative_energy,iza_code # Output density|energy_per_tsite|n_tsites|relative_energy|iza_code 19.021| -31.703| 40| 9.021|AEL 19.228| -31.707| 40| 8.632|AFO 19.385| -31.695| 24| 9.802|ATV 18.778| -31.702| 34| 9.061|DOH 19.570| -31.693| 24| 9.959|EWO 18.401| -31.698| 32| 9.451|GON 18.551| -31.695| 112| 9.807|IHW 17.778| -31.693| 288| 9.972|IMF 19.154| -31.695| 6| 9.762|JBW 18.187| -31.695| 96| 9.734|MFI 19.278| -31.709| 48| 8.443|MRE 18.035| -31.698| 90| 9.481|MSO 20.417| -31.724| 44| 7.003|MTF 19.227| -31.704| 136| 8.898|MTN 18.542| -31.693| 28| 9.966|MTW 19.137| -31.695| 60| 9.798|PCR 20.037| -31.709| 144| 8.464|PSI 18.843| -31.703| 64| 9.004|SAF 18.371| -31.703| 112| 8.975|STO 19.894| -31.706| 17| 8.671|VET Rows: 20 (showing first 20) Keys: density, energy_per_tsite, iza_code, n_tsites, relative_energy</code></pre> <ul> <li>The quantities shown above are available with the keys (besides standard ASE database keys):</li> </ul> <table> <thead> <tr> <th scope="col">Key</th> <th scope="col">Quantity</th> <th scope="col">Unit</th> </tr> </thead> <tbody> <tr> <td>id</td> <td>Identifier</td> <td> </td> </tr> <tr> <td>formula</td> <td>Chemical formula of the unit cell</td> <td> </td> </tr> <tr> <td>natoms</td> <td>Number of atoms</td> <td> </td> </tr> <tr> <td>volume</td> <td>Unti cell volume</td> <td>Å<sup>3</sup></td> </tr> <tr> <td>mass</td> <td>Atomic mass of the unit cell</td> <td>amu</td> </tr> <tr> <td>density</td> <td>Framework density</td> <td>Si/nm<sup>3</sup></td> </tr> <tr> <td>energy_per_tsite</td> <td>NNPscan energy</td> <td>eV</td> </tr> <tr> <td>n_tsites</td> <td>Number of T-sites</td> <td> </td> </tr> <tr> <td>relative_energy</td> <td>Energy with respect to quartz</td> <td>kJ/(mol Si)</td> </tr> <tr> <td>iza_code</td> <td>only for 'IZA_NNPscan.db'</td> <td> </td> </tr> </tbody> </table> <ul> <li> Comma separated csv files for the quantities listed above: "DEEM_NNPscan.csv" and "IZA_NNPscan.csv"</li> </ul> <p><em>2. Neural network potentials (NNP) for silica</em></p> <ul> <li>SchNet [Schütt18,Schütt19] NNP files trained on DFT data at the PBE+D3 (NNPpbe) and SCAN+D3 level (NNPscan)</li> <li>Simulations can be performed using <a href="https://schnetpack.readthedocs.io/en/stable/getstarted/getstarted.html#references">SchNetPack</a> with its ASE calculator</li> <li>This example shows a simple single-point calculation</li> </ul> <pre><code class="language-python">import ase.io import torch from schnetpack.interfaces import SpkCalculator from schnetpack.environment import AseEnvironmentProvider # check if GPU(s) are available if torch.cuda.is_available(): device = "cuda" else: device = "cpu" # load the NNP model model = torch.load('SiOscan1', map_location=device) # read some structure atoms = ase.io.read( ... ) # define SchNetPack calculator calc = SpkCalculator(model=model, device=device, energy='energy', forces='forces', environment_provider=AseEnvironmentProvider(6.) ) # attach calculator to atoms object atoms.set_calculator(calc) # perform simulations, e.g., single-point calculation energy = atoms.get_potential_energy() print(energy)</code></pre> <p><em>3. Test set used for accuracy evaluation (ASE database: test_set_NNPscan.db)</em></p>
Marine plastics alter the organic matter composition of the air-sea boundary layer, with influences on CO2 exchange: a large-scale analysis method to explore future ocean scenarios
<p>Microplastics are substrates for microbial activity and can influence biomass production. This has potentially important implications in the sea-surface microlayer, the marine boundary layer that controls gas exchange with the atmosphere and where biologically produced organic compounds can accumulate. In the present study, we used six large scale mesocosms to simulate future ocean scenarios of high plastic concentration. Each mesocosm was filled with 3 m3 of seawater from the oligotrophic Sea of Crete, in the Eastern Mediterranean Sea. A known amount of standard polystyrene microbeads of 30 μm diameter was added to three replicate mesocosms, while maintaining the remaining three as plastic-free controls. Over the course of a 12-day experiment, we explored microbial organic matter dynamics in the sea-surface microlayer in the presence and absence of microplastic contamination of the underlying water. Our study shows that microplastics increased both biomass production and enrichment of carbohydrate-like and proteinaceous marine gel compounds in the sea-surface microlayer. Importantly, this resulted in a 3 % reduction in the concentration of dissolved CO2 in the underlying water. This reduction was associated to both direct and indirect impacts of microplastic pollution on the uptake of CO2 within the marine carbon cycle, by modifying the biogenic composition of the sea's boundary layer with the atmosphere.</p>
Supplementary table 1 for 'Imperial timber? Dendrochronological evidence for large-scale road building along the Roman limes in the Netherlands' (2015)
<p>This supplementary table to Visser(2015) was not openly available. This dataset provides the supplementary table in the open ODS-format and also as XLS and CSV.</p> <div> <div>Publication: Visser, RM. 2015 Imperial timber? Dendrochronological evidence for large-scale road building along the Roman limes in the Netherlands. <em>Journal of Archaeological Science</em> 53: 243–254. DOI: <a href="https://doi.org/10.1016/j.jas.2014.10.017">https://doi.org/10.1016/j.jas.2014.10.017</a>.</div> </div>
Road-deposited sediment wash-off experiments on a large-scale laboratory
<p><span>This dataset includes raw and processed data from a series of large-scale laboratory tests that were conducted to assess and study the wash-off process of RDS (Road deposited sediments) considering variations in rainfall intensity, for two scenarios: 30 mm/h and 50 mm/h; and modifying RDS loads applied on BLOCK for three scenarios: 100 g/m<sup>2</sup>, 150 g/m<sup>2</sup>, 200 g/m<sup>2</sup>. First, the hydraulic was detailly characterized including rainfall intensity maps, water flows, surface water depths and surface water velocities for both rainfall intensities tested. A synthetic granulometric of RDS was homogeneously distributed on the physical model surface and then washed-off by the simulated rainfall. A total of 31 water samples were collected at the manhole discharge per each experiment. Total RDS mass that remain on the surface and inside the gully was collected by a wet vacuum after the rainfall event. A mass balance considering the initial RDS applied and the total RDS recollected in the three samples locations, was calculated. TUR (Turbidity), EC (Conductivity), TS (Total Solids), TSS (Total suspended solids), TDS (Total dissolved solids), RDS mass by flow, and RDS mass flow variables were measured for the RDS samples recollected in the Manhole discharge. The behaviour of each RDS fraction was also analysed through laser diffraction (</span>Beckman Coulter LS 13 320, Aqueous Liquids Module<span>). This work is part of a Transnational Access developed by the Universidad Distrital Francisco José de Caldas (Colombia) and Universidade da Coruña (Spain) within the scope of Co-UDlabs project. Data may be used to increase knowledge on road-deposited sediment wash-off process, allowing also for calibrating, developing, and validating new and existing urban wash-off models.</span></p>
TBPos: Dataset for Large-Scale Precision Visual Localization (database files)
<p>Large-scale dataset for visual localization, provided in the format of the well-known InLoc dataset (Taira et al, 2018). Contains co-registered RGB point clouds and a script for generating the rest of the 'database' files for visual localization by the InLoc algorithm. Note: query images are provided in a separate repository.</p>
EMBERSim: A Large-Scale Databank for Boosting Similarity Search in Malware Analysis
<p>In recent years there has been a shift from heuristics-based malware detection towards machine learning, which proves to be more robust in the current heavily adversarial threat landscape. While we acknowledge machine learning to be better equipped to mine for patterns in the increasingly high amounts of similar-looking files, we also note a remarkable scarcity of the data available for similarity-targeted research. Moreover, we observe that the focus in the few related works falls on quantifying similarity in malware, often overlooking the clean data. This one-sided quantification is especially dangerous in the context of detection bypass. We propose to address the deficiencies in the space of similarity research on binary files, starting from EMBER — one of the largest malware classification data sets. We enhance EMBER with similarity information as well as malware class tags, to enable further research in the similarity space. Our contribution is threefold: (1) we publish EMBERSim, an augmented version of EMBER, that includes similarity-informed tags; (2) we enrich EMBERSim with automatically determined malware class tags using the open-source tool AVClass on VirusTotal data and (3) we describe and share the implementation for our class scoring technique and leaf similarity method.</p>
Data from a Large-Scale Experiment to Evaluate the Effects of Trapping to Control Muskrats (Ondatra zibethicus) in The Netherlands
<p>This data set supports the publication 'A Large-Scale Experiment to Evaluate the Effects of Trapping to Control Muskrats (Ondatra zibethicus) in The Netherlands' by Daan Bos, Emiel van Loon, Erik Klop and Ron Ydenberg. (the paper was accepted for publication in Wildlife Society Bulletin in 2020)</p> <p>The Muskrat is an invasive species in Europe and in the Netherlands muskrat burrowing can compromise the integrity of dykes and hence poses a public safety threat. For that reason a control programme has been in effect since the arrival of the species in 1941. To investigate the relation between catch and effort and enhance prediction models, a large randomized controlled experiment was designed and conducted from 2013 till 2016. The publication by Bos et al. (2020) analyses the experimental results and here we present and document the experimental data. See the readme.md file for further information.</p>
Replication Data & Code - Large-scale land acquisitions exacerbate local land inequalities in Tanzania
<h3><strong>Reference</strong></h3><p>Sullivan J.A., Samii, C., Brown, D., Moyo, F., Agrawal, A. 2023. Large-scale land acquisitions exacerbate local farmland inequalities in Tanzania. Proceedings of the National Academy of Sciences 120, e2207398120. <a href="https://doi.org/10.1073/pnas.2207398120">https://doi.org/10.1073/pnas.2207398120</a> </p><h3><strong>Abstract</strong></h3><p>Land inequality stalls economic development, entrenches poverty, and is associated with environmental degradation. Yet, rigorous assessments of land-use interventions attend to inequality only rarely. A land inequality lens is especially important to understand how recent large-scale land acquisitions (LSLAs) affect smallholder and indigenous communities across as much as 100 million hectares around the world. This paper studies inequalities in land assets, specifically landholdings and farm size, to derive insights into the distributional outcomes of LSLAs. Using a household survey covering four pairs of land acquisition and control sites in Tanzania, we use a quasi-experimental design to characterize changes in land inequality and subsequent impacts on well-being. We find convincing evidence that LSLAs in Tanzania lead to both reduced landholdings and greater farmland inequality among smallholders. Households in proximity to LSLAs are associated with 21.1% (<i>P</i> = 0.02) smaller landholdings while evidence, although insignificant, is suggestive that farm sizes are also declining. Aggregate estimates, however, hide that households in the bottom quartiles of farm size suffer the brunt of landlessness and land loss induced by LSLAs that combine to generate greater farmland inequality. Additional analyses find that land inequality is not offset by improvements in other livelihood dimensions, rather farm size decreases among households near LSLAs are associated with no income improvements, lower wealth, increased poverty, and higher food insecurity. The results demonstrate that without explicit consideration of distributional outcomes, land-use policies can systematically reinforce existing inequalities.</p><h3><strong>Replication Data</strong></h3><p>We include anonymized household survey data from our analysis to support open and reproducible science. In particular, we provide i) an anoymized household dataset collected in 2018 (n=994) for households nearby (treatment) and far-away from (control) LSLAs and ii) a household dataset collected in 2019 (n=165) within the same sites. For the 2018 surveys, several anonymized extracts are provided including an imputed (n=10) dataset to fill in missing data that was used for the main analysis. This data can be found in the <i>hh_data</i> folder and includes:</p><ul><li><i>hh_imputed10_2018:</i> anonymized household dataset for 2018 with variables used for the main analysis where missing data was imputed 10 times</li><li><i>hh_compensation_2018:</i> anonymized household extract for 2018 representing household benefits and compensation directly received from LSLAs</li><li><i>hh_migration_2018:</i> anonymized household extract for 2018 representing household migration behavior following LSLAs</li><li><i>hh_rsdata_2018:</i> extracted remote sensing data at the household geo-location for 2018</li><li><i>hh_land_<strong>2019</strong>:</i><strong> </strong> anonymized household extract for <strong>2019 </strong>of land variables</li></ul><p>Our analysis also incorporates data from the Living Standards Measurement Survey (LSMS) collected by the World Bank (found in <i>lsms_data</i> folder). We've provide sub-modules from the LSMS dataset relevant to our analysis but the full datasets can be access through the World Bank's Microdata Library (https://microdata.worldbank.org/index.php/home). </p><p>Across several analyses we use the LSLA boundaries for our four selected sites. We provide a shapefile for the LSLA boundaries in the <i>gis_data</i> folder.</p><p>Finally, our data replication includes several model outputs (found in <i>mod_outputs)</i>, particularly those that are lengthy to run in R. These datasets can optionally be loaded into R rather than re-running analysis using our <i>main_analysis.Rmd</i> script. </p><h3><strong>Replication Code</strong></h3><p>We provide replication code in the form of R Markdown (.Rmd) or R (.R) files. Alongside the replication data, this can be used to reproduce main figures, table, supplementary materials, and results reported in our article. Scripts include:</p><ul><li><i>main_analysis.Rmd:</i> main analysis supporting the finding, graphs, and tables reported in our main manuscript</li><li><i>compensation.R:</i> analysis of benefits and compensation received directly by households from LSLAs</li><li><i>landvalue.R:</i> analysis of household land values as a function of distance from LSLAs</li><li><i>migration.R:</i> analysis of migration behavior following LSLAs</li><li><i>selection_bias.R:</i> analysis of LSLA selection bias between control and treatment enumeration areas</li></ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.