Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,146
datasets available to search
ShareScore release 0.9.0
Dataset results
13,146 results for “Integration”
Data on the Digital Economy and Society Index (DESI), the ASEAN Digital Integration Index (ADII), and the Digital Intelligence Index (DII)
<p>This dataset contains the quantitative measurement of the Digital Economy and Society Index (DESI), the ASEAN Digital Integration Index (ADII), and the Digital Intelligence Index (DII) in 2019.</p>
An Integrated Usability Framework for Evaluating Open Government Data Portals and Analysis of EU and GCC OGD Portals
<p><span>This dataset contains data collected during a study (<em><strong>"<a href="https://arxiv.org/ftp/arxiv/papers/2403/2403.08451.pdf">An Integrated Usability Framework for Evaluating Open Government Data Portals: Comparative Analysis of EU and GCC Countries</a>"</strong></em>) conducted by Fillip Molodtsov and Anastasija Nikiforova (University of Tartu).</span></p> <p><span> </span><span>It being made public both to act as supplementary data for the paper and in order for other researchers to use these data in their own work potentially contributing to the improvement of current data ecosystems and develop user-friendly, collaborative, robust, and sustainable open data portals.</span></p> <p><span>***Purpose of the study***</span></p> <p><span>This paper develops an integrated framework for evaluating OGD portal effectiveness that accommodates user diversity (regardless of their data literacy and language), evaluates collaboration and participation, and the ability of users to explore and understand the data provided through them. </span></p> <p><span>The framework is validated by applying it to 33 national portals across European Union (EU) and Gulf Cooperation Council (GCC) countries, as a result of which we rank OGD portals, identify some good practices that lower-performing portals can learn from, and common shortcomings.</span></p> <p><span>***Methodology***</span></p> <p><span>(1) systematic literature review to establish a knowledge base and identify frameworks have been used to evaluate OGD portals, we conducted a systematic literature review - Dataset_ Usability_Framework_SLR;</span></p> <p><span>(2) development of the Integrated Usability Framework for Evaluating Open Government Data Portals, which content is based on the outputs of the first step, along with selected articles of experts in portal design, and an exploratory assessment of the French, Irish, Estonian and Spanish portals - Dataset_Integrated_Usability_Framework;</span></p> <p><span>(3) data collection, that is a completion of the protocol developed in the previous step by analysing 34 national OGD portals of the EU and GCC countries. When all individual protocols were collected, the total score are calculated using the weighting system. The average scores are calculated for the EU and GCC. The portals are ranked. The top portals (best performers) are determined for each dimension - Dataset_EU_GCC_OGDportal_Usability_results_clustering.</span></p> <p><span>(4) identification of relationships and patterns among different portals based on their performance metrics as a result of the cluster analysis. By calculating the average dimensional scores of portals from both types of clusters, their performance across multiple dimensions is evaluated - Dataset_EU_GCC_OGDportal_Usability_results_clustering.</span></p> <p> </p> <p><strong><em><span>For more details see Molodtsov, F., Nikiforova, A. (2024). “An Integrated Usability Framework for Evaluating Open Government Data Portals: Comparative Analysis of EU and GCC Countries”. In Proceedings of the 25th Annual International Conference on Digital Government Research (DGO 2024), June 11--14, 2024, Taipei, Taiwan, 10.1145/3657054.3657159</span></em></strong></p> <p><span>***Format of the file***</span></p> <p><span>.xls, .csv</span></p> <p><span>***Licenses or restrictions***</span></p> <p><span>CC-BY</span></p>
Data for "emIAM v1.0: an emulator for Integrated Assessment Models using marginal abatement cost curves"
<p>This dataset contains codes, data, tables, andd figures (high resolution) related to the following publication: Xiong, W., K. Tanaka, P. Ciais, D. J. A. Johansson, M. Lehtveer (2022) emIAM v1.0: an emulator for Integrated Assessment Models using marginal abatement cost curves. Submitted to arXiv on 23 December 2022.</p>
Integrated Datasets for analyses on potentially hazardous locations for women in Valencia, Dublin, San Francisco, and Toluca
<p>This dataset provides a compilation of the data used to analyze and identify potentially dangerous<br>places for women. Multiple data collection techniques, including official data downloads, web<br>scraping, and participatory mapping, were combined for integration, applying specific processing.<br>The datasets refer to four cities: Valencia (Spain), Dublin (Ireland), San Francisco (United States),<br>and Toluca (Mexico).<br>Depending on the availability and context of each city, the datasets are classified into three<br>categories: DATA, TWT, and MAP. The DATA prefix refers to files containing the results of the<br>analysis of socioeconomic variables downloaded from official sources; for the mapping, the<br>standard territorial unit was a 25x25 m grid for Valencia and 50x50 m for Dublin and San Francisco.<br>The files with the prefix TWT are composed of datasets containing tweets collected through web<br>scraping and analyzed using natural language processing (NLP) algorithms and neural networks;<br>the purpose is to identify and classify tweets related to gender violence, feelings of fear, or<br>perceptions of insecurity. For MAP files, participants gathered them through participatory<br>mapping processes, using specific calls to public space users and a supporting web application<br>designed for this purpose. The files with the prefix POL contain datasets used for crime prediction based on crime density for the city of Valencia.</p>
Single aerosol measurements from a wideband integrated bioaerosol sensor, collected during the Antarctic Circumnavigation Expedition (ACE).
<p><strong>Dataset abstract</strong></p> <p>This data set contains the time series of single particle data, measured by the wideband integrated bioaerosol sensor (WIBS-4, University of Hertfordshire, Hatfield, UK), during the Antarctic Circumnavigation Expedition (ACE), which was conducted between 20th of December 2016 and 19th of March 2017. WIBS provides aerosol optical diameter (5 μm - 14 μm), asymmetry factor and fluorescent signals on three different channels. WIBS measures single aerosol particles at a sampling rate of 125 Hz. More technical details about WIBS could be found in Kaye et al. (2005).</p> <p><strong>Dataset contents</strong></p> <ul> <li>part_1_Cape_Town_Kerguelen.csv, data file, comma-separated values</li> <li>part_2_Kerguelen_Hobart.csv, data file, comma-separated values</li> <li>part_3_Hobart_Mertz.csv, data file, comma-separated values</li> <li>part_4_Mertz_Punta Arenas.csv, data file, comma-separated values</li> <li>part_5_Punta_Arenas_Cape_Town.csv, data file, comma-separated value</li> <li>part_6_Cape_Town_Bremerhaven.csv, data file, comma-separated values</li> <li>data_file_header.txt, metadata, text</li> <li>README.txt, metadata, text</li> </ul> <p><strong>Dataset license</strong></p> <p>This aerosol measurement dataset collected using a WIBS during ACE is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
Sensor deployment to support the integrated energy management system in residential buildings in ReCO2ST LoRa Dataset
<p>LoRa Radio Testing Datasets for preliminary performance tests. These datasets were taken in order to ensure that the LoRa radios were capable of transmitting through concrete and testing various preamble settings of the radio. As per the paper,</p> <p>"Although these testing methodologies were indicative but not exact or perfect, to test in a manner that was qualitative would have been both costly and beyond the scope of the project." </p> <p>These tests were to help us verify feasibility of the chosen LoRa Radio</p>
Example of datasets processed to demonstrate a multisource data integration methodology
<p>This dataset contains the data processed to demonstrate the multi-source spatial data integration methodology proposed in the paper "Multisource spatial data integration for use cases applications".</p> <p>It contains:</p> <p>- the building footprint extracted from the IFC model of a newly designed building in WKT format, by using the GeoBIM_Tool (<a href="https://github.com/twut/GEOBIM_Tool">https://github.com/twut/GEOBIM_Tool</a>);</p> <p>- the extrusion of the footprint until the measured height measured with the same GeoBIM_Tool;</p> <p>- a portion of the Rotterdam 3D city model generated with 3dfier and available at https://3d.bk.tudelft.nl/opendata/3dfier/, converted in CityJSON with the citygml-tools (https://www.cityjson.org/tutorials/conversion/), developed to convert data between CityGML and CityJSON.</p>
Dataset for the study Late development of audio-visual integration in the vertical plane
<p>It is not clear how multisensory skills develop and how visual experience impacts on multisensory spatial development. Conflicting results show that visual calibration precedes multisensory integration for the audio-visual spatial bisection task (Gori et al., 2012a, 2012b) while in other tasks such as spatial localization, visual calibration occurs after multisensory development (Rohlf et al., 2020). Results in blind individuals can say something about the role of vision on perceptual development. Scientific evidences show that blind individuals have impairments in bisecting the auditory space (Gori et al., 2014) but not in localizing auditory sources (Lessard et al., 1998). Such results suggest that sensory calibration and impairment are linked. We studied the development of audio-visual multisensory localization in the vertical plane in sighted individuals from 5 years to adulthood to address this hypothesis. We hypothesize that typical children would show late audio-visual integration for the vertical plane, preceded by visual dominance. Unimodal and bimodal audio-visual thresholds and PSEs were measured and compared with the Bayesian optimal-integration model (maximum likelihood estimation). Results show that the development of multisensory integration in the vertical plane is not evident at 5 years, suggesting visual dominance for vertical audio-visual localization. These results support the idea that multisensory perception in the vertical domain depends on sensory calibration. We discuss these scientific results proposing that the process of cross-sensory calibration is task-specific and highlighting the importance of linking the impairment and development to better determine how our brain works.</p> <p>Data are in textual tab delimited format. Columns report for each subject: age, age_bin, condition, jnd.</p> <p> </p>
Dataset of "Towards Artefact Aware Human Motion Capture using Inertial Sensors Integrated into Loose Clothing"
<p>This dataset was used in the publication:<br> <strong>Towards Artefact Aware Human Motion Capture using Inertial Sensors Integrated into Loose Clothing</strong><br> presented at the IEEE International Conference on Robotics and Automation 2022</p> <p><strong>Abstract:</strong><br> Inertial motion capture has become an attractive alternative to optical motion capture for human joint angle estimation outside the laboratory. Usually inertial sensors are assumed to be tightly fixed to the body segments, which can be cumbersome regarding setup-time and ease-of-use. However, integrating the sensors directly into clothing, usually, results in additional clothing motion relative to the motion of the underlying bones that should be captured.<br> In this work we propose the <em>Difference Mapping</em> distributions approach that corrects the segment orientations of a given inertial motion capture system that assumes tightly coupled sensors.<br> The approach allows to reduce the joint angle errors due to clothing artefacts by at least 77.2 percent for people with similar morphology performing a similar task as seen in the training data, including an ergonomic assessments scenario at work places with 10 participants. <br> Moreover, we show that the uncertainty of the distribution can be used to measure the reliability of the predicted map if e.g. the motion is further away from the training data to allow for an artefact aware inertial motion tracking approach.<br> The experimental data for this study is available online</p> <p> </p> <p><strong>Data structure:</strong><br> The data contains trials of 12 subjects for different motions, wearing at the same time a tight setup with inertial sensors and a loose working suit with integrated inertial sensors. It contains the raw IMU data, raw Magnetometer data and the estimated segment orientations using a Sensor Fusion engine provided by Sci-Track.<br> Please note, that in the publication only the first 10 subjects were used and the upper body information was used only. The Sternum sensor of the tight setup of subjects 11, 12 and 13 tilted slowly during the long-term measurements. For this reason only 10 subjects were included in the study. However all remaining sensor of the tight setup were not tilted during recording. In particular the lower body recordings of all subjects are not corrupted.<br> <br> Code samples, a visualizer and further useful information is provided under the following git repository:<br> https://github.com/lorenzcsunikl/Dataset-of-Artefact-Aware-Human-Motion-Capture-using-Inertial-Sensors-Integrated-into-Loose-Clothing</p>
Multi-aspect Integrated Migration Indicators (MIMI) dataset
<p>Nowadays, new branches of research are proposing the use of non-traditional data sources for the study of migration trends in order to find an original methodology to answer open questions about cross-border human mobility. The Multi-aspect Integrated Migration Indicators (MIMI) dataset is a new dataset to be exploited in migration studies as a concrete example of this new approach. It includes both official data about bidirectional human migration (traditional flow and stock data) with multidisciplinary variables and original indicators, including economic, demographic, cultural and geographic indicators, together with the Facebook Social Connectedness Index (SCI). It is built by gathering, embedding and integrating traditional and novel variables, resulting in this new multidisciplinary dataset that could significantly contribute to nowcast/forecast bilateral migration trends and migration drivers.</p> <p>Thanks to this variety of knowledge, experts from several research fields (demographers, sociologists, economists) could exploit MIMI to investigate the trends in the various indicators, and the relationship among them. Moreover, it could be possible to develop complex models based on these data, able to assess human migration by evaluating related interdisciplinary drivers, as well as models able to nowcast and predict traditional migration indicators in accordance with original variables, such as the strength of social connectivity. Here, the SCI could have an important role. It measures the relative probability that two individuals across two countries are friends with each other on Facebook, therefore it could be employed as a proxy of social connections across borders, to be studied as a possible driver of migration. </p> <p>All in all, the motivations for building and releasing the MIMI dataset lie in the need of new perspectives, methods and analyses that can no longer prescind from taking into account a variety of new factors. The heterogeneous and multidimensional sets of data present in MIMI offer an all-encompassing overview of the characteristics of human migration, enabling a better understanding and an original potential exploration of the relationship between migration and non-traditional sources of data.</p> <p> </p> <p>The MIMI dataset is made up of one single CSV file that includes 28,821 rows (records/entries) and 876 columns (variables/features/indicators). Each row is identified uniquely by a pairs of countries, built from the joining of the two ISO-3166 alpha-2 codes for the origin and destination country, respectively. The dataset contains as main features the country-to-country bilateral migration flows and stocks, together with multidisciplinary variables measuring cultural, demographic, geographic and economic variables for the two countries, together with the Facebook strength of connectedness of each pair. </p> <p> </p> <p><strong>Related paper: </strong>Goglia, D., Pollacci, L., Sirbu, A. (2022). Dataset of Multi-aspect Integrated Migration Indicators. <a href="https://doi.org/10.5281/zenodo.6500885">https://doi.org/10.5281/zenodo.6500885</a></p>
Multi-stakeholder research data management training as a tool to improve the quality, integrity, reliability and reproducibility of research: Quantitative data of the post-course surveys
<p>Data contains doctoral students' and postdoc researchers' (n=168) self-ratings of their RDM competencies before and after the 3 ECTS credits "Basics of Research Data Management" (BRDM) trainings held 2019-2021 in the University of Turku and Åbo Akademi University, Finland. Moreover, data contains respondents' self-reported further learning needs.</p>
QRNG module integrated on a polymer photonic-platform (polyboard)
<p>This dataset includes measured random number distribution and generated randomness evaluation results on the Polyboard QRNG module with 4 output paths (1x4) as well as characterization measurements of the Polyboard QRNG module with 16 output paths and integrated SPADs including dark-count rates and detector efficiency evaluation.</p>
Coefficients for Tight Logarithmic Approximations and Bounds for Generic Capacity Integrals
<p>This is a supplementary dataset for the publication:</p> <p>I. M. Tanash and T. Riihonen, "Tight Logarithmic Approximations and Bounds for Generic Capacity Integrals and Their Applications to Statistical Analysis of Wireless Systems," in <em>IEEE Transactions on Communications</em>, 2022, doi: 10.1109/TCOMM.2022.3198435.</p> <p>The dataset contains the sets of optimized coefficients for the novel minimax approximations of the Nakagami and lognormal capacity integrals in terms of absolute error. The proposed approximations have the form of a weighted sum of logarithmic functions. The optimized coefficients are found for a wide range of the corresponding fading parameters, namely m for the Nakagami capacity integral and σ (standard deviation) for the lognormal capacity integral. Please note that the optimized coefficients in the provided dataset for the lognormal capacity integral are calculated for σdB (standard deviation in decibels) so σ=0.1 log_e(10) σdB in Eq. 5.</p> <p>The Matlab function (func_extract_coef.m) extracts the required set of optimal coefficients from the provided dataset according to the selected capacity integral, the parameter's value, and the number of terms. See help func_extract_coef for more information.</p> <p>The Matlab script (general_any_func) implements the theory presented in the corresponding journal paper: More specifically, it implements solving Eq. 22 to calculate the optimized coefficients of Eq. 7 for the Nakagami capacity integral. The code also provides general comments on how to generalize it to obtain the optimized coefficients of any communication system in terms of absolute error. Number of supplementary Matlab functions (general_any_func, func_abs_gen_any_func, calc_d_gen, calc_Cappr_gen, calc_d_gen_derivative, calc_Cappr_gen_derivative, Gauss_Laguerre, and peakseek) are provided herein and are used in the main Matlab script.</p> <p>A Matlab script (Example.m) is also provided as an example to illustrate the use of the provided Matlab function (func_extract_coef.m) in extracting the required coefficients from the dataset, to calculate and plot the corresponding absolute error which is shown by figure Example.jpg.</p>
Data, scripts, and R Notebook for Carneiro et al 2023. Flight performance and wing morphology in the bat Carollia perspicillata: biophysical models and energetics. Integrative Zoology DOI:10.1111/1749-4877.12707
<p>Files provided as supporting information for the paper by Carneiro et al. 2023. Flight performance and wing morphology in the bat <em>Carollia perspicillata</em>: biophysical models and energetics. Integrative Zoology. DOI:10.1111/1749-4877.12707</p> <p>File descriptions</p> <p>ArmTA.txt - Temperature and surface areas for arms of <em>C. perspicillata</em> after flight experiment<br> BodyTA.txt - Temperature and surface areas for body of <em>C. perspicillata</em> after flight experiment<br> HeadTA.txt - Temperature and surface areas for head of <em>C. perspicillata</em> after flight experiment<br> WingTA.txt - Temperature and surface areas for wings (patagium) of <em>C. perspicillata</em> after flight experiment<br> WingMorph.txt - Morphological variables measured in the body and wings of <em>C. perspicillata</em><br> HeatLoss.R - Function to estimate heat loss (Qt)<br> PowFlight.R - Function to estimate minimum power required to fly<br> Script-HeatLoss-FlightPerformance.R - R script with set of analyses performed<br> SupportingInformationFile.docx - R notebook with set of analyses performed, word format<br> SupportingInformationFile.nb.html - R notebook with set of analyses performed, html format<br> SupportingInformationFile.Rmd - R notebook with set of analyses performed (R markdown)</p> <p>For the R scripts (Script-HeatLoss-FlightPerformance.R) and notebook (<br> SupportingInformationFile.Rmd) to work and be compiled, all files need to be copied to the same folder.</p>
Life cycle inventories for the article: Circular Battery Production in the EU: Insights from integrating Life Cycle Assessment into System Dynamics Modeling on Recycled Content and Environmental Impacts
<p>This repository provides the unregionalized life cycle inventories to the paper "<span>Ginster, R.</span>, <span>Blömeke, S.</span>, <span>Popien, J. L.</span>, <span>Scheller, C.</span>, <span>Cerdas, F.</span>, <span>Herrmann, C.</span>, & <span>Spengler, T. S.</span> (<span>2024</span>). <span>Circular battery production in the EU: Insights from integrating life cycle assessment into system dynamics modeling on recycled content and environmental impacts</span>. <em>Journal of Industrial Ecology</em>, <span>1</span>–<span>18</span>. <a href="https://doi.org/10.1111/jiec.13527">https://doi.org/10.1111/jiec.13527</a>".</p> <h2>Contents</h2> <p>The repository is split into 2 parts and comprises the following files:</p> <p><strong>01_production: </strong>contains the necessary life cycle inventories for battery production.</p> <ul> <li><strong>01_primary</strong>: contains the life cycle inventories for battery production from primary materials.</li> <li><strong>02_secondary</strong>: contains the life cycle inventories for battery production from secondary materials.</li> <li><strong>03_active_material</strong>: contains the life cycle inventories for the active battery materials from primary materials.</li> <li><strong>04_active_material</strong>: contains the life cycle inventories for the active battery materials from secondary materials.</li> </ul> <p> </p> <p><strong>02_recycling: </strong>contains the necessary inventories for battery recycling.</p> <ul> <li><strong>01_process</strong>: contains the life cycle inventories for battery recycling.</li> <li><strong>02_intermediate</strong>: contains the life cycle inventories for the intermediate system for battery recycling.</li> <li><strong>03_output</strong>: contains the life cycle inventories for the resulting substances from battery recycling.</li> </ul> <h2>Summary</h2> <p>These files allow to reproduce the results of our study. Each file contains the life cycle inventory of one distinct battery capacity (20, 45, 68, 85, 95, 100 kWh) with a specific cell chemistry (LFP, NCA, NMC333, NMC532, NMC622, NMC811, NMC955) for battery production (based on Knehr et al. 2022) or for battery recycling (based on Blömeke et al. 2023).</p> <h2>Related publication</h2> <p>More details on the scientific context is provided in the publication itself:</p> <p><span>Ginster, R.</span>, <span>Blömeke, S.</span>, <span>Popien, J. L.</span>, <span>Scheller, C.</span>, <span>Cerdas, F.</span>, <span>Herrmann, C.</span>, & <span>Spengler, T. S.</span> (<span>2024</span>). <span>Circular battery production in the EU: Insights from integrating life cycle assessment into system dynamics modeling on recycled content and environmental impacts</span>. <em>Journal of Industrial Ecology</em>, <span>1</span>–<span>18</span>. <a href="https://doi.org/10.1111/jiec.13527">https://doi.org/10.1111/jiec.13527</a></p> <h2>Funding</h2> <p>This publication (Raphael Ginster and Steffen Blömeke) was created within the Research Training Group CircularLIB, supported by the Ministry of Science and Culture of Lower Saxony with funds from the program zukunft.niedersachsen of the Volkswagen Foundation (MWK | ZN3678).</p> <p>The publication on which this dataset is based were funded by the German Federal Ministry of Education and Research within the Competence Cluster Recycling & Green Battery (greenBatt) under the grant numbers 03XP0302A (Christian Scheller) and 03XP0331A (Jan-Linus Popien). The authors are responsible for the contents of this publication.</p>
The tectonic evolution of the Arctic since Pangea breakup: Integrating constraints from surface geology and geophysics with mantle structure
<div>Description of Resources - Shephard et al. (2013)</div> <div> </div> <div>This file provides a detailed description of all of the files that make up the data collection associated with the publication: Shephard, G. E., Müller, R. D., & Seton, M. (2013). The tectonic evolution of the Arctic since Pangea breakup: Integrating constraints from surface geology and geophysics with mantle structure. Earth-Science Reviews, 124(0), 148-183. doi: <a href="https://doi.org/10.1016/j.earscirev.2013.05.012" target="_blank" rel="noopener">10.1016/j.earscirev.2013.05.012</a></div> <div> </div> <div>Note: For information on file formats and what programs to use to interact with various file formats, see "File Formats and Recommended Programs”.</div> <div> </div> <div>Note: This paper is based on a global model (Seton et al., 2012), which should also be referenced if looking globally or regions other than the Arctic or northern Panthalassa.</div> <div> </div> <div>The files that make up the tectonic reconstruction model include:</div> <div>• <strong>Rotations </strong>- This is a global rotation model (based on Seton et al., 2012) that includes the new rotations for the Arctic.</div> <div>* Shephard_etal_ESR2013.rot (373 KB)</div> <div> </div> <div>• <strong>Coastlines </strong>- These are present day coastlines that have been assigned plate reconstruction ids to allow them to be reconstructed using the rotation file.</div> <div>* Shephard_etal_ESR2013_Coastlines.gpml (34.1 MB)</div> <div>* Shephard_etal_ESR2013_Coastlines.txt (3.2 MB)</div> <div>* Shephard_etal_ESR2013_Coastlinesc.kml (6.3 MB; datum - WGS 1984)</div> <div>* Shephard_etal_ESR2013_Coastlines.shp (3.2 MB inc auxiliary files; datum - WGS 1984)</div> <div> </div> <div>• <strong>Static polygons </strong>- These are closed polygons that split present day Earth's surface into regions that can be assigned to a given plate id, and therefore reconstructed back through time using the rotation file. These polygons can be used to cookie-cut and assign plate ids to geometry and raster data (for more information on this feature please visit http://gplates.org or http://earthbyte.org).</div> <div>* Shephard_etal_ESR2013_staticpolygons.gpml (19.4 MB)</div> <div>* Shephard_etal_ESR2013_staticpolygons.txt (2.7 MB)</div> <div>* Shephard_etal_ESR2013_staticpolygons.kml (4.4 MB; datum - WGS 1984)</div> <div>* Shephard_etal_ESR2013_staticpolygons.shp (2.3 MB inc auxiliary files; datum - WGS 1984)</div> <div> </div> <div>• <strong>Plate boundary geometries and resolved topologies</strong> – Resolved topologies comprise ridges, transforms, subduction zones and other plate boundary geometries. These boundaries intersect to form closed plate polygons ('resolved topologies') that are valid at 1 Myr intervals (0-200 Ma). The plate boundary geometries and plate polygons have been assigned plate reconstruction ids to allow them to be reconstructed using the rotation file.</div> <div>* Shephard_etal_ESR2013_platebounds.gpml (27.7 MB) - contains both plate boundaries and resolved topological plate polygons</div> <div>* Resolved topologies:</div> <div>- topology_*.00Ma.txt (20.6 MB)</div> <div>- topology_*.00Ma.shp (12.5 MB inc auxiliary files; datum - WGS 1984)</div> <div> </div> <div> </div> <div>References</div> <div> </div> <div>M. Seton, R.D. Müller, S. Zahirovic, C. Gaina, T.H. Torsvik, G. Shephard, A. Talsma, M. Gurnis, M. Turner, S. Maus, M. Chandler, (2012). Global continental and ocean basin reconstructions since 200 Ma. Earth-Science Reviews, 113(3–4), 212-270. doi:<a href="https://doi.org/10.1016/j.earscirev.2012.03.002" target="_blank" rel="noopener">10.1016/j.earscirev.2012.03.002</a></div>
Integrated Agent-based Modelling and Simulation of Transportation Demand and Mobility Patterns in Sweden
<h2>About</h2> <p><span>The Synthetic Sweden Mobility (SySMo) model provides a simplified yet statistically realistic microscopic representation of the real population of Sweden. The agents in this synthetic population contain socioeconomic attributes, household characteristics, and corresponding activity plans for an average weekday. This agent-based modelling approach derives the transportation demand from the agents’ planned activities using various transport modes (e.g., car, public transport, bike, and walking).</span></p> <div> <p>This open data repository contains four datasets: </p> <p>(1) Synthetic Agents, </p> </div> <div> <p>(2) Activity Plans of the Agents, </p> </div> <div> <p>(3) Travel Trajectories of the Agents, and </p> </div> <div> <p>(4) Road Network (EPSG: 3006)</p> <p><span>(OpenStreetMap data were retrieved on August 28, 2023, from https://download.geofabrik.de/europe.html, and GTFS data were retrieved on September 6, 2023 from https://samtrafiken.se/)</span></p> <p><span>The database can serve as input to assess the potential impacts of new transportation technologies, infrastructure changes, and policy interventions on the mobility patterns of the Swedish population.</span></p> </div> <h2>Methodology</h2> <p>This dataset contains statistically simulated 10.2 million agents representing the population of Sweden, their socio-economic characteristics and the activity plan for an average weekday. For preparing data for the MATSim simulation, we randomly divided all the agents into 10 batches. Each batch's agents are then simulated in MATSim using the multi-modal network combining road networks and public transit data in Sweden using the package pt2matsim (https://github.com/matsim-org/pt2matsim). </p> <p>The agents' daily activity plans along with the road network serve as the primary inputs in the MATSim environment which ensures iterative replanning while aiming for a convergence on optimal activity plans for all the agents. Subsequently, the individual mobility trajectories of the agents from the MATSim simulation are retrieved.</p> <p>The activity plans of the individual agents extracted from the MATSim simulation output data are then further processed. All agents with negative utility score and negative activity time corresponding to at least one activity are filtered out as the ‘infeasible’ agents. The dataset ‘<strong>Synthetic Agents</strong>’ contains all synthetic agents regardless of their <span>‘<em>feasibility</em>’ (0=excluded & 1=included in plans and trajectories). In the other datasets, only agents with feasible activity plans are included. </span></p> <p>The simulation setup adheres to the MATSim 13.0 benchmark scenario, with slight adjustments. The strategy for replanning integrates BestScore (60%), TimeAllocationMutator (30%), and ReRoute (10%)— the percentages denote the proportion of agents utilizing these strategies. In each iteration of the simulation, the agents adopt these strategies to adjust their activity plans. The "BestScore" strategy retains the plan with the highest score from the previous iteration, selecting the most successful strategy an agent has employed up until that point. The "TimeAllocationMutator" modifies the end times of activities by introducing random shifts within a specified range, allowing for the exploration of different schedules. The "ReRoute" strategy enables agents to alter their current routes, potentially optimizing travel based on updated information or preferences. These strategies are detailed further in W. Axhausen et al. (2016) work, which provides comprehensive insights into their implementation and impact within the context of transport simulation modeling. </p> <h2>Data Description</h2> <h3>(1) Synthetic Agents</h3> <p>This dataset contains all agents in Sweden and their socioeconomic characteristics. </p> <p>The attribute ‘<span><em>feasibility</em></span>’ has two categories: <em>feasible</em><em> agents </em>(73%), and <em>infeasible agents</em> (27%). <span>Infeasible agents are agents with negative utility score and negative activity time corresponding to at least one activity.</span> </p> <p>File name: 1_syn_pop_all.parquet</p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>PId</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td>Deso</td> <td>Zone code of Demographic statistical areas (DeSO)<sup>1</sup></td> <td>String</td> <td>-</td> </tr> <tr> <td> <pre>kommun</pre> </td> <td>Municipality code</td> <td>Integer</td> <td>-</td> </tr> <tr> <td> <pre>marital </pre> </td> <td>Marital Status (single/ couple/ child)</td> <td>String</td> <td>-</td> </tr> <tr> <td> <pre>sex </pre> </td> <td>Gender (0 = Male, 1 = Female)</td> <td>Integer</td> <td>-</td> </tr> <tr> <td> <pre>age</pre> </td> <td>Age</td> <td>Integer</td> <td>-</td> </tr> <tr> <td> <pre>HId</pre> </td> <td>A unique identifier for households</td> <td>Integer</td> <td>-</td> </tr> <tr> <td> <pre>HHtype </pre> </td> <td>Type of households (single/ couple/ other)</td> <td>String</td> <td>-</td> </tr> <tr> <td> <pre>HHsize </pre> </td> <td>Number of people living in the households</td> <td>Integer</td> <td>-</td> </tr> <tr> <td> <pre>num_babies</pre> </td> <td>Number of children less than six years old in the household</td> <td>Integer</td> <td>-</td> </tr> <tr> <td>employment</td> <td>Employment Status (0 = Not Employed, 1 = Employed)</td> <td>Integer</td> <td>-</td> </tr> <tr> <td>studenthood</td> <td>Studenthood Status (0 = Not Student, 1 = Student)</td> <td>Integer</td> <td>-</td> </tr> <tr> <td>income_class</td> <td>Income Class (0 = No Income, 1 = Low Income, 2 = Lower-middle Income, 3 = Upper-middle Income, 4 = High Income)</td> <td>Integer</td> <td>-</td> </tr> <tr> <td>num_cars</td> <td>Number of cars owned by an individual </td> <td>Integer</td> <td>-</td> </tr> <tr> <td>HHcars</td> <td>Number of cars in the household</td> <td>Integer</td> <td>-</td> </tr> <tr> <td> <pre>feasibility</pre> </td> <td>Status of the individual (1=feasible, 0=infeasible)</td> <td>Integer</td> <td>-</td> </tr> </tbody> </table> <p>1 <a href="https://www.scb.se/vara-tjanster/oppna-data/oppna-geodata/deso--demografiska-statistikomraden/">https://www.scb.se/vara-tjanster/oppna-data/oppna-geodata/deso--demografiska-statistikomraden/</a></p> <h3>(2) Activity Plans of the Agents</h3> <p>The dataset contains the car agents’ (agents that use cars on the simulated day) activity plans for a simulated average weekday. </p> <p>File name: <span>2_plans_i.parquet, i = 0, 1, 2, ..., 8, 9. (10 files in total)</span></p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>act_purpose</p> </td> <td> <p>Activity purpose (work/ home/ school/ other)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>PId</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>act_end </p> </td> <td> <p>End time of activity (0:00:00 – 23:59:59)</p> </td> <td> <p>String</p> </td> <td> <p>hour:minute:seco</p> <p>nd</p> </td> </tr> <tr> <td> <p>act_id</p> </td> <td> <p>Activity index of each agent</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>mode</p> </td> <td> <p>Transport mode to reach the activity location</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>POINT_X </p> </td> <td> <p>Coordinate X of activity location (SWEREF99TM)</p> </td> <td> <p>Float</p> </td> <td> <p>metre</p> </td> </tr> <tr> <td> <p>POINT_Y</p> </td> <td> <p>Coordinate Y of activity location (SWEREF99TM)</p> </td> <td> <p>Float</p> </td> <td> <p>metre</p> </td> </tr> <tr> <td> <p>dep_time </p> </td> <td> <p>Departure time (0:00:00 – 23:59:59)</p> </td> <td> <p>String</p> </td> <td> <p>hour:minute:seco</p> <p>nd</p> </td> </tr> <tr> <td> <p>score</p> </td> <td> <p>Utility score of the simulation day as obtained from MATSim</p> </td> <td> <p>Float</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>trav_time </p> </td> <td> <p>Travel time to reach the activity location</p> </td> <td> <p>String</p> </td> <td> <p>hour:minute:seco</p> <p>nd</p> </td> </tr> <tr> <td> <p>trav_time_min </p> </td> <td> <p>Travel time in decimal minute</p> </td> <td> <p>Float</p> </td> <td> <p>minute</p> </td> </tr> <tr> <td> <p>act_time </p> </td> <td> <p>Activity duration in decimal minute</p> </td> <td> <p>Float</p> </td> <td> <p>minute</p> </td> </tr> <tr> <td> <p>distance</p> </td> <td> <p>Travel distance between the origin and the destination</p> </td> <td> <p>Float</p> </td> <td> <p>km</p> </td> </tr> <tr> <td> <p>speed</p> </td> <td> <p>Travel speed to reach the activity location</p> </td> <td> <p>Float</p> </td> <td> <p>km/h</p> </td> </tr> </tbody> </table> <h3>(3) Travel Trajectories of the Agents</h3> <p>This dataset contains the driving trajectories of all the agents on the road network, <span>and the public transit vehicles used by these agents, including buses, ferries, trams etc. The files are produced by MATSim simulations and organised into 10 *.parquet’ files (representing different batches of simulation) corresponding to each plan file.</span></p> <p>File name: <span>3_events_i.parquet, i = 0, 1, 2, ..., 8, 9. (10 files in total)</span></p> <p> </p> <table> <tbody> <tr> <td> <div> <div> <p><strong>Column </strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Description </strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Data type </strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Unit </strong></p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>time </p> </div> </div> </td> <td> <div> <div> <p>Time in second in a simulation day (0-86399) </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>second </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>type </p> </div> </div> </td> <td> <div> <div> <p>Event type defined by MATSim simulation* </p> </div> </div> </td> <td> <div> <div> <p>String </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>person </p> </div> </div> </td> <td> <div> <div> <p>Agent ID </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>link </p> </div> </div> </td> <td> <div> <div> <p>Nearest road link consistent with the road network </p> </div> </div> </td> <td> <div> <div> <p>String </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>vehicle </p> </div> </div> </td> <td> <div> <div> <p>Vehicle ID identical to person </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>from_node </p> </div> </div> </td> <td> <div> <div> <p>Start node of the link </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>to_node </p> </div> </div> </td> <td> <div> <div> <p>End node of the link </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> </tbody> </table> <p>* One typical episode of MATSim simulation events: Activity ends (actend) -> Agent’s vehicle enters traffic (vehicle enters traffic) -> Agent’s vehicle moves from previous road segment to its next connected one (left link) -> Agent’s vehicle leaves traffic for activity (vehicle leaves traffic) -> Activity starts (actstart) </p> <h3>(4) Road Network</h3> <p>This dataset contains the road network.</p> <p>File name: 4_network.shp</p> <table> <tbody> <tr> <td> <div> <div> <p><strong>Column </strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Description </strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Data type </strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Unit </strong></p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>length </p> </div> </div> </td> <td> <div> <div> <p>The length of road link </p> </div> </div> </td> <td> <div> <div> <p>Float </p> </div> </div> </td> <td> <div> <div> <p>metre </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>freespeed </p> </div> </div> </td> <td> <div> <div> <p>Free speed </p> </div> </div> </td> <td> <div> <div> <p>Float </p> </div> </div> </td> <td> <div> <div> <p>km/h </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>capacity </p> </div> </div> </td> <td> <div> <div> <p>Number of vehicles </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>permlanes </p> </div> </div> </td> <td> <div> <div> <p>Number of lanes </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>oneway </p> </div> </div> </td> <td> <div> <div> <p>Whether the segment is one-way (0=no, 1=yes) </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>modes </p> </div> </div> </td> <td> <div> <div> <p>Transport mode </p> </div> </div> </td> <td> <div> <div> <p>String </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>from_node </p> </div> </div> </td> <td> <div> <div> <p>Start node of the link </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>to_node </p> </div> </div> </td> <td> <div> <div> <p>End node of the link </p> </div> </div> </td> <td> <div> <div> <p>Integer </p> </div> </div> </td> <td> <div> <div> <p>- </p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>geometry </p> </div> </div> </td> <td> <div> <div> <p>LINESTRING (SWEREF99TM) </p> </div> </div> </td> <td> <div> <div> <p>geometry </p> </div> </div> </td> <td> <div> <div> <p>metre </p> </div> </div> </td> </tr> </tbody> </table> <p> </p> <p><strong><span>Additional Notes</span></strong></p> <p><span>This research is funded by the RISE Research Institutes of Sweden, the Swedish Research Council for Sustainable Development (Formas, project number 2018-01768), and Transport Area of Advance, Chalmers.</span></p> <p><strong><span>Contributions</span></strong></p> <p><span>YL designed the simulation, analyzed the simulation data, and, along with CT, executed the simulation. CT, SD, FS, and SY conceptualized the model (SySMo), with CT and SD further developing the model to produce agents and their activity plans. KG wrote the data document. All authors reviewed, edited, and approved the final document.</span></p>
Using machine learning to integrate genetic and environmental data to model genotype-by-environment interactions
<p>Files generated from the study described in <a href="https://doi.org/10.1101/2024.02.08.579534">Fernandes et. al (2024)</a> .</p> <p>The file "cvs_h2s.csv" comprises the coefficient of variation and the Cullis heritability for each environment.</p> <p>The file "all_predictions.csv" contains the predictions from all the models evaluated, in different cross-validation (CV) scenarios.</p> <p>The file "coincidence_index.csv" has the Coincidence Index (CI) for each CV and models evaluated in our study.</p> <p>Our study used the multi-environment maize yield trials data from the Genomes to Fields 2022 initiative (<a href="https://doi.org/10.1186/s13104-023-06421-z">Lima et. al 2024</a>).</p>
Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants
<p>This dataset consists of the reference data files, metadata and processed results files for the paper "Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants," which investigates clonality in normal human dermal fibroblast cell populations in 32 cell lines from distinct donors, using bulk whole-exome sequencing and single-cell RNA-sequencing data.</p> <p>This dataset contains everything required to reproduce the results presented in the paper from processed data and results of our data processing workflows. Our analyses can be reproduced using the <a href="https://github.com/davismcc/fibroblast-clonality">source code</a> and instructions available at our <a href="https://davismcc.github.io/fibroblast-clonality/">project website</a>.</p> <p>The <em>entire</em> analysis workflow from raw data to final results is also reproducible but is substantially more complicated and computationally intensive. It also requires large datasets to be obtained from other repositories. Specifically, single-cell RNA-seq data have been deposited in the ArrayExpress database at EMBL-EBI under accession number E-MTAB-7167. Whole-exome sequencing data is available through the HipSci portal (www.hipsci.org). Combined with the dataset in this repository and following the instructions on the project website, it is possible to run our entire analysis pipeline.</p> <p> </p>
ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics
<p>ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics</p> <p>Supplementary file (full dataset download)</p> <p>- v2 with SMILES errors fixed (19.01.2019)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.