Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,235
datasets available to search
ShareScore release 0.9.0
Dataset results
2,235 results for “engineering”
Research Beyond the Lab, Spring Term 2022, Global Health Engineering, ETH Zurich. Raw data and analysis-ready derived data on waste management in public spaces in Zurich, Switzerland.
<p>This repository contains all raw and derived data produced as part of the <a href="https://rbtl-fs22.github.io/website/">ETH Zurich course "Research Beyond the Lab: Open Science and Research Methods for a Global Engineer" (151-8102-00L)</a> offered in spring term 2022.</p> <p>Students were assigned teams of four to conduct a collaborative research project broadly addressing the theme of “Trash in the Public Spaces of Zurich” in collaboration with <a href="https://www.stadt-zuerich.ch/ted/de/index/entsorgung_recycling.html">Entsorgung & Recycling Zürich (ERZ)</a>, the waste management department at Stadt Zürich.</p> <p>Research methods and design are taught in the first half of the course. Surveys and a waste characterisation study are then designed based on the research questions students have developed in their respective teams. The collected raw data is used in the course to teach principles of research data management, tidy data structures, reproducible research with R & RStudio, and collaboration and version control with Git & GitHub.</p>
Sentinel-2: Cloud Probability in Earth Engine
<p>Links:</p> <ul> <li><a href="https://developers.google.com/earth-engine/datasets/catalog/COPERNICUS_S2_CLOUD_PROBABILITY">Sentinel-2: Cloud Probability</a> in Earth Engine's Public Data Catalog</li> <li><a href="https://radiantearth.github.io/stac-browser/#/external/storage.googleapis.com/earthengine-stac/catalog/COPERNICUS/COPERNICUS_S2_CLOUD_PROBABILITY.json">Sentinel-2: Cloud Probability</a> in Earth Engine STAC viewed with STAC Browser</li> </ul> <p>The S2 cloud probability is created with the <a href="https://github.com/sentinel-hub/sentinel2-cloud-detector">sentinel2-cloud-detector</a> library (using <a href="https://github.com/microsoft/LightGBM">LightGBM</a>). All bands are upsampled using bilinear interpolation to 10m resolution before the gradient boost base algorithm is applied. The resulting <code>0..1</code> floating point probability is scaled to <code>0..100</code> and stored as a UINT8. Areas missing any or all of the bands are masked out. Higher values are more likely to be clouds or highly reflective surfaces (e.g. roof tops or snow).</p> <p>Sentinel-2 is a wide-swath, high-resolution, multi-spectral imaging mission supporting Copernicus Land Monitoring studies, including the monitoring of vegetation, soil and water cover, as well as observation of inland waterways and coastal areas.</p> <p>The Level-2 data can be found in the collection <a href="https://radiantearth.github.io/stac-browser/COPERNICUS_S2_SR">COPERNICUS/S2_SR</a>. The Level-1B data can be found in the collection <a href="https://radiantearth.github.io/stac-browser/COPERNICUS_S2">COPERNICUS/S2</a>. Additional metadata is available on assets in those collections.</p> <p>See <a href="https://developers.google.com/earth-engine/tutorials/community/sentinel-2-s2cloudless">this tutorial</a> explaining how to apply the cloud mask.</p>
Design Methodologies and Engineering Applications for Ecosystem Biomimicry: An Interdisciplinary Review Spanning Cyber, Physical, and Cyber-Physical Systems
<p>This is the data for the interdisciplinary review on ecosystem biomimicry for engineering applications. </p>
Impact of Passive Voice in Requirements Engineering
<p>This repository contains instrumentation material and results for the experiment described in the paper "On The Impact of Passive Voice Requirements on Domain Modelling" by Henning Femmer, Jan Kucera, and Antonio Vetrò from Technische Universität München.</p>
Diversity Awareness in Software Engineering Participant Research
<p>This dataset contains the result of a classification of three ICSE venues namely, ICSE 2019, 2020, and 2021 technical tracks, as stated in the methodology of the paper “Diversity awareness in software engineering participant studies” by Dutta et al. (2023).</p>
Testing Database Engines via Query Plan Guidance
<p>This artifact is the supplementary material for the paper "Testing Database Engines via Query Plan Guidance" that is published in ICSE'23.</p> <p>In the paper, we propose a new concept called Query Plan Guidance (QPG) for testing Database Management Systems (DBMSs). QPG tackles the test case generation problem by guiding test case generation towards exploring a variety of unique query plans. We found 50+ unique, previously-unknown bugs in SQLite, TiDB, and CockroachDB.</p> <p>The prototype of QPG is a Java project, and we provide Dockerfiles to reproduce our three key empirical results.</p>
Mapping ecosystem types and land cover types in the Seychelles granitic islands, using Earth Engine and Sentinel-2
<p>We share here maps produced using Earth Engine: https://code.earthengine.google.com/?accept_repo=users/bsenterre/gis</p> <p>The maps include a land cover classification based on Sentinel-2, at 10m resolution, using an Object-Based Image Analysis approach, for the Seychelles granitic islands. Based on the land cover, landform (modeled using TauDEM), altitude and expert knowledge, we then derived a model of ecosystem types, with 3 maps: current distribution, potential distribution and prehuman distribution.</p> <p>A report exists (18th May 2022) that describes in detail the methodology, and it is being used for the preparation of a publication. The maps uploaded here are in raster format (geotif), crs=4326, and are accompanied by QGIS legend files (.qml), so they should load in QGIS with their legend automatically.</p>
Chinese Engineers Relational Database (CERD) Bi-monthly Export
<p><strong>This is a bi-monthly export.</strong></p> <p>CERD is a database of engineers from the Chinese Republican period (1912–1949). Based on various digitised historical sources, it is a prosopographic catalogue of individuals, their education and their employment, and the institutions connected with it. Most biographical events have geographical information attached to them. The data can be used freely by researchers to answer individual research questions.</p> <p><strong>Citation recommendation:</strong></p> <p>Pelzer, Thorben, et al., eds. (2021–2023). Chinese Engineers Relational Database (CERD) (Version 1.7.0). Zenodo. http://doi.org/10.5281/zenodo.4075601.</p> <p><strong>Changelog:</strong></p> <p>1.7.0 (February 2023): Approx. 17,600 individuals, two additional sources<br> 1.6.0 (August 2022): Approx. 17,400 entries, completed sources, minor corrections, mergers, translations<br> 1.5.0 (June 2022): Approx. 17,300 entries, added additional memberships<br> 1.4.0 (April 2022): Approx. 16,800 entries, added additional sources, schooling<br> 1.3.0 (February 2022): Approx. 16,500 entries, added documentation, additional memberships<br> 1.2.0 (December 2021): Approx. 16,300 entries, added frequent CSV exports, source annotations, 5 missing <em>minglu</em> pages, additional association memberships<br> 1.1.0 (October 2021): Approx. 15,700 entries, added selected association memberships<br> 1.0.0 (August 2021): Approx. 15,400 entries [complete <em>gongchengren minglu</em> dataset milestone]<br> 0.5.0 (June 2021): Approx. 13,000 entries<br> 0.4.0 (April 2021): Approx. 10,500 entries<br> 0.3.0 (February 2021): Approx. 7,500 entries, as well as various corrections, mergers, translations<br> 0.2.0 (December 2020): Approx. 5,000 entries, as well as various corrections, mergers, translations<br> 0.1.0 (October 2020): Early version with approx. 3,000 entries</p> <p><strong>Online Access:</strong></p> <p>Via Heurist: <a href="https://home.uni-leipzig.de/cerd/">https://home.uni-leipzig.de/cerd/</a></p>
Water displacement device, engines characterization
<p>This dataset contains 44 formulations (rows), 23 features and 1 supervisor factors (24 columns).</p> <p>This data have been extracted from a water displacement device that measures the volume of submerged samples. The displaced water passes through a flow sensor that produces pulses. The features in the dataset are the times between the first 20 pulses (t0 - t19), the total pulse counting (pc), the total operation time (tt) and the DC (dc) component on each experiment (formulation). The supervising value is the volume of marbles with different diameters.</p> <p>The work consist on tuning the water displacement device so it yields accurate volumes from pulse patterns.</p>
Heterogeneous environmental seascape across a biogeographic break influences the thermal physiology and tolerances to ocean acidification in an ecosystem engineer
<p>Dataset for the metabolic rates of limpets under two different pCO2/pH conditions</p> <p>MR are in O2 mg h−1g−1</p>
replicAnt - Plum2023 - 3D Models - Unreal Engine 5
<p>This dataset contains the 3D models used to generate all synthetic data presented in the <em>replicAnt - generating annotated images of animals in complex environments using Unreal Engine </em>manuscript. The models have been generated with the open-source photogrammetry platform <em>scAnt</em> <a href="https://peerj.com/articles/11155/">peerj.com/articles/11155</a>/ and have been pre-processed and converted into Unreal Engine 5 compatible .uasset files, to be used with the associated <em>replicAnt</em> project available from <a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created <em>replicAnt</em>, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. <em>replicAnt</em> places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with <em>replicAnt</em> can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, <em>replicAnt</em> may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College’s President’s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>
Rapid Review Dataset for Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team
<p>A collection of evidence briefings produced through a rapid literature review protocol the Department of Software Engineering and Research at Sandia National Laboratories. These briefings are described in our research paper, "Seeking Enlightenment: Incorporating Evidence-Based Practice Techniques in a Research Software Engineering Team", which was accepted for publication at the 1st Annual Conference of the United States Research Software Engineer Association (US-RSE'23).</p> <p>Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525. SAND2023-06549O.</p>
CRISPR-based engineering of RNA viruses
<p>CRISPR RNA-guided endonucleases have enabled precise editing of DNA. However, options for editing RNA remain limited. Here, we combine sequence-specific RNA cleavage by CRISPR ribonucleases with programmable RNA repair to make precise deletions and insertions in RNA. This work establishes a new recombinant RNA technology with immediate applications for the facile engineering of RNA viruses.</p> <p> </p> <p>This dataset contains code for analyzing sequencing data and generating figures in the manuscript.</p>
A Thematic Synthesis on Empathy in Software Engineering based on the Practitioners' Perspective - Supplementary Material
<p>This repository contains the supplementary material of the paper "A Thematic Synthesis on Empathy in Software Engineering based on the Practitioners' Perspective" <br> (DOI https://doi.org/10.1145/3613372.3613407) accepted at the Research Track of the <br> XXXVII Brazilian Symposium on Software Engineering (SBES 2023).</p> <p>The artifacts are a result of a thematic synthesis of grey literature <br> performed to investigate the meaning, importance, practices, and effects of empathy <br> from the perspective of software practitioners. <br> The analysis was based on web articles from DEV, an online community used by software developers. <br> The data were collected and stored in the repository to preserve the evidence and ensure the study’s replicability.<br> <br> The repository contains the following material:</p> <p>1- <all codes.ods> and <all codes.xlsx><br> All codes generated in the data extraction process, considering research questions RQ1-RQ5:<br> The two files have the same content in different formats - ODS and XLSX.</p> <p>2 - <dataset.csv> <br> The list of web articles collected from the DEV in CSV format with all inclusion and <br> exclusion information, plus demographic data.</p> <p>3 - <empathy-framework.jpg><br> Figure 3 of the paper: A conceptual map of the meaning (boxes in orange) and <br> the value (boxes in blue) of empathy according to the software practitioners</p> <p>4 - <empathy-model.jpg> <br> Figure 4 of the paper: A conceptual framework for communication and collaboration (A), <br> management and leadership (B), coding (C), and code review (D).</p> <p>5 - <extraction.ods> and <extraction.xlsx>. The data extracted from the web articles, <br> including codes and quotes for each research question. <br> The two files have the same content in different formats - ODS and XLSX.</p>
Dataset - Survey results - Applying Model-based Requirements Engineering in AIDOaRt Collaborative Project
<p>This dataset and its associated report contain the results of an online survey on using a model-based requirements engineering approach in AIDOaRT project in 2022.</p>
Buried Interface Engineering Enables Efficient and 1,960-hour Isos-L-2i Stable Inverted Perovskite Solar Cells
<p>High-performance perovskite solar cells (PSCs) typically require interfacial passivation, yet this is challenging for the buried interface, owing to the dissolution of passivation agents during the deposition of perovskites. Here, we overcome this limitation with in-situ buried interface passivation – achieved via directly adding a cyanoacrylic acid-based molecular additive, namely BT-T, into the perovskite precursor solution. Classical and ab-initio molecular dynamics simulations reveal that BT-T spontaneously may self-assemble at the buried interface during the formation of the perovskite layer on a nickel oxide hole transporting layer. The preferential buried interface passivation results in facilitated hole transfer and suppressed charge recombination. In addition, residual BT-T molecules in the perovskite layer enhance its stability and homogeneity. We report a power-conversion efficiency (PCE) of 23.48% for 1.0 cm2 inverted-structure PSCs. The encapsulated PSC retains 95.4% of its initial PCE following 1,960-hour maximum power point tracking under continuous light illumination at 65°C (i.e., ISOS-L-2I protocol). Our demonstration of operating-stable PSCs under accelerated ageing conditions represents a step closer to the commercialization of this emerging technology.</p>
Cognition in Social Engineering Empirical Research: a Systematic Literature Review
<p># Description of contents</p> <p>This repository contains the codebook, dataset and analysis scripts in R used for the following publication: Pavlo Burda, Luca Allodi, Nicola Zannone. "Cognition in Social Engineering Empirical Research: a Systematic Literature Review". ACM Transactions on Computer-Human Interaction (TOCHI).<br> The repository consists of the following files:</p> <p>## dataset_and_codebook.xlsx contains the dataset, the codebook and a detailed description of contents.</p> <p>## scripts/ contains the R scripts used for the analysis</p> <p>## readme.txt contains this readme</p> <p><br> # Dataset and codebook</p> <p>The dataset_and_codebook.xlsx contains the following sheets:</p> <p>## Codebook<br> Contains a detailed description of the dataset (tables, columns, fields, etc.).<br> The codebook describes the concepts and variables that are present in the dataset. This includes explanations on meaning, numerical values, classification schemes and labels.</p> <p>## Hypotheses table <br> Contains the hypotheses for all analyzed papers and is used in the results of the paper. <br> Each row is a hypothesis of an included paper, cells contain one or more values (e.g., value1,value2,...) with or without sub values (e.g., value1,value2(sub-value1, ...)). Any content that is in square brackets [] is ignored in the analysis. Empty cells mean that there is no applicable value for that column.</p> <p>## Papers table<br> Contains the analyzed papers and is used in the results of the paper. It is also used in the overview table (Table 6) in Appendix C.<br> Each row is an included paper, cells contain up to two values (e.g., value1, value2) or the word 'multiple' in case of more than two values. Empty cells mean that there is no applicable value for that column.</p> <p>## Values table <br> Contains the description of cell values of 'Hypotheses' and 'Papers' and sub-values (specific variables in a study) that belong to a value. <br> It has a hierarchical structure from left to right where each field on the right column falls under the first non-empty field on the immediate left-top.</p> <p><br> # Reproducing results with scripts/ (generate figures)<br> To run the R scripts included in the 'scripts' directory it is sufficient to follow the instructions in and run the 'RUN_ALL_SCRIPTS.R' in 'scripts' directory.<br> The scripts use the TSV (Tab Separated Values) format of the dataset, which is the exact copy of the 'Papers' and 'Hypotheses' tables in the 'dataset_and_codebook.xlsx' file.<br> The resulting figures are stored in the 'scripts/results' directory.</p>
Systematic Multi-Trait AAV Capsid Engineering for Efficient Gene Delivery
<p>Datasets for "Systematic Multi-Trait AAV Capsid Engineering for Efficient Gene Delivery", Eid et al., <em>Nature Communications. </em></p>
Dataset for study "Band gap engineering by cationic substitution in Sn(Zr1-xTix)Se3 alloy for bottom sub-cell application in solar cells"
<p>This dataset contains raw and processed data that were used to for the study entilted "Band gap engineering by cationic substitution in Sn(Zr1-xTix)Se3 alloy for bottom sub-cell application in solar cells". </p>
Experimental Results of "Bioprinting Cell- and Spheroid-Laden Protein-Engineered Hydrogels as Tissue-on-Chip Platforms"
<p>This repository contains the experimental results of the article "Bioprinting Cell- and Spheroid-Laden Protein-Engineered Hydrogels as Tissue-on-Chip Platforms" by Duarte Campos, D., Lindsay, C., Roth, J., LeSavage, B., Seymour, A., Krajina, B., Ribeiro, R., Costa, P., Heilshorn, S., published in <em>Front. bioeng. biotechnol. </em><strong>8, 374</strong> (2020). https://doi.org/10.3389/fbioe.2020.00374</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.