Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,235
datasets available to search
ShareScore release 0.9.0
Dataset results
2,235 results for “engineering”
Research data supporting "Raman spectroscopic imaging for quantification of depth-dependent and local heterogeneities in native and engineered cartilage"
<p>Research data supporting the publication: Albro M. et al., 2018, npj Regenerative Medicine, DOI: https://doi.org/10.1038/s41536-018-0042-7.</p>
Word Embeddings for the Software Engineering Domain
<p>A .bin file for a word2vec model pre-trained on 15GB of Stack Overflow posts. </p> <p>For more details refer to the following paper:</p> <p>Efstathiou, V., Chatzilenas, C., Spinellis, D., 2018. "Word Embeddings for the Software Engineering Domain". In <em>Proceedings of the 15th International Conference on Mining Software Repositories.</em> ACM</p>
Phase transitions as intermediate steps in the formation of molecularly engineered protein fibers
<p>This upload contains raw and unprocessed data sets including: Tensile test, diffraction, simulation, surface tension measurement, viscosity measurement, amino acid sequence and videos.</p>
GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin
<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>
Frame Embeddings for Software and Requirements Engineering Domain
<p>This project is aimed to identify semantic relatedness of <a href="https://framenet2.icsi.berkeley.edu/">FrameNet </a>semantic frames in the domain of software and requirements engineering. The folder contains the frame embeddings that are obtained using the <strong>context-based method</strong> described in our ESEM paper*.</p> <p>Waad Alhoshan, Liping Zhao, and Riza Batista-Navarro. 2018. Using Semantic Frames to Identify Related Textual Requirements: An Initial Validation. In ACM / IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) (ESEM ’18), October 11–12, 2018, Oulu, Finland. ACM, New York, NY, USA, 2 pages. https://doi.org/10.1145/3239235.3267441 </p> <p> </p> <p> </p>
Research data supporting "Engineering anisotropic muscle tissue using acoustic cell patterning"
<p>Raw research data supporting the publication:</p> <p>Armstron, JPK et al., "Engineering anisotropic muscle tissue using acoustic cell paterning", Advanced Materials, DOI: 10.1002/adma.201802649 (2018)</p>
Observations of groundwater fluctuations and surface moisture content on a medium-grained, planar beach (Sand Engine, the Netherlands)
<p>These data are groundwater and beach surface moisture values collected during the MegaPex campaign between October 11 and 20, 2014 at the Sand Engine, The Hague, the Netherlands by MSc students and staff of the Coastal Research Group at Utrecht University, the Netherlands. The data were obtained at 8 locations in a cross-shore array on the intertidal and upper beach. During the measurements the beach was planar (1:30) and the median grain size was 0.365 mm. The data are supplemented with bed profiles along the instrument array. For further information and meta-data, please consult the readme.txt and the header of the individual text files in the zip-file.</p>
CS2_ITD_ENG_6_Engine torque
<p>Piston engines generate an output torque with very high instantaneous variations, leading to severe constraints on engine, propeller and mounting frame structure (on aeronautical applications). This information is mandatory to be able to design a mounting frame able to cope with such an engine.</p>
Metadata on Articles Published at the Requirements Engineering Conference, REFSQ conference, or Requirements Engineering Journal from 2009 until 2018
<p>In the 1990s, it was recognized that Requirements Engineering lays the foundation for high quality software. A substantial research community has formed that set out to enable practitioners of the 21st century to systematically adopt proven strategies to common development challenges and to enable the engineering of innovative solutions and product features. But is contemporary RE Research delivering what it set out to deliver? In the article at IEEE Software 36(4) with DOI 10.1109/MS.2019.2909127, we provide a brief overview over the accomplishments of the past 10 years and identify open opportunities. The work at hand is the raw dataset of metadata, specifically keywords and author names, of articles published at the Requirements Engineering Conference, REFSQ conference, or Requirements Engineering Journal from 2009 until 2018 and supplements our article.</p>
Survey Data Set Part 1 - Attitudes Towards Videos as a Documentation Option for Communication in Requirements Engineering
<p>In 2017, we conducted an online survey to explore software professionals' attitudes towards videos as a documentation option for communication in requirements engineering. The survey covered the following topics:</p> <ul> <li>Demographics</li> <li>Attitude towards videos as a medium in RE including its strengths, weaknesses, opportunities, and threats</li> <li>Current production and use of videos in RE, respectively the obstacles that prevent the production and use of videos</li> </ul> <p>64 out of 106 software professionals from industry and academia completed the survey. The survey was implemented in LimeSurvey and distributed across several communication channels such as LinkedIn, ResearchGate, and a mailing list of a German RE professionals group.</p> <p>This dataset includes the following files:</p> <ul> <li>"Raw and analyzed data.xlsx" contains the raw and analyzed survey responses which are anonymized <ul> <li>This data includes <em>demographics </em>and <em>attitude</em>.</li> <li>The data on <em>video production and use</em> are included in: <a href="https://zenodo.org/record/4064741">Survey Data Set Part 2 - Attitudes Towards Videos as a Documentation Option for Communication in Requirements Engineering</a>.</li> </ul> </li> <li>"Survey - Offline version.docx" contains the questions and possible answers of the survey</li> <li>"Survey - Offline version.pdf" contains the questions and possible answers of the survey</li> </ul> <p>This survey was designed, conducted, and analyzed by Oliver Karras (<a href="https://twitter.com/KarrasOliver">@KarrasOliver</a>).</p>
Replication package for "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Analysis of the DLR Knowledge Exchange Workshop Series on Software Engineering
<p>This repository is used to analyze the workshops of the DLR internal workshop series on software<br> engineering. These workshops are two-day events of the DLR software engineering community and<br> focus on different main topics every year.</p>
Injectable, Scalable 3D Tissue-Engineered Model of Marrow Hematopoiesis
<p>Raw data associated with the publication "<strong>Injectable, Scalable 3D Tissue-Engineered Model of Marrow Hematopoiesis"</strong></p> <p><a href="https://www.sciencedirect.com/science/article/pii/S0142961219307641"><strong>DOI: 10.1016/j.biomaterials.2019.119665</strong></a></p>
Evaluation data used in "An innovative STEM outreach model (OH-Kids) to foster the next generation of geoscientists, engineers, and technologists"
<p>This repository contains all data of the evaluation questionnaire used to assess modifications in pupils’ perceptions of same water resources concepts and science and scientist resulting from the application of OH-Kids outreach model in six Mexican primary schools (n=344 pupils).</p>
DATASET of Large-scale Neural Recordings for DENOISING Engine
<p><span>30 sec raw data (.brw) was recorded with BrainWave SW and detected LFP events and spikes were stored in (.bxr). These extracellular recordings were obtained from acute hippocampal-cortical slices and were collected at 14KHz/electrode sampling frequency.</span></p>
Levels of a Research Software Engineer
<p><strong>Levels of a Research Software Engineer: </strong>The diverse role of the RSE can be captured by the degree or level to which they work with researchers, and in what scope. Level 1 of RSE "domain" are closest to researchers, working directly on their behalf. Level 2 "generalist" RSE work on core technologies needed across the scientific community, and level 3 "researcher" take this a step further, researching the space or models underlying the software itself.</p>
San Diego Earthquake Dataset with Feature-Engineered Variables
<p>For the San Diego region, using data from the Southern California Earthquake Data Center (SCEDC), we filtered events by latitude 32.715, longitude -117.1611 within a 150 km radius, focusing on earthquake events from August 1, 2004, 00:00:00 to August 1, 2024, 00:00:00. All magnitude types and depths were included, and 21 variables were feature-engineered to enhance predictive modeling. This dataset provides a robust foundation for earthquake prediction in the San Diego area, incorporating both raw seismic data and advanced engineered features.</p>
Spatial Feature Engineering Dataset for Forest Aboveground Biomass Estimation Using Landsat Imagery
<p><strong>Study Area:</strong><br>The dataset covers forested regions in Oregon, Washington, Idaho, and eastern Montana, characterized by diverse climatic conditions due to orographic effects. The forests in the Coast Range and western slopes of the Cascades, with high precipitation (800-3000 mm annually), contrast with the drier forests in Idaho and Montana, which receive over 400 mm annually. The dataset includes highly productive Douglas-fir and western hemlock forests, with aboveground biomass (AGB) densities exceeding 1200 Mg ha⁻¹, as well as fire-adapted lodgepole and ponderosa pine forests in the rainshadow regions.</p> <p><strong>LiDAR AGB Estimates:</strong><br>The dataset includes 176 lidar-derived AGB maps from 2002 to 2016, covering various regions in Oregon, Washington, Idaho, and Montana. A Random Forest (RF) model was used to estimate AGB at a 30m² resolution, utilizing lidar height features, DEM features, and climate data. Non-forested areas and buildings were masked using binary forest cover maps from the LCMS dataset and the Microsoft Building Footprints dataset.</p> <p><strong>Reference Dataset:</strong><br>A composited AGB map, derived from the 176 lidar maps, was created to develop Landsat-based AGB models, covering 9,361,622 ha of forested land. The AGB layer was stratified into 30 bins, and training, development, and testing sets were constructed for model validation. The dataset includes 7500 test samples and 300,000 training and development samples, with a 500m buffer around test set locations to prevent spatial autocorrelation.</p> <p><strong>Landsat Satellite Imagery:</strong><br>Landsat imagery from 1990 to 2022 was utilized, with three time series derived: all scenes, scenes from May to November, and annual medoid composites. The imagery was processed using the Google Earth Engine (GEE) platform, focusing on periods of maximum phenological activity.</p> <p><strong>Feature Engineering:</strong><br>Extensive feature engineering was performed, generating spectral, spatial, temporal, and topographic features from Landsat imagery and DEM data. Features were extracted over the reference AGB map's domain, synchronized with the lidar acquisition dates.</p> <ul> <li><strong>LandTrendr Fitted Imagery:</strong> Spectral features were derived from LandTrendr-fitted imagery, smoothing variations in the time series.</li> <li><strong>LandTrendr Disturbance and Recovery Features:</strong> Temporal features were derived from LandTrendr models, characterizing disturbance and recovery events.</li> <li><strong>CCDC Disturbance and Recovery Features:</strong> CCDC algorithm-derived features characterized disturbances and recovery using harmonic models.</li> <li><strong>Buffer Features:</strong> Local variations were captured using buffer statistics around each pixel.</li> <li><strong>GLCM Features:</strong> GLCM texture features summarized the joint distribution of gray-tone values.</li> <li><strong>Edge Detectors:</strong> Various edge detection operators captured spatial derivatives and edges.</li> <li><strong>Morphological Operations:</strong> Morphological features were derived using multi-channel image processing techniques.</li> <li><strong>Neighborhood Vectorization:</strong> Direct vectorization of satellite measurements in pixel neighborhoods.</li> <li><strong>Neighborhood Similarity:</strong> Similarity features characterized the relationship between pixel neighborhoods and their centroids.</li> <li><strong>Topographic Features:</strong> Topography was characterized using elevation, slope, aspect embeddings, and hillshade layers from the NED DEM.</li> </ul> <p>This comprehensive dataset enables robust analysis of AGB models and their performance across diverse forested landscapes in the Pacific Northwest</p>
Eumelanin-Enhanced Photothermal Disinfection of Contact Lenses Using a Sustainable Marine Nanoplatform Engineered with Electrospun Nanofibers_(antibacterial study - S.aureus)
<p>Eumelanin-Enhanced Photothermal Disinfection of Contact Lenses Using a Sustainable Marine Nanoplatform (antibacterial study - S.aureus)</p>
Figure 6 in Les engins et techniques de pêche utilisés dans la baie de Loango, République du Congo, et leurs incidences sur les prises accessoires
Figure 6. - Courbe réponse de la proportion de tortues retrouvées mortes par événement de pêche en fonction du temps de calée (en heures). L'intervalle de confiance à 95% apparaît en pointillés. +: valeurs observées. [Response curve of the proportion of sea turtles found dead per fishing event according to soaking time (in hours). The 95% confidence interval appears in dotted lines. +: observed values.]
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.