Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
16
datasets available to search
ShareScore release 0.7.1
Dataset results
16 results for “Feature Engineering”
A Hybrid Feature Location Technique for Re-engineering Single Systems into Software Product Lines
<p>The dataset used for evaluating the hybrid feature location technique presented in the paper: "A Hybrid Feature Location Technique for Re-engineering Single Systems into Software Product Lines". This enables reproducibility, evaluation, and comparison of our study.</p> <p>_________________________________________________________________________________________________________</p> <p>Folder "Dataset" contains for each subject system used:</p> <p>(i) the artificial variants and their configurations;</p> <p>(ii) the ECCO repository containing the traces;</p> <p>(iii) the ground truth and composed variants;</p> <p>(iv) the metrics results.</p> <p>_________________________________________________________________________________________________________</p> <p>Folder "Scenarios" contains for each subject system used:</p> <p>(i) the videos recorded from exercising features on GUI.</p>
A Google Earth Engine code to analyze residential buildings' real estate values, summer surface thermal anomaly patterns and urban features: a Florence (Italy) case study
<ol> </ol> <p>The layers included in the code were from the study conducted by the research group of CNR-IBE (Institute of BioEconomy of the National Research Council of Italy) and ISPRA (Italian National Institute for Environmental Protection and Research), published by the Sustainability journal (<strong>https://doi.org/10.3390/su14148412</strong>).</p> <p>Link to the <strong>Google Earth Engine (GEE) code</strong> <strong>(link: <a href="https://code.earthengine.google.com/715aa44e13b3640b5f6370165edd3002">https://code.earthengine.google.com/715aa44e13b3640b5f6370165edd3002</a></strong>)</p> <p>You can analyze and visualize the following spatial layers by accessing the GEE link: </p> <ol> <li><strong>Daytime summer land surface temperature</strong> (raster data, horizontal resolution 30 m, from Landsat-8 remote sensing data, years 2015-2019)</li> <li><strong>Surface thermal hot-spot </strong>(raster data, horizontal resolution 30 m) was obtained by using a statistical-spatial method based on the Getis-Ord Gi* approach through the ArcGIS Pro tool.</li> <li><strong>Surface albedo</strong> (raster data, horizontal resolution 10 m, Sentinel-2A remote sensing data, year 2017)</li> <li><strong>Impervious area</strong> (raster data, horizontal resolution 10 m, ISPRA data, year 2017)</li> <li><strong>Tree cover</strong> (raster data, horizontal resolution 10 m, ISPRA data, year 2018)</li> <li><strong>Grassland area</strong> (raster data, horizontal resolution 10 m, ISPRA data, year 2017)</li> <li><strong>Water bodies</strong> (raster data, horizontal resolution 2 m, Geoscopio Platform of Tuscany, year 2016)</li> <li><strong>Sky View Factor</strong> (raster data, horizontal resolution 1 m, lidar data from the OpenData platform of Florence, year 2016)</li> <li><strong>Buildings' units</strong> of Florence (shapefile from the OpenData platform of Florence) include data on the residential real estate value from the Real Estate Market Observatory (OMI) of the National Revenue Agency of Italy (source: https://www1.agenziaentrate.gov.it/servizi/Consultazione/ricerca.htm, accessed on 14 July 2022). Data on the characterization of the buffer area (50 m) surrounding the buildings are included in this shapefile [the names of table attributes are reported in the square brackets]: averaged values of the daytime summer land surface temperature [LST_media], thermal hot-spot pattern [Thermal_cl], mean values of sky view factor [SVF_medio], surface albedo [alb_medio], and average percentage areas of imperviousness [ImperArea%], tree cover [TreeArea%], grassland [GrassArea%] and water bodies [WaterArea%]. </li> </ol> <p>Here attached the .txt file of the <strong>GEE code</strong>. </p> <p> </p> <p><em>E-mail</em></p> <p>Giulia Guerri, CNR-IBE, giulia.guerri@ibe.cnr.it</p> <p>Marco Morabito, CNR-IBE, marco.morabito@cnr.it</p> <p>Alfonso Crisci, CNR-IBE, alfonso.crisci@ibe.cnr.it</p>
San Diego Earthquake Dataset with Feature-Engineered Variables
<p>For the San Diego region, using data from the Southern California Earthquake Data Center (SCEDC), we filtered events by latitude 32.715, longitude -117.1611 within a 150 km radius, focusing on earthquake events from August 1, 2004, 00:00:00 to August 1, 2024, 00:00:00. All magnitude types and depths were included, and 21 variables were feature-engineered to enhance predictive modeling. This dataset provides a robust foundation for earthquake prediction in the San Diego area, incorporating both raw seismic data and advanced engineered features.</p>
Spatial Feature Engineering Dataset for Forest Aboveground Biomass Estimation Using Landsat Imagery
<p><strong>Study Area:</strong><br>The dataset covers forested regions in Oregon, Washington, Idaho, and eastern Montana, characterized by diverse climatic conditions due to orographic effects. The forests in the Coast Range and western slopes of the Cascades, with high precipitation (800-3000 mm annually), contrast with the drier forests in Idaho and Montana, which receive over 400 mm annually. The dataset includes highly productive Douglas-fir and western hemlock forests, with aboveground biomass (AGB) densities exceeding 1200 Mg ha⁻¹, as well as fire-adapted lodgepole and ponderosa pine forests in the rainshadow regions.</p> <p><strong>LiDAR AGB Estimates:</strong><br>The dataset includes 176 lidar-derived AGB maps from 2002 to 2016, covering various regions in Oregon, Washington, Idaho, and Montana. A Random Forest (RF) model was used to estimate AGB at a 30m² resolution, utilizing lidar height features, DEM features, and climate data. Non-forested areas and buildings were masked using binary forest cover maps from the LCMS dataset and the Microsoft Building Footprints dataset.</p> <p><strong>Reference Dataset:</strong><br>A composited AGB map, derived from the 176 lidar maps, was created to develop Landsat-based AGB models, covering 9,361,622 ha of forested land. The AGB layer was stratified into 30 bins, and training, development, and testing sets were constructed for model validation. The dataset includes 7500 test samples and 300,000 training and development samples, with a 500m buffer around test set locations to prevent spatial autocorrelation.</p> <p><strong>Landsat Satellite Imagery:</strong><br>Landsat imagery from 1990 to 2022 was utilized, with three time series derived: all scenes, scenes from May to November, and annual medoid composites. The imagery was processed using the Google Earth Engine (GEE) platform, focusing on periods of maximum phenological activity.</p> <p><strong>Feature Engineering:</strong><br>Extensive feature engineering was performed, generating spectral, spatial, temporal, and topographic features from Landsat imagery and DEM data. Features were extracted over the reference AGB map's domain, synchronized with the lidar acquisition dates.</p> <ul> <li><strong>LandTrendr Fitted Imagery:</strong> Spectral features were derived from LandTrendr-fitted imagery, smoothing variations in the time series.</li> <li><strong>LandTrendr Disturbance and Recovery Features:</strong> Temporal features were derived from LandTrendr models, characterizing disturbance and recovery events.</li> <li><strong>CCDC Disturbance and Recovery Features:</strong> CCDC algorithm-derived features characterized disturbances and recovery using harmonic models.</li> <li><strong>Buffer Features:</strong> Local variations were captured using buffer statistics around each pixel.</li> <li><strong>GLCM Features:</strong> GLCM texture features summarized the joint distribution of gray-tone values.</li> <li><strong>Edge Detectors:</strong> Various edge detection operators captured spatial derivatives and edges.</li> <li><strong>Morphological Operations:</strong> Morphological features were derived using multi-channel image processing techniques.</li> <li><strong>Neighborhood Vectorization:</strong> Direct vectorization of satellite measurements in pixel neighborhoods.</li> <li><strong>Neighborhood Similarity:</strong> Similarity features characterized the relationship between pixel neighborhoods and their centroids.</li> <li><strong>Topographic Features:</strong> Topography was characterized using elevation, slope, aspect embeddings, and hillshade layers from the NED DEM.</li> </ul> <p>This comprehensive dataset enables robust analysis of AGB models and their performance across diverse forested landscapes in the Pacific Northwest</p>
Los Angeles, California, Earthquake Dataset with Feature-Engineered Variables
<p>This dataset includes detailed records of seismic events in Southern California, such as magnitudes, depths, and locations, filtered to focus on a 100 km radius around Los Angeles from January 1, 2012, to September 1, 2024. It also includes a target variable representing the maximum earthquake magnitude within 30 days of each event, along with additional engineered features for use in machine learning and neural network algorithms to improve earthquake forecasting.</p>
Engineering Machine Learning features to predict adsorption of carbon dioxide and nitrogen in metal-organic frameworks
<p>This repository contains CIF files for metal-organic frameworks and Grand canonical Monte Carlo (GCMC) simulation results for the article <em>Engineering Machine Learning features to predict adsorption of carbon dioxide and nitrogen in metal-organic frameworks</em> by Zijun Deng and Lev Sarkisov.</p>
Data for: An advanced systems biology framework of feature engineering for cold tolerance genes discovery from integrated omics and non-omics data in soybean
<p><span>Soybean [<em>Glycine max (L.) Merr.</em>] </span><span>serves as one of the most economically valuable crops globally, but it is sensitive to low temperatures during the crop growing season. Currently, agriculture around the world has faced more serious abiotic stresses due to climate change, so there is an urgent need to breed cold-tolerant cultivars to resist the changing environment. The cold-tolerant trait is a complex and quantitative trait controlled by multiple genes, environmental factors, and their interaction. A total of 56 soybean samples were used, including 28 resistant varieties and 28 susceptible varieties, in the field experiments. We selected 55 SNPs (which were mapped to 39 CTgenes) from the CTgenes for distinguishing cold-tolerant lines from cold-susceptible lines. The SNP data can be applied for soybean's cold-tolerant experiment, such as soybean marker-assisted selection, soybean varieties clustering, the systems biology analysis, and further validation. </span></p>
Data for: An advanced systems biology framework of feature engineering for cold tolerance genes discovery from integrated omics and non-omics data in soybean
Open the record for dataset details and reuse information.
Experimental Data for "What Makes a Top-Performing Precision Medicine Search Engine? Tracing Main System Features in a Systematic Way" at SIGIR2020
<p>This deposit contains data used for the experiments reported in the paper "<a href="https://doi.org/10.1145/3397271.3401048">What Makes a Top-Performing Precision Medicine Search Engine? Tracing Main System Features in a Systematic Way</a>", most notably the ElasticSearch 5.4 indices used for the reported experiments.</p> <p>To load the indices into an ElasticSearch cluster of your own, use the restore function described in the <a href="https://www.elastic.co/guide/en/elasticsearch/reference/5.4/modules-snapshots.html">ElasticSearch documentation</a>.</p> <p>The names of the index snapshots contained here are</p> <ul> <li>ct1718 for the indexed ClinicalTrials data used in the TREC-PM challenges in <a href="http://www.trec-cds.org/2017.html">2017</a> and <a href="http://www.trec-cds.org/2018.html">2018</a>.</li> <li>ct19 for the indexed ClinicalTrials data used in the TREC-PM challenge in <a href="http://www.trec-cds.org/2019.html">2019</a>.</li> <li>ba1718 for the indexed PubMed data used in the TREC-PM challenges in <a href="http://www.trec-cds.org/2017.html">2017</a> and <a href="http://www.trec-cds.org/2018.html">2018</a>.</li> <li>ba19 for the indexed PubMed data used in the TREC-PM challenge in <a href="http://www.trec-cds.org/2019.html">2019</a>.</li> </ul> <p>The other file contains the original output that <a href="https://www.automl.org/automated-algorithm-design/algorithm-configuration/smac/">SMAC</a> wrote to disc during the parameter optimization process. There are directories for the biomedical abstracts (BA) and clinical trials (ct) and for each respective 10 fold cross validation split. Those file contain the exact parameter configurations and their evalation score (the infNDCG metric was used) in live-runXX.json files.</p> <p>The code to these files is located in <a href="https://zenodo.org/record/3856403">this Zenodo deposit</a>.</p>
Feature Engineered Dataset from HPPC Files for Lithium-Ion Battery Cells
<p>This dataset contains time-series data from Hybrid Pulse Power Cycle (HPPC) tests of lithium-ion battery cells, specifically designed for evaluating cell performance and degradation over multiple cycles. The dataset originally published by (Popp et al., 2024) includes information from both charge and discharge tests for each cell, with additional engineered features derived from the raw data. The cells in this dataset were evaluated under different States of Health (SoH), providing insights into performance across various degradation levels.</p> <h2>Dataset Overview</h2> <p>The dataset comprises charge and discharge cycles from 256 individual cells. Each test involves the following procedures:</p> <ol> <li>Cells were equilibrated at 25 °C inside a thermal chamber to ensure a stable temperature before testing.</li> <li>Cells were charged using constant current/constant voltage (CC/CV) mode with a charging rate of C/2 and a cutoff current of 98 mA.</li> <li>After a resting period to allow thermal stabilization, the cells were discharged at a rate of C/5 until the voltage dropped to 2.5 V, capturing the full discharge capacity.</li> <li>A second resting period was conducted until the cells reached thermal equilibrium.</li> <li>HPPC cycles were performed at multiple stages of the charge and discharge cycles to assess cell performance at different State of Charge (SOC) levels.</li> </ol> <p>The dataset also includes a micro HPPC cycle performed at every 10% SOC decrement, starting from 100% SOC down to 10% SOC, with a final HPPC test at 0% SOC. Cells were tested using a combination of charge and discharge pulses to simulate real-world usage patterns. All cycles were recorded at 100 Hz for standard charge/discharge cycles and 1 kHz for the HPPC tests, allowing for high-resolution analysis of the cell behavior.</p> <h2>Data Features</h2> <p>The dataset includes the following features recorded by (Popp. et al., 2024) during the tests:</p> <ul> <li><strong>Time</strong>: Timestamp for each data point.</li> <li><strong>ClimaTemp</strong>: Temperature recorded from the climate chamber.</li> <li><strong>I</strong>: Current applied to the cell (in Amps).</li> <li><strong>Itarget</strong>: Setpoint for the applied current (in Amps).</li> <li><strong>P</strong>: Power output of the cell (in Watts).</li> <li><strong>Q</strong>: Total charge accumulated in the cell (in Amp-seconds).</li> <li><strong>Qneg</strong>: Negative charge accumulated (in Amp-seconds).</li> <li><strong>Qpos</strong>: Positive charge accumulated (in Amp-seconds).</li> <li><strong>Temp_Cell</strong>: Temperature measured at the cell (in °C).</li> <li><strong>U</strong>: Cell voltage (in Volts).</li> </ul> <p>In addition to these core features, engineered features have been included to facilitate analysis of battery performance and degradation trends:</p> <ul> <li><strong>Cumulative_Cycles</strong>: Running count of cycles completed by the cell.</li> <li><strong>Avg_Voltage</strong>: The average voltage of the cell up to the current cycle.</li> <li><strong>Capacity_Fade_Rate</strong>: The rate of capacity degradation over time.</li> <li><strong>Avg_Temperature</strong>: Average temperature experienced by the cell up to the current cycle.</li> <li><strong>Temp_Variation</strong>: Maximum difference in cell temperature over time.</li> <li><strong>High_Temp_Flag</strong>: Binary flag indicating whether the cell temperature exceeded 40°C.</li> <li><strong>Internal_Resistance</strong>: Internal resistance of the cell, calculated from voltage and current data.</li> <li><strong>Power_Consumption_Rate</strong>: Rate at which power is consumed by the cell.</li> <li><strong>Energy_Efficiency</strong>: Efficiency of energy storage, calculated from the ratio of positive to negative charge.</li> <li><strong>Rolling_Avg_Voltage</strong>: Rolling average of the cell voltage over recent cycles.</li> <li><strong>Std_Dev_Voltage</strong>: Standard deviation of cell voltage across cycles.</li> <li><strong>Max_Voltage</strong>: Maximum voltage observed during the test.</li> <li><strong>Min_Voltage</strong>: Minimum voltage observed during the test.</li> <li><strong>Dynamic_Resistance</strong>: Resistance change calculated from voltage and current differentials.</li> <li><strong>Impedance</strong>: Impedance calculated from voltage and current changes.</li> <li><strong>Temp_Coefficient</strong>: Rate of change in power with respect to temperature.</li> <li><strong>Thermal_Runaway_Risk</strong>: Flag indicating the potential for thermal runaway conditions.</li> <li><strong>Effective_Capacity</strong>: Net charge retained by the cell after each cycle.</li> <li><strong>Energy_Throughput</strong>: Total energy delivered by the cell over time, calculated from power and time.</li> </ul> <p>This dataset is suitable for research into lithium-ion battery degradation, performance modeling, and State of Health prediction. The engineered features provide a robust foundation for predictive modeling, including the estimation of Remaining Useful Life (RUL) and optimization of battery usage in practical applications.</p> <p>The dataset has been structured for easy integration with machine learning workflows, with all features provided in CSV format for each test cycle.</p> <h3>Citation:</h3> <p>Popp, A., Spaeth, U., & Schmuelling, B. (2024). Samsung INR21700-50E Capacity and HPPC tests (V1.0) [Data set]. 2024 IEEE Transportation Electrification Conference and Expo (ITEC), Rosemont, IL, USA. Zenodo. <a href="https://doi.org/10.5281/zenodo.10891871" target="_new" rel="noopener">https://doi.org/10.5281/zenodo.10891871</a></p>
The Best of Both Worlds: Combining Learned Embeddings with Engineered Features for Accurate Prediction of Correct Patches
<p>Dataset for Panther</p>
Direct probing of germinal center responses reveals immunological features and bottlenecks for nAb responses to an engineered HIV trimer
GEO Series GSE89148. Macaca mulatta. 24 samples. Type: Expression profiling by high throughput sequencing.
Comparison of Different Feature Engineering Methods for Automated ICD Coding
ClinicalTrials.gov study NCT04849195. IPD Sharing: NO. Countries: 1. Publications: 0.
Visualization Engineering Platform for Pulse Diagnosis of Traditional Chinese Medicine-The Research of Similar Moiré Feature Analyzing Approach Based on Recurrent Neural Network to Process the Measure
ClinicalTrials.gov study NCT04661605. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Visualization Engineering Platform for TCM Pulse Diagnosis - Pulse Diagnosis Based on Federated Learning to Diagnose Slippery and Choppy and Other Pulses Waveform Image Features to Assist in the Study
ClinicalTrials.gov study NCT05630248. IPD Sharing: NO. Countries: 1. Publications: 0.
Chromatin extrusion explains key features of loop and domain formation in wild-type and engineered genomes
GEO Series GSE74072. Homo sapiens. 74 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.