Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data-set for "Reliability and operation cost of underdamped memories during cyclic erasures" (Research Article, No. apxr.202300074) to Advanced Physics Research.
<p><strong>Dataset for the article: "Reliability and operation cost of underdamped memories during cyclic erasures" (Research Article, No. apxr.202300074) to Advanced Physics Research. </strong></p> <ul> <li><strong>FigX.fig: </strong>Matlab format figure source used to draw all of the plots of the article, embedding the all data.</li> <li><strong>evolution_K.m:</strong> Matlab code file that implement the Repeated Erasure model and compute the success rate of the erasure from the temperature evolution computed from the model.</li> <li><strong>Model_Landauer_enchaine.m:</strong> Matlab code file. Function which uses the SE Model to calculate the evolution of all the energy quantities (average) during step 1 and step 2 starting at an initial temperature CI, and also renders the final temperature conditions to be used iteratively in the RE Model.</li> <li><strong>Toy_Model.m: </strong>Theoretical toy model for explaning the Landauer repeted erasures temperature evolution.</li> <li><strong>Analysis_repeated_erasure.m:</strong> Matlab code file to treat raw data .mat files of 45 repeated erasures to compute the energy, work and heat evolution. The average is done on the full raw data samples corresponding to thousands of 45 repeated erasures.</li> <li><strong>XmsRawData.zip:</strong> Raw data of 45 repeated erasures of duration X ms each. Those Matlab data files are meant to be analysed with the Analysis_repeated_erasure code. Among the usefull variables contained in those files we identifiy :</li> </ul> <ol> <li>V: applied protocol, position of center of the wells in nm</li> <li>x: position signal in nm</li> <li>sigma_m: calibration of the standard deviation used by the feedback</li> <li>x0: threshold of the virtual potential</li> <li>tau1: duration of step1</li> <li>tau2: duration of step 2</li> <li>tau4: half resting time between step1 and step2</li> <li>X1: initial half distance between the wells</li> <li>dt: acquisition sampling time</li> </ol> <p><br> </p> <p> </p>
Data set for the article: Modifying twist algorithms for determining the key length of a Vigenere cipher
<p>Data supporting the work in the article: Modifying twist algorithms for determining the key length of a Vigen\`{e}re cipher.</p>
ADVEX - The CarboEurope-Integrated Project Advection Experiment Data Set
<p>Extensive field measurements have been performed at three CarboEurope-Integrated Project forest sites with different topography (Renon/Ritten, Italian Alps, Italy; Wetzstein, Thuringia, Germany; Norunda, Uppland, Sweden) to evaluate the relevant terms of the carbon balance by measuring CO2 concentrations [CO2] and the wind field in a 3D multitower cube setup. The same experimental setup (geometry and instrumentation) and the same methodology were applied to all the three experiments. Refer to Feigenwinter et al. (2008), <a href="https://doi.org/10.1016/j.agrformet.2007.08.013">https://doi.org/10.1016/j.agrformet.2007.08.013</a> for more details of the ADVEX campaign.</p>
Simulated cycling data set and musculoskeletal models
<p><span>This study used musculoskeletal modelling to explore the relationship between cycling conditions (power output and cadence) and muscle activation and metabolic power. We hypothesized that the cadence that minimized the simulated average active muscle volume would be higher than that which minimized the simulated metabolic power. We validated the simulation by comparing predicted muscle activation and fascicle velocities with experimental electromyography and ultrasound images. We found strong correlations for averaged muscle activations and moderate to good correlations for fascicle dynamics. These correlations tended to weaken when analyzed at the individual participant level. Our study revealed a curvilinear relationship between average active muscle volume and cadence, with the minimum active volume being aligned to the self-selected cadence. The simulated metabolic power was consistent with previous results and was minimized at lower cadences than that which minimized active muscle volume across power outputs. Whilst there are some limitations to the musculoskeletal modelling approach, the findings suggest that minimizing active muscle volume may be a more important factor than minimizing metabolic power for self-selected cycling cadence preferences. Further research is warranted to explore the potential of an active muscle volume-based objective function for control schemes across a wider range of cycling conditions.</span></p>
Data set for model validation in "Simulating ice segregation and thaw consolidation in permafrost environments with the CryoGrid community model"
<p>This upload contains the data set for model validation in the manuscript "Simulating ice segregation and thaw consolidation in permafrost environments with the CryoGrid community model".</p>
Data Set for Enhanced Performance Prediction of ATL Model Transformations
<p>Model transformation languages are domain-specific languages, which are designed to comfortably define transformations. With the increasing use of transformations in various domains, the complexity and size of input models are also increasing. However, developers often lack suitable models for performance testing. We have therefore conducted experiments in which we predict the performance of model transformations based on characteristics of input models using machine learning approaches. In particular, we focused on how to predict the performance of transformations that also transform attributes whose values can have arbitrary size. This dataset contains our raw and processed input data, the scripts necessary to repeat our experiments, and the results we obtained.</p> <p>Our input data consists of the time measurements for six different transformations defined in the Atlas Transformation Language (ATL), as well as the collected characteristics of the real-world input models we used. In this data set, we provide the script that implements our experiments. We predict the execution time of ATL transformations using the machine learning approaches linear regression, random forests and support vector regression using a radial basis function kernel. We also investigate different sets of characteristics of input models as input for the machine learning approaches. These are described in detail in the provided documentation.pdf. The results of the experiments are provided as raw data in individual cvs files. Furthermore, we provide our Eclipse plugin, which collects the characteristics for a set of given models.</p> <p>A detailed documentation is available in documentaion.pdf.</p>
Data set for "Protracted localization of metamorphism and deformation in a heterogeneous lower-crustal shear zone"
<p>Electron micro probe raw data used in the publication: "Protracted localization of metamorphism and deformation in a heterogeneous lower-crustal shear zone" (Zertani et al., 2023).</p> <p>Zertani, S., Menegon, L., Pennacchioni, G., Buisman, I., Corfu, F., Jamtveit, B. (2023). Protracted localization of metamorphism and deformation in a heterogeneous lower-crustal shear zone. Journal of Structural Geology, 104960. https://doi.org/10.1016/j.jsg.2023.104960</p>
Data set "Structural Dependence of Extended Amide III Vibrations in Two-Dimensional Infrared Spectra"
<p>This data set accompanies the publication "Structural Dependence of Extended Amide III Vibrations in Two-Dimensional Infrared Spectra" by Julia Brüggemann, Maria Chekmeneva, Mario Wolter, and Christoph R. Jacob (TU Braunschweig, Germany)</p> <p>It contains the following files:</p> <p>Directory 01_structures:</p> <p> - xyz files of all considered molecular structures.</p> <p>Directory 02_2d-ir_data_and_code:</p> <p> - updated version of the code of calculating 2D-IR spectra, initially published in DOI: 10.5281/zenodo.7328312, which can be used to obtain the 2D-IR spectra shown in the publication and its supporting information.</p>
Kenai Mountains to Sea: Using Thermal Infrared Imagery to Implement Long-Term Salmon Conservation - thermal imagery data set
<p>Cook Inletkeeper contracted NV5 Geospatial (formerly, Quantum Spatial Incorporated) to collect thermal infrared (TIR) during the summer of 2020 along four streams on the Kenai Peninsula in southern Alaska: Beaver Creek, Crooked Creek, Funny River, and Moose River under a project name “Kenai Rivers”. All streams were contracted to be flown in the summer of 2020 and during the afternoon hours in order to maximize the thermal contrast between the river’s water and the banks. The survey extends for a total length of 59.1 km miles of the Kenai Rivers. The Data were collected to aid the Cook Inletkeeper team to identify the spatial variability in surface temperatures as well as thermal influence of point sources, tributaries, and surface springs. The data will also be used to identify high-value habitats for the salmonids population within the four streams.</p> <p>Note: These data and related items of information have not been formally disseminated by NOAA and do not represent any agency determination, view, or policy.</p> <p>Funding for this project came, in part, from the Alaska Sustainable Salmon Fund (AKSSF Project #53003).</p>
Data sets for "Dynamical bulk-boundary correspondence and dynamical quantum phase transitions in higher-order topological insulators"
<p>Data sets for the return rates and Loschmidt eigenvalues for the paper "Dynamical bulk-boundary correspondence and dynamical quantum phase transitions in higher-order topological insulators". Included are data for both bulk calculations and finite open systems labelled OBC. Data files are named with the quench parameters, see the article for details.</p>
Data set of solar irradiation from Svalsat, Svalbard
<p>The data collected with a set of SP110 Apogee pyranometers.</p>
Data sets and machine learning models for: Machine learning from quantum chemistry to predict experimental solvent effects on reaction rates
<p>The datasets and final machine learning model files for the manuscript "Machine learning from quantum chemistry to predict experimental solvent effects on reaction rates". Citation should refer directly to the manuscript:</p> <ul> <li>Chung, Y.; Green, W. H. Machine learning from quantum chemistry to predict experimental solvent effects on reaction rates. <em>Chemical Science </em><strong>2024,</strong> doi: <a href="https://doi.org/10.1039/D3SC05353A">10.1039/D3SC05353A</a></li> </ul> <p>To use the machine learning models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/RxnSolvKSE_ML">https://github.com/yunsiechung/chemprop/tree/RxnSolvKSE_ML</a>. </p> <p>Detailed information can be found in README.md file.</p> <p><br><strong>Details on the files</strong></p> <p>In the pretraining and finetuning set csv files, each column represents:</p> <ol> <li>rxn_smiles: atom-mapped reaction SMILES</li> <li>solvent_smiles: solvent SMILES</li> <li>ddGsolv: solvation free energy of activation of a reaction-solvent pair at 298K in kcal/mol (main prediction target)</li> <li>ddHsolv: solvation enthalpy of activation of a reaction-solvent pair at 298K in kcal/mol (main prediction target)</li> <li>dGsolv_reactant: solvation free energy of reactant(s) at 298K in kcal/mol (additional feature)</li> <li>dGsolv_product: solvation free energy of product(s) at 298K in kcal/mol (additional feature)</li> <li>dHsolv_reactant: solvation enthalpy of reactant(s) at 298K in kcal/mol (additional feature)</li> <li>dHsolv_product: solvation enthalpy of product(s) at 298K in kcal/mol (additional feature)</li> </ol> <p><strong>Data sets under 'RxnSolvKSE_dataset_v1.1.zip'</strong></p> <ul> <li>pretraining_set: contains the dataset used for pre-training <ul> <li>all_data: contains all calculated data <ul> <li>pretraining_rxn_solvent_ddGsolv_ddHsolv_with_features_all.csv: contains both main prediction targets and additional feature for reaction-solvent pairs</li> <li>pretraining_solvent_info.csv: list of all solvents</li> <li>pretraining_unique_rxn.csv: list of all reactions, both forward and reverse directions</li> </ul> </li> <li>chosen_500k_data: contains the chosen 500k data <ul> <li>pretraining_rxn_solvent_ddGsolv_ddHsolv_500k.csv: contains main prediction targets for reaction-solvent pairs</li> <li>pretraining_features_react_prod_dGsolv_dHsolv_500k.csv: contains additional features for reaction-solvent pairs</li> <li>train_test_split: contains the 5-fold random split training and test sets.</li> </ul> </li> </ul> </li> <li>finetuning_set: contains the dataset used for fine-tuning <ul> <li>all_data: contains all calculated data <ul> <li>finetuning_rxn_solvent_ddGsolv_ddHsolv_with_features_all.csv: constains both main prediction targets and additional features for reaction-solvent pairs. The rxn_key column indicates whether the reaction is bimolecular hydrogen abstraction (bihabs), unimolecular hydrogen migration (intrahabs), or radical addition to a multiple bond (raddition). The 'fwd' and 'rev' each indicate forward and reverse reactions.</li> <li>finetuning_solvent_info.csv: list of all solvents</li> <li>finetuning_unique_rxn.csv: list of all reactions, both forward and reverse directions</li> </ul> </li> <li>chosen_data: contains chosen data <ul> <li>finetuning_rxn_solvent_ddGsolv_ddHsolv_chosen.csv: contains main prediction targets for reaction-solvent pairs</li> <li>finetuning_features_react_prod_dGsolv_dHsolv_chosen.csv: contains additional features for reaction-solvent pairs</li> </ul> </li> </ul> </li> <li>experimental_set: contains the experimental rate constant data used to test the model. The original experimental data can be found at <a href="../record/7747557">https://zenodo.org/record/7747557</a>. <ul> <li> expt_rxn_atom_mapped_smiles.csv: contains the atom-mapped reaction SMILES used for the experimental data.</li> <li>expt_data_collected.xlsx: contains all experimental data and detailed information</li> <li>expt_rxn_solv_smiles_with_features_all.csv: contains the computed additional features for the experimental reaction-solvent pairs.</li> </ul> </li> </ul> <p><strong>Machine learning model files under 'RxnSolvKSE_ML_model_files.zip'</strong></p> <ul> <li>Contains the Chemprop machine learning model files for predicting ddGsolv and ddHsolv for a reaction-solvent pair. It takes atom-mapped reaction SMILES and solvent SMILES as inputs.</li> <li>To use these ML models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/RxnSolvKSE_ML">https://github.com/yunsiechung/chemprop/tree/RxnSolvKSE_ML</a></li> </ul>
Data sets of measured cross-sectional area coordinates and material properties of an old Vindeby blade
<p>Measured raw data for determining cross-sectional area coordinates of an old Vindeby blade. Material properties are also appended.</p>
Data set: IEC 60270 Calibration Uncertainty in Gas-Insulated Substations
<p>Data set for the publication named: IEC 60270 Calibration Uncertainty in Gas-Insulated Substations.</p>
Chimeras benchmark data set
<p>The Chimera data set is a benchmark for astronomical software working with multi-wavelength photometry of distant galaxies and quasars. The Chimera data set provides ground truth galaxy and AGN properties, and realistic broad-band photometry (Ultraviolet to Far-infrared). It is a benchmark for validating SED fitting codes that try to infer properties of the host galaxy with an active galactic nucleus or quasar.</p>
Data set for Metrological Qualification of PD Analysers for Insulation Diagnosis of HVDC and HVAC Grids
<p>Dataset for the publication named: "Metrological Qualification of PD Analysers for Insulation Diagnosis of HVDC and HVAC Grids"</p>
Data sets for "On the detection of a solar radio burst event that occurred on 28 August 2022 and its effect on GNSS signals as observed by ionospheric scintillation monitors distributed over the American sector" by Wright et al.
<p>Copy of data described in 'On the detection of a solar radio burst event that occurred on 28 August 2022 and its effect on GNSS signals as observed by ionospheric scintillation monitors distributed over the American sector' by Wright et. al and to appear in the Journal of Space Weather and Space Climate (JSWSC), 2023.</p><p> </p>
Dataset for "Enabling Machine Learning Models in Alarm Fatigue Research: Creation of a Large Relevance-annotated Oxygen Saturation Alarm Data Set"
<p>Chromik and Flint et al. (2024) (under review) propose an algorithm that uses clinical alarm logs, an annotation guideline (Klopfenstein et al. 2023), and routinely collected intensive care data to create a data set of relevance-annotated oxygen saturation alarms. We provide the algorithm's source code and data set of annotated oxygen saturation alarms as supplementary material to the publication.</p> <ul> <li>The algorithm's implementation is open-source and can be re-used on similar data sets.</li> <li>Our implementation used airway management data mappings to identify airway devices (AD), ventilation devices (VD), and ventilation modes (VM). These mappings can be found here: <a href="../doi/10.5281/zenodo.7511031">https://zenodo.org/doi/10.5281/zenodo.7511031</a></li> <li>The data set suggests that the majority of oxygen saturation alarms in the intensive care unit is non-actionable.</li> <li>We are the first to provide such an extensive data set of annotated oxygen saturation alarms.</li> </ul>
DATA SET: In vivo characterization of the optical and hemodynamic properties of the human sternocleidomastoid muscle through ultrasound-guided hybrid near-infrared spectroscopies.
<p>This repository contains the data sets of the article:</p><p>L. Cortese et al., "In vivo characterization of the optical and hemodynamic properties of the human sternocleidomastoid muscle through ultrasound-guided hybrid near-infrared spectroscopies."</p>
FT-IR data set for hierarchical cluster-based deep learning
<p>FT-IR of raw materials described in Kristoffersen et al. (2019) doi:<a href="https://doi.org/10.1016/j.talanta.2019.06.084">10.1016/j.talanta.2019.06.084</a>. Predicts average molecular weight after hydrolysis. The data is pre-processed using Savitzky-Golay 2nd derivative smoothing (with window width 11 pt and 3rd order polynomial smoothing) followed by extended multiplicative signal correction (EMSC). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.