Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
DATA SET HI VOLTAGE LINE LOADING AND REGIONAL POWER PRODUCTION
<p>DATA SET CORRESPONDING TO</p> <pre>10.5281/zenodo.3570503</pre> <p> </p> <p>Local Balancing of Low-Voltage Networks by Utilizing Distributed Flexibilities as Part of the InterFlex Field Trial</p>
Biobox YAML file for the CAMI 2 Mouse Gut Toy data set, samples 0-63
<p>Biobox YAML file for the CAMI 2 Mouse Gut Toy data set, samples 0-63</p>
High-CAPE summer convection in large-domain large-eddy simulations with ICON - model and observational data sets
<p>Data sets including all observational and ICON model data for publication in Atmosperic Chemistry and Physics Journal (ACP) - "High-CAPE summer convection in large-domain large-eddy simulations with ICON"</p>
Average genome coverage of the CAMI 2 Mouse Gut Toy data set
<p>Genome coverage (short reads) averaged over the 64 samples of the CAMI 2 Mouse Gut Toy data set</p>
Periodic dynamic induction control of wind farms: proving the potential in simulations and wind tunnel experiments - data sets
<p>Data sets of the wind tunnel experiments described in "Periodic dynamic induction control of wind farms: proving the potential in simulations and wind tunnel experiments". DOI: https://doi.org/10.5194/wes-2019-50</p>
Data Set to 'Sodium-induced population shift drives activation of thrombin'
<p>The data set 10.5281/zenodo.3688506 contains the raw data used for the preparation of the manuscript 'Sodium-induced population shift drives activation of thrombin' (doi:10.1038/s41598-020-57822-0). To limit the required storage space, the trajectories are deposited only as coordinates of the analysed loop. The data set is separated into following parts:</p> <ul> <li>X-ray: Data used in the analysis of the PDB structures of thrombin, including a list of the PDB IDs, results of the PCA and values of the introduced features in the structures.</li> <li>cMD: Input parameters, topologies and starting coordinates of classical MD simulation based on different PDB structures (3bei, 3lu9 – with Na<sup>+</sup> ions and without Na<sup>+</sup> ions), resulting trajectories, projection on X-ray PCA and values of the introduced features during the simulations.</li> <li>TMD: Input parameters, topologies and starting coordinates of the targeted MD simulations, resulting trajectories, projection on X-ray PCA and RMSD values.</li> <li>Seeded-Simulations: Input parameters, topologies and starting coordinates of classical MD simulations that are started from cluster representatives of the TMD simulations and resulting trajectories (3*100 simulations, each lasting 200 ns), separated into simulations with Na+ and without Na+, additionally values of the introduced features during the simulations.</li> <li>MSM: For both systems (with and without Na<sup>+</sup>), files are provided that can be loaded with pyemma (TICA objects, kmeans clusterings, calculated implied timescales and MSMs); additionally TICs of the seeded simulations, psi-dihedrals of the seeded simulations, RMSD values of the E and E<sup>*</sup> states and representative structures of the E and E<sup>*</sup> states</li> </ul> <p>A more detailed description of the included files can be found in the README files in each folder.</p>
Arsenic adsorption modelling: TiO2, Fe2O3 and composite TiO2-Fe2O3 sorbents - detailed characterisation and adsorption data sets
<p>Raw data, example files and templates for the research paper provisionally titled 'Improved Accuracy in the Surface Complexation Modelling of Arsenic on Multicomponent Sorbents Using Low Energy Ion Scattering'</p> <p>Included is</p> <ul> <li>materials characterisation data (TiO2, Fe2O3 and a TiO2-Fe2O3 composite)</li> <li>arsenic(III) and arsenic(V) adsorption data (both adsorption isotherms and pH adsorption edges)</li> <li>potentiometric titration data</li> <li>templates for preparing FITEQL input files from titration data and pH adsorption edges</li> <li>examples of FITEQL input and output files</li> </ul> <p>Characterisation techniques used include BET, XRD, FTIR, zeta potential, DLS, low energy ion scattering (LEIS), XPS and XRF. Data was primarily collected by Jay Bullen between 2016 and 2019, with assistance from named collaborators.</p>
Data set to ''Volcano growth versus deformation by strike-slip faults: morphometric characterization through analogue modelling'
<p>This data set is the supplementary material to Grosse et al. (2020) 'Volcano growth versus deformation by strike-slip faults: morphometric characterization through analogue modelling', published in Tectonophysics (https://doi.org/10.1016/j.tecto.2020.228411). The data set consists of (1) 249 digital elevation models (DEMs) of each step of the the analogue experiments carried out, in standard ENVI format, zipped; and (2) an Excel file containing the DEM-derived morphometric parameters for each of the analogue models.</p> <p>Experiments were carried out at the analogue modelling lab of the Department of Geography at the Vrije Universiteit Brussel (Belgium). A granular mixture of fine-grained quartz sand and kaolin clay was used as analogue material. Experiments were conducted on a fixed table, on which a basal layer of granular material was placed. A basal plate attached to a step-motor was used to simulate pure strike-slip displacements of the basal layer. Volcano growth was simulated by depositing loads of granular material on top of the basal layer from a point source. The analogue models were photographed at regular time intervals during the experiments using four digital cameras. The photographs were used to generate synthetic digital elevation models (DEMs) with 0.2 mm spatial resolution of each step of the analogue models by applying the MICMAC digital stereo-photogrammetry software. The ENVI software was used to re-sample the DEMs to a 0.5 mm spatial resolution and apply the noise-reduction Lee filter. Morphometric data were then extracted from the DEMs by applying two IDL-language algorithms: NETVOLC, used to automatically calculate the volcano edifice basal outline, and MORVOLC, used to extract a set of morphometric parameters.</p>
MS data set: Identification of Microorganisms by Liquid Chromatography-Mass Spectrometry (LC-MS1) and in silico Peptide Mass Data
<p>Data set consisting of raw LC-MS2 data, LC-MS1 peak data and a description</p> <p>For unreviewed publication preprint: <strong>Identification of Microorganisms by Liquid Chromatography-Mass Spectrometry (LC-MS<sup>1</sup>) and <em>in silico </em>Peptide Mass Data</strong></p> <p>ABSTRACT</p> <p>Over the past decade, modern methods of mass spectrometry (MS) have emerged that allow reliable, fast and cost-effective identification of pathogenic microorganisms. While MALDI-TOF MS has already revolutionized the way microorganisms are identified, recent years have witnessed also substantial progress in the development of liquid chromatography (LC)-MS based proteomics for microbiological applications. For example, LC-tandem mass spectrometry (LC-MS<sup>2</sup>) has been proposed for microbial characterization by means of multiple discriminative peptides that enable identification at the species, or sometimes at the strain level. However, such investigations can be very time-consuming, especially if the experimental LC-MS<sup>2</sup> data are tested against sequence databases covering a broad panel of different microbiological taxa.</p> <p>In this proof of concept study, we present an alternative bottom-up proteomics method for microbial identification. The proposed approach involves efficient extraction of proteins from cultivated microbial cells, digestion by trypsin and LC-MS measurements. MS<sup>1</sup> data are then extracted and systematically tested against an in silico library of peptide mass data compiled in house. The library has been computed from the UniProt Knowledgebase Swiss-Prot and TrEMBL databases and comprises more than 12,000 strain-specific in silico profiles, each containing tens of thousands of peptide mass entries. Identification analysis involves computation of score values derived from spectral distances between experimental and in silico peptide mass data and compilation of score ranking lists. The taxonomic positions of the microbial samples are then determined by using the best-matching database entries. The suggested method is computationally efficient – less than two minutes per sample - and has been successfully tested by a set of 19 different microbial pathogens. The approach is rapid, accurate and automatable and holds great potential for future microbiological applications.</p> <p><em>For details see the following preprint: Lasch, P. Schneider, A. Blumenscheit, C. and Doellinger, J. “Identification of Microorganisms by Liquid Chromatography-Mass Spectrometry (LC-MS1) and in silico Peptide Mass Data”. bioRxiv preprint, http://dx.doi.org/10.1101/870089</em></p> <p> </p>
Reproduction data set and code for the results in Reichl et al., 2020
<p>This is the cleaned estimation dataset used to reproduce the results in Reichl et al., 2020. The data are contained in "ClimateCertaintyRaw.csv". The R file is the Bayesian estimation of the econometric model. The .txt file gives the Mplus 8.2 code for reproducing the psychometric structural equation model. The full survey text and programming instructions are included as a PDF for reference.</p>
DATA SET FOR: Active faulting, submarine surface rupture and seismic migration along the Liquiñe-Ofqui fault system, Patagonian Andes
<p>Data description: These data corresponde to high-resolution bathymetry and seismic reflection profiles obtained in the inner fjord west of Puerto Aysén (between 73.13°- 72.68°W and 45.32°-45.47°S; Figs. 1 and 2). The data set was obtained during a geophysical study as part of the DETSUFA project (Deslizamientos Tsunamigénicos en el Fiordo de Aysén; Lastras et al., 2013), which took place between March 4th and 17th, 2013, aboard the R/V BIO Hésperides.<br> <br> KONGSBERG SIMRAD multibeam EM-1002S was used to obtain bathymetric data, and it works with 111 beams at a 96 kHz sonar frequency and with a maximum ping rate of >10 Hz. Equidistant mode was used for swath bathymetry acquisition. This array maximized the number of beams facilitating data acquisition and obtaining a homogenized final grid with improved resolution, with tracks separated every 150 m. The swath thickness was the same regardless of width, generating a 50% overlap between each track, with the exception of areas located near the coast. Expendable Bathythermograph (XBT) probes were used at specific sites to measure changes in water sound velocity due to eventual changes in fresh water circulation, tides, and sediment.<br> <br> Seismic reflection data were acquired using an array of two BOLT air guns (165 and 175 inches3), which were towed behind the vessel stern. The configuration used in the seismic sources was 2,000 psi, a depth of 3 m for the gun, with a firing rate of 15 m over the seafloor. A 100 m long mini-streamer with a 25 m active section, corresponding to one single channel, recovered the shots. The seismic data were recorded by using the DELPH SEISMICPLUS system with a recording length of 4.0 s and a preamplifier gain of 8 Hz. The raw seismic data were processed aboard the SMT Kingdom Suite, including the navigation and standard processes of electrical noise removing (50 Hz filter), gain amplifier and bandpass filtering, to improve data visualization. Postprocessing included the migration of the sea bottom diffractions and the muting of the water column performed in Seismic-Unix.</p> <p>Files:</p> <p>Raw Seismic reflection data for lines 05, 06 and 07 (SU & SEG files)</p> <p>Masked Seismic profiles for lines 05, 06 and 07 (SU, PDF & PS files)</p> <p>Bathymetry of inner and outer Aysén Fjord (ASCII file)</p>
Experimental Data Sets for the study "Benchmarking a $(\mu+\lambda)$ Genetic Algorithm with Configurable Crossover Probability"
<p>This is the experimental result of the study "Benchmarking a (μ+λ) Genetic Algorithm with Configurable Crossover Probability". A novel (μ+λ) GA is proposed and benchmarked, in which we stochastically determine whether to apply the crossover operator either for each individual or generation with a crossover probability <span class="math-tex">\(p_c\)</span>. This data set consists of two parts:</p> <ol> <li>The results of (μ+λ) GA on 25 pseudo-Boolean problems defined in <em>IOHprofiler </em>(<a href="https://iohprofiler.github.io/">https://iohprofiler.github.io/</a>) with the following setup: <span class="math-tex">\(\mu \in \{10, 50, 100\}, \lambda \in \{1, \lceil\mu/2\rceil, \mu\}, p_c\in\{0, 0.5\}.\)</span> <ul> <li>'IOHprofiler_Problems_standard_bit_mutation.csv' --> the (μ+λ) GA with standard bit mutation.</li> <li>'IOHprofiler_Problems_fast_mutation.csv' --> the (μ+λ) GA with fast mutation.</li> </ul> </li> <li>The results of (μ+λ) GA on OneMax and LeadingOnes problems with the following setup: <span class="math-tex">\(n \in \{64,100,150,200,250,500\}, \mu \in \{2,3,5,8,10,20,30,...,100\}, \\ \lambda \in \{1, \lceil \mu/2 \rceil, \mu\}, \text{and }p_c \in \{0.1 k \mid k \in [0..9]\}\cup\{0.95\}.\)</span> <ul> <li>'OneMax_raw.csv' --> the fixed-target running time/first hitting time from 100 independent runs for target values in <span class="math-tex">\([1..n]\)</span>.</li> <li>'OneMax_summary.csv' --> the mean, median, standard deviation, some quantiles, expected running time (ERT), the number of successful runs, and the success rate from 100 independent runs for target values in <span class="math-tex">\([1..n]\)</span>.</li> <li>'LeadingOnes_raw.csv' --> the same with 'OneMax_raw.csv' for LeadingOnes.</li> <li>'LeadingOnes_summary.csv' --> the same with 'OneMax_summary.csv' for LeadingOnes.</li> </ul> </li> </ol> <p><strong>Contact</strong>: if you have any questions or suggestions, please feel free to contact <a href="https://www.universiteitleiden.nl/en/staffmembers/furong-ye#tab-1">Furong Ye</a> or <a href="http://www-ia.lip6.fr/~doerr/">Carola Doerr</a>.</p>
Data set for: Adjustable Deterministic Pseudonymization of Speech Listening Experiment, Report of listening experiments
<p>Data set used in "Adjustable Deterministic Pseudonymization of Speech Listening Experiment". Includes Rmarkdown script.</p> <p> </p>
DATA SET FOR PUBLICATION: Structure Determination of Hen Egg-White Lysozyme Aggregates Adsorbed to Lipid/Water and Air/Water Interfaces
<p>The data set collected for the publication: "Structure Determination of Hen Egg-White Lysozyme Aggregates Adsorbed to Lipid/Water and Air/Water Interfaces" (<a href="https://doi.org/10.1021/acs.langmuir.9b03826">https://doi.org/10.1021/acs.langmuir.9b03826</a>).</p>
MarTREC Data Set for Report: Developing and Applying an Analysis Methodology to Identify Flow Generation Influences between Vessel and Truck Shipments
<p>Truck activity is logically connected to vessel activity at a port. In turn, vessel activity is also influenced by truck shipments. Although one might expect a direct and straightforward relation between these two types of shipments, that is rarely the case. For instance, many maritime containers carry consolidated cargos that have multiple and different final destinations. Also, different truck capacities, customs clearance and regulations play a critical role in determining the actual relation between these two types of shipments. This project aims at shedding light on the nuances of maritime and roadway flow relations by quantitatively analyzing the linkages between these two types of shipments.</p> <p>The study performed a statistical analysis to determine the probability distributions of vessel and truck activity, and then explore the correlation of each activity with the other. The analysis yielded coefficients that function as explanatory values for specific truck flows.</p> <p>The ultimate purpose of this study is to provide a clearer and quantitative understanding of the relationship between maritime and truck shipments, and by doing so, to provide tools to develop a system for managing trucks that maximizes efficiency for industry, while minimizing industry’s negative impacts on a region.</p> <p>For this purpose, the study selected the Port Freeport as a case study.</p>
Data Sets for Measuring and Modeling the Performance Configurations of Distributed DBMS
<p>These data sets contain the performance measurements and additional metadata as accompanying material for the research paper <strong>Baloo: Measuring and Modeling the Performance Configurations of Distributed DBMS</strong><em> </em>that is published in the <em>Symposium on Modelling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS) 2020.</em></p> <p>The attached readme describes the data set structure.</p>
Dale_Wonsho demographic data set
<p>These are the data sets used for analyzing demographic data for the manuscript entitled " <strong>Births and deaths in Sidama in southern Ethiopia: Findings from the 2018 Dale-Wonsho Health and Demographic Surveillance System (HDSS)". This data set is used for the first demographic study done in the newly established HDSS site by Hawassa University.</strong></p>
Statistics of turbulent tracer dispersion from UV camera observations of SO 2, Data set
<p>LES and radiation transport data sets used to produce the figures in the manuscript:</p> <p>Kylling, A., Ardeshiri, H., Cassiani, M., Dinger, A. S., Park, S.-Y., Pisso, I., Schmidbauer, N., Stebel, K., and Stohl, A.: Can statistics of turbulent tracer dispersion be inferred from camera observations of SO2 in the ultraviolet?, Atmos. Meas. Tech. Discuss., https://doi.org/10.5194/amt-2019-286, in review, 2019</p> <p>See README file for further description of files.</p>
Data set for "Anatomically and functionally distinct thalamocortical inputs to primary and secondary mouse whisker somatosensory cortices"
<p>Data set for: El-Boustani S, Sermet BS, Foustoukos G, Oram TB, Yizhar O, Petersen CCH (2020) Anatomically and functionally distinct thalamocortical inputs to primary and secondary mouse whisker somatosensory cortices. Nature Communications 11: 3342. doi: 10.1038/s41467-020-17087-7</p> <p>There are 4 files in this upload:</p> <p>1. The file named "2020_El-Boustani_NCOMMS.pdf" is the Open Access pdf file of the manuscript published in Nature Communications.</p> <p>2. The file named "2020_El-Boustani_NCOMMS_SupMovie1.avi" is Supplementary Movie 1 in .avi format, accompanying the Nature Communications publication.</p> <p>2. The file named "2020_El-Boustani_NCOMMS_SupMovie2.avi" is Supplementary Movie 2 in .avi format, accompanying the Nature Communications publication.</p> <p>4. The file named "El-Boustani_data_code.zip" (~36 GB) is a zipped version of a folder "El-Boustani_data_code" (~45 GB), which contains the data analysed in the study along with the Matlab code used to generate the published figures. To access the data and the code, first unzip the file. In the main folder, data for all experiments are stored within folders starting by the prefix “SB”. The code for generating population data and plotting figures from the paper are in the folder “Matlab_code”. In this folder, several Matlab scripts are named after the panels or figures they will plot such as “Plot_Fig3d_Axon_GCaMP6s_traces_example.m”. After opening each file, executing the script will automatically plot the panels and name them accordingly. In some files, the type of data to plot should be specified at the very beginning of the script: “VPM” for VPM data, “POMf” for POm-FO data and “Layer1” for POm-HO data in layer 1. For figure 2, the code is located in a dedicated folder where a Matlab file “Plot_Fig2d_g_Populatin_Plot_POm_FO_HO.m” is used to generate the figures. Finally, other Matlab files are included that are used to create population .mat files or for additional analysis related to the manuscript.</p>
M4DB Sample Data Set
<p>A sample data set for the MicroMagnetic-Metadata DataBase (M4DB).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.