Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
MER1 APXS Xray Spectra (Partially Processed) Data Bundle
This bundle contains Xray Spectra (Partially Processed) data from the Alpha Particle X-ray Spectrometer on Mars Exploration Rover 1 (Opportunity). Raw APXS data have been converted to ASCII tables and spectra from previous integrations have been removed.
BRIC-23 GeneLab Process Verification Test: Bacillus subtilis transcriptomic proteomic and metabolomic data
Microbes interact with humans in complex ways and understanding how they respond to the spaceflight environment is important to the success of future manned spaceflight missions. The BRIC-23 mission was designed to measure the response of Bacillus subtilis and Staphylococcus aureus to the spaceflight environment. This experiment aimed to produce high quality omics data from B. subtilis and S. aureus grown aboard the International Space Station (ISS) to allow comparison to matched ground controls. There were two primary objectives for this experiment: (1) Demonstrate all post-flight processes and operations required for successful completion of GeneLab Reference Missions conducted on ISS and (2) Generate high quality GeneLab Reference Mission omics data sets for two prokaryotic model organisms Bacillus subtilis and Staphylococcus aureus. Freezing Control Experiment: The BRIC hardware has significant thermal inertia thus the freezing rate of samples placed at -80 C is quite slow. This could affect RNA-sequencing proteomic and metabolic data sets. In an effort to understand how slow freezing could affect these data sets a control experiment was designed in which B. subtilis and S. aureus were grown in petri plates and either slow frozen to -80 C at a rate matching the BRIC-23 spaceflight samples or processed immediately to harvest RNA and protein.
BRIC-23 GeneLab Process Verification Test: Staphylococcus aureus transcriptomic, proteomic, and metabolomic data
Microbes interact with humans in complex ways and understanding how they respond to the spaceflight environment is important to the success of future manned spaceflight missions. The BRIC-23 mission was designed to measure the response of Bacillus subtilis and Staphylococcus aureus to the spaceflight environment. This experiment aimed to produce high quality omics data from B. subtilis and S. aureus grown aboard the International Space Station (ISS) to allow comparison to matched ground controls. There were two primary objectives for this experiment: (1) Demonstrate all post-flight processes and operations required for successful completion of GeneLab Reference Missions conducted on ISS, and (2) Generate high quality GeneLab Reference Mission omics data sets for two prokaryotic model organisms, Bacillus subtilis and Staphylococcus aureus. Freezing Control Experiment: The BRIC hardware has significant thermal inertia, thus the freezing rate of samples placed at -80 C is quite slow. This could affect RNA-sequencing, proteomic and metabolic data sets. In an effort to understand how slow freezing could affect these data sets, a control experiment was designed in which B. subtilis and S. aureus were grown in petri plates and either slow frozen to -80 C at a rate matching the BRIC-23 spaceflight samples or processed immediately to harvest RNA and protein. B.subtilis omics data is deposited in GLDS-138.
BRIC-23 GeneLab Process Verification Test: Staphylococcus aureus transcriptomic proteomic and metabolomic data
Microbes interact with humans in complex ways and understanding how they respond to the spaceflight environment is important to the success of future manned spaceflight missions. The BRIC-23 mission was designed to measure the response of Bacillus subtilis and Staphylococcus aureus to the spaceflight environment. This experiment aimed to produce high quality omics data from B. subtilis and S. aureus grown aboard the International Space Station (ISS) to allow comparison to matched ground controls. There were two primary objectives for this experiment: (1) Demonstrate all post-flight processes and operations required for successful completion of GeneLab Reference Missions conducted on ISS and (2) Generate high quality GeneLab Reference Mission omics data sets for two prokaryotic model organisms Bacillus subtilis and Staphylococcus aureus. Freezing Control Experiment: The BRIC hardware has significant thermal inertia thus the freezing rate of samples placed at -80 C is quite slow. This could affect RNA-sequencing proteomic and metabolic data sets. In an effort to understand how slow freezing could affect these data sets a control experiment was designed in which B. subtilis and S. aureus were grown in petri plates and either slow frozen to -80 C at a rate matching the BRIC-23 spaceflight samples or processed immediately to harvest RNA and protein.
Mars Atmosphere and Volatile Evolution (MAVEN) Imaging Ultraviolet Spectrograph (IUVS) Processed-level Data Product Bundle
Mars Atmosphere and Volatile Evolution (MAVEN) Imaging Ultraviolet Spectrograph (IUVS) Processed-level Data Product Bundle
MER2 APXS Xray Spectra (Partially Processed) Data Bundle
This bundle contains Xray Spectra (Partially Processed) data from the Alpha Particle X-ray Spectrometer on Mars Exploration Rover 2 (Spirit). Raw APXS data have been converted to ASCII tables and spectra from previous integrations have been removed.
BRIC-23 GeneLab Process Verification Test: Bacillus subtilis transcriptomic, proteomic, and metabolomic data
Microbes interact with humans in complex ways and understanding how they respond to the spaceflight environment is important to the success of future manned spaceflight missions. The BRIC-23 mission was designed to measure the response of Bacillus subtilis and Staphylococcus aureus to the spaceflight environment. This experiment aimed to produce high quality omics data from B. subtilis and S. aureus grown aboard the International Space Station (ISS) to allow comparison to matched ground controls. There were two primary objectives for this experiment: (1) Demonstrate all post-flight processes and operations required for successful completion of GeneLab Reference Missions conducted on ISS, and (2) Generate high quality GeneLab Reference Mission omics data sets for two prokaryotic model organisms, Bacillus subtilis and Staphylococcus aureus. Freezing Control Experiment: The BRIC hardware has significant thermal inertia, thus the freezing rate of samples placed at -80 C is quite slow. This could affect RNA-sequencing, proteomic and metabolic data sets. In an effort to understand how slow freezing could affect these data sets, a control experiment was designed in which B. subtilis and S. aureus were grown in petri plates and either slow frozen to -80 C at a rate matching the BRIC-23 spaceflight samples or processed immediately to harvest RNA and protein. S. aureus omics data is deposited in GLDS-145.
SPATIALLY ADAPTIVE SEMI-SUPERVISED LEARNING WITH GAUSSIAN PROCESSES FOR HYPERSPECTRAL DATA ANALYSIS
SPATIALLY ADAPTIVE SEMI-SUPERVISED LEARNING WITH GAUSSIAN PROCESSES FOR HYPERSPECTRAL DATA ANALYSIS GOO JUN * AND JOYDEEP GHOSH* Abstract. A semi-supervised learning algorithm for the classification of hyperspectral data, Gaussian process expectation maximization (GP-EM), is proposed. Model parameters for each land cover class is first estimated by a supervised algorithm using Gaussian process regressions to find spatially adaptive parameters, and the estimated parameters are then used to initialize a spatially adaptive mixture-of-Gaussians model. The mixture model is updated by expectationmaximization iterations using the unlabeled data, and the spatially adaptive parameters for unlabeled instances are obtained by Gaussian process regressions with soft assignments. Two sets of hyperspectral data taken from the Botswana area by the NASA EO-1 satellite are used for experiments. Empirical evaluations show that the proposed framework performs significantly better than baseline algorithms that do not use spatial information, and the results are also better than any previously reported results by other algorithms on the same data.
Block-GP: Scalable Gaussian Process Regression for Multimodal Data
Regression problems on massive data sets are ubiquitous in many application domains including the Internet, earth and space sciences, and finances. In many cases, regression algorithms such as linear regression or neural networks attempt to fit the target variable as a function of the input variables without regard to the underlying joint distribution of the variables. As a result, these global models are not sensitive to variations in the local structure of the input space. Several algorithms, including the mixture of experts model, classification and regression trees (CART), and others have been developed, motivated by the fact that a variability in the local distribution of inputs may be reflective of a significant change in the target variable. While these methods can handle the non-stationarity in the relationships to varying degrees, they are often not scalable and, therefore, not used in large scale data mining applications. In this paper we develop Block-GP, a Gaussian Process regression framework for multimodal data, that can be an order of magnitude more scalable than existing state-of-the-art nonlinear regression algorithms. The framework builds local Gaussian Processes on semantically meaningful partitions of the data and provides higher prediction accuracy than a single global model with very high confidence. The method relies on approximating the covariance matrix of the entire input space by smaller covariance matrices that can be modeled independently, and can therefore be parallelized for faster execution. Theoretical analysis and empirical studies on various synthetic and real data sets show high accuracy and scalability of Block-GP compared to existing nonlinear regression techniques.
TuxNet: A simple interface to process RNA sequencing data and infer gene regulatory networks
GEO Series GSE112563. Arabidopsis thaliana. 6 samples. Type: Expression profiling by high throughput sequencing.
Measurement data evaluating stream processing on HPC systems
<p>Measurement data evaluating different setups and options for stream processing on the Cori supercomputer at NERSC.</p> <p>To access the data, download the files and use the following commands:</p> <pre><code>cat stream_hpc.tar.gz_part_* > stream_hpc.tar.gz tar xzf stream_hpc.tar.gz</code></pre> <p> </p>
Data associated with 'Processes at the margins of supraglacial debris cover: quantifying dirty ice ablation and debris redistribution'
<p>Data associated with Fyffe, C. L., Woodget, A. S., Kirkbride, M. P., Deline, P., Westoby, M. J. and Brock, B. W. (2020) Processes at the margins of supraglacial debris cover: quantifying dirty ice ablation and debris redistribution, Earth Surface Processes and Landforms, doi:10.1002/esp.4879.</p> <p>The file All_quadrat_data_Z contains the data associated within each of the quadrats which were positioned on the ice surface. Some of the values are extracted from the UAV derived data within the area of the quadrat, while other data is based on measurements from the ground truth images (taken at ground level). The main data types are given below. The data sub-types and units are given in the file header.</p> <ul> <li>Quadrat: Quadrat name</li> <li>Elevation: Elevation from the UAV DEM</li> <li>Aspect: Aspect from the UAV DEM</li> <li>Ablation: Ablation derived from the UAV products (see the paper for details)</li> <li>Albedo: Measured albedo</li> <li>Rock_stats_image: Statistics on the clasts sampled from the ground truth image of the quadrat</li> <li>Debris_cover: Percentage debris cover derived from both the ground truth image (point) and UAV derived orthophotos</li> <li>Rock_stats_meas: Statistics on the clasts measured in the field</li> </ul> <p>The file Feature_Tracking_Data_Z contains data on clast movement, derived by locating the same clasts in the July and August orthophotos. The August orthophoto location was corrected for ice flow. The main data types are given below. The data sub-types and units are given in the file header.</p> <ul> <li>Coordinates: Locations of the points on the clasts in the July and August orthophotos</li> <li>Elevation: Elevation from the UAV derived DEMs</li> <li>Slope_from_slope_map: Slope from the UAV derived DEMs</li> <li>Surface: Either dirty ice (partial debris cover, or debris (continuous cover)</li> <li>Clast_length: A-axis length of the tracked clast</li> <li>Track_lines_stats: Characteristics of the line joining the clast locations</li> <li>Clast_velocity: Velocity of the tracked clast. Note the velocities derived from the distance measured over the surface were used for the analysis in the paper.</li> </ul> <p>The file Profiles_Data_Z contains the elevation and ablation data extracted following four cross-profiles which crossed the study area. Tables are also included giving the calculated moraine slope gradients and their change over time. The main data types are:</p> <ul> <li>Distance: The distance along the cross-section</li> <li>Elevation: The elevation from either the July or August DEM</li> <li>Ablation: The ablation value from the map of spatially continuous ablation.</li> </ul> <p>The folder GIS_Data_Z contains a range of GIS data associated with the paper. Included in the Zip file is an ArcMap document 'GIS_Data_Dirty_Ice' which includes all the data, along with a Word document ('Metadata for GIS data') which describes each of the files included within the ArcMap document. Briefly, the included data covers ground truth data (georeferenced ground truth photos, sample points and their classification, clast sizes, quadrat georeferencing information), clast track data (over the study area and the small scale clast tracks), stake data (measured and modelled ablation, calculated horizontal velocities and combined slope and emergence values), mapped ablation, percentage debris cover (and the difference in percentage cover), orthophotos, elevation data (and the cross-sections used to extract this) and the boundaries between the dirty and debris-covered ice. All data sets are in the WGS 1984 UTM Zone 32N coordinate system.</p> <p>The folder Percent_Debris_Analysis_H_Z contains the data used to derive Figure 5 (ablation and percentage debris cover data) in Matlab format (.mat file). It also includes the code used to make the figure. More details on file names will not be given here as the code documents the analysis steps.</p> <p>If any further information is required about the data then please get in contact with Catriona Fyffe.</p> <p> </p>
Data for: The influence of microbially induced calcite precipitation on subsurface transport processes in cement at the pore scale
<p>This is the dataset for the paper "The influence of microbially induced calcite precipitation on subsurface transport processes in cement at the pore scale" in AWR. </p>
SX Data Processing Results with CTDD Variations
<p>Data Processing Files</p> <p>Data: Data processing results of lysozyme obtained by serial synchrotron crystallography<br>Data processing software: CrystFEL<br>Indexing algorithms: MOSFLM, DirAx, XGANDALF<br>Note: The numbers in the file names indicate the crystal-to-detector distance.</p>
SubsurfaceBreaks v. 1.0: A supervised detection of fault-related structures on triangulated models of subsurface homoclinal interfaces: Input and Processed Data
<p>This companion dataset relates to the manuscript "<strong>SubsurfaceBreaks</strong> <strong>v. 1.0: A supervised detection of fault-related structures on triangulated models of subsurface homoclinal interfaces"</strong>, by Michał Michalak, Christian Gerhards and Peter Menzel.</p> <p>There are several groups of files:</p> <ul> <li>a file with parameters (params.txt) of the generated homoclinal interfaces (slopes) such as dip angle, dip direction, level of noise).</li> <li>files 0-999 are generated using the code from GitHub. (https://github.com/michalmichalak997/SubsurfaceBreaks/blob/main/Broken_synthetic_subsurface_slopes) for generating synthetic slopes. Every slope is in a separate file (.txt files) and it is possible to upload the slope to ParaView for further inspection: Delaunay triangulation, normal vectors and dip vectors have their own .vtu files. The .txt files (0-999) can be uploaded for training using the Python script (https://github.com/michalmichalak997/SubsurfaceBreaks/blob/main/Broken_subsurface_slopes_training_testing_evaluating_revision.ipynb).</li> <li>KSH_input.txt corresponds to real data from Kraków-Silesian Homocline. Every row corresponds to a point representing a geological horizon separating Middle Jurassic geological units: Kościeliska sandstones from ore-bearing clays. This data set can be used to calculate geometric attributes using the code from GitHub (https://github.com/michalmichalak997/SubsurfaceBreaks/blob/main/Broken_real_subsurface_slopes).</li> <li>KSH_input_output_0 corresponds to an output file from processing the KSH_input.txt file using the code from GitHub (https://github.com/michalmichalak997/SubsurfaceBreaks/blob/main/Broken_real_subsurface_slopes). This file should be uploaded to the Python script to identify fault-related features on a real subsurface slope.</li> </ul>
Pre-processed functional data for the github repo: gecthomas/Spectral_DCM_in_PD_VH/tree/spm_course
<p>This repository was created for the purposes of a practical demonstration for the May 2024 SPM course at UCL.</p> <p>This upload contains subject-level anonymised functional data that have undergone the following pre-processing:</p> <ul> <li>first 5 volumes discarded</li> <li>spatial realingment</li> <li>unwarping</li> <li>normalisation to MNI space</li> <li>smoothing</li> <li>denoising with ICA-AROMA</li> </ul> <p>There are also anatomical data in the form of brain masks.</p> <p>These data are to be used in conjuction with the code in the repositry found <a href="https://github.com/gecthomas/Spectral_DCM_in_PD_VH/tree/spm_course" target="_blank" rel="noopener"><strong>here</strong></a>.</p>
Data and code used in the article "Integrating multiphysics processes with deep learning for the Prediction of coupled water-vapor-heat water fluxes in unsaturated zone"
<p>《Integrating multiphysics processes with deep learning for the Prediction of coupled water-vapor-heat water fluxes in unsaturated zone》 (the paper has been submitted to <em>Water Resources Research</em>). The data and code used in the paper, specifically, the water content data and temperature data from in-situ observations from July 1, 2020 to August 29, 2020 (with a frequency of observations of 5 min), and the sample code for implementing cpinn using python (mainly the tensorflow library). These data can help readers better understand and replicate our study. All data and code have been uploaded.</p> <p>The paper has been submitted to <em>Water Resources Research</em>, if the paper is accepted for publication, all data will be made freely available.</p> <p>.</p> <p> </p>
Supplementary data repository for publication "Machine hammer peening as a post-processing treatment for powderbed manufactured maraging steel X3NiCoMoTi18-9-5 (1.2709)"
<p>Supplemantary data repository for investigation on powderbed 3D printed tesnile specimens in raw and postprocessed conditions.</p> <p>This readme contains a brief description of the chosen materials, parts and testing devices as well as the perfomed testing methods, including tensile testing, hardness testing and microstructural analysis, with references to the respective supplementary data in the repository. An explanation of the data aquisition and processing methods, together with the used nomenclature is given as well.</p> <p>The dataset is compressed as `Dataset.zip` to preserve the folder structure.</p> <p>See README.MD or README.PDF for detailed description of the dataset.</p> <p> </p>
Magnetic Bead Processing Enables Sensitive Ligation-based Detection of HIV Drug Resistance Mutations [Data Release]
<p>Raw data files for the manuscript "Magnetic Bead Processing Enables Sensitive Ligation-based Detection of HIV Drug Resistance Mutations".</p>
TuxNet: A simple interface to process RNA sequencing data and infer gene regulatory networks [XVE:PAN]
GEO Series GSE112564. Arabidopsis thaliana. 9 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.