Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
Patient-specific processed data, code and visualisations for "Fluctuations in EEG band power at subject-specific timescales over minutes to days explain changes in seizure evolutions"
<p>Processed data and code for reproducing the main results and figures of the paper "<strong>Fluctuations in EEG band power at subject-specific timescales over minutes to days explain changes in seizure evolutions</strong>".</p> <p>We analysed publicly available data from subjects with drug-resistant focal epilepsy. A total of 2656 hours of long-term intracranial electroencephalography (iEEG) from 18 subjects was obtained using the "The SWEC-ETHZ iEEG Database and Algorithms" (available at <a href="http://ieeg-swez.ethz.ch">http://ieeg-swez.ethz.ch</a>) (Burrello et al., 2019).</p> <p>Reference<br> A. Burrello, L. Cavigelli, K. Schindler, L. Benini, A. Rahimi, <strong>‘‘</strong>Laelaps: An Energy-Efficient Seizure Detection Algorithm from Long-term Human iEEG Recordings without False Alarms<strong>’’</strong> <em>in proceedings of the</em> <em>ACM/IEEE Design, Automation, and Test in Europe Conference (DATE)</em>, Florence, Italy, March 25-29, 2019. </p>
Process Controls on Flood Seasonality in Brazil - link to data and plots
<p>This data set and plots of circular histograms accompany the paper "Process Controls on Flood Seasonality in Brazil" at Geophysical Research Letters (<a href="https://doi.org/10.1029/2021GL096754">https://doi.org/10.1029/2021GL096754</a>). In this paper, we investigate the relationship between the seasonality of floods, maximum annual rainfall, and maximum annual soil moisture data of 886 basins in Brazil for 1980-2015 to shed light on process controls of flood generation.</p>
CSV files and R script: writing process data of typed picture description by 15 cognitively impaired patients and 15 healthy controls
<p>Writing process data of 15 cognitively impaired patients and 15 age- and gender-matched healthy controls were obtained. Each of them completed two typed picture description tasks that were logged with Inputlog, a keystroke logging tool. Variables included time on task; number of characters, pauses and Pause-bursts per minute; proportion of pause time; duration of Pause-bursts; and pause time between words. For pause time between words, also the effect of pauses preceeding specific word categories was analyzed.</p> <p>The data were used to explore if the observation of writing behavior can assist in the screening and follow-up of mild cognitive impairment (MCI) and mild dementia due to Alzheimer’s disease (AD). This data set contains the CSV files that were used for the analyses and the corresponding R script.</p>
Supplementary GIS data - Potential and implications of automated pre-processing of LiDAR-based digital elevation models for large-scale archaeological landscape analysis
<p>A supplementary dataset related to the paper discussing preparation of a digital elevation model derived from DMR 5G (LiDAR-based DEM of the Czech Republic) cleaned of modern artificial features. It includes data used as a clipping mask and data produced during the testing phase.</p> <p>Contents:</p> <ul> <li>..\clipping_buffers.gdb\ - Clipping buffers based on ZABAGED dataset used for masking the original data stored as ESRI geodatabase.</li> <li>..\drainages\ - Drainages with Strahler order higher than four (potential watercourses) for the original and filtered DEMs. <ul> <li>drainages_filtered - Drainges identified in the filtered DEM stored as GeoTIFF.</li> <li>drainages_original - Drainges identified in the original DEM stored as GeoTIFF. </li> </ul> </li> <li>..\LSC\ - Locations with significant land surface curvature for the original and filtered DEMs. <ul> <li>LSC_filtered - Significant LSC identified in the filtered DEM stored as GeoTIFF. </li> <li>LSC_original - Significant LSC identified in the original DEM stored as GeoTIFF. </li> </ul> </li> <li>..\visibility\ - Viewsheds computed over the original and filtered DEMs. <ul> <li>Libice\ - Sample viewsheds computed for the early medieval hillfort of Libice. <ul> <li>Libice_visibility_filtered - Viewshed based on the filtered DEM stored as GeoTIFF. </li> <li>Libice_visibility_original - Viewshed based on the original DEM stored as GeoTIFF. </li> <li>observer_points - Observer points used for calculating the viewsheds.</li> </ul> </li> <li>regular_grid\ - Cumulative viewsheds calculated for regularly spaced points in a 10 x 10 km grid with a visibility radius of 5 km and an observer height of 2 m; a total of 574 viewsheds. <ul> <li>visibility_filtered - Cumulative viewshed for the filtered DEM stored as GeoTIFF.</li> <li>visibility_original - Cumulative viewshed for the original DEM stored as GeoTIFF. </li> <li>visibility_test_buffers - Buffers used for the viewshed calculations stored as ESRI shapefile.</li> <li>visibility_test_observers - Observer points used for the viewshed calculations stored as ESRI shapefile.</li> </ul> </li> </ul> </li> </ul> <p> </p> <p>Preprint version of the related paper:</p> <p>Novák, David and Pružinec, Filip, Potential and Implications of Automated Pre-Processing of Lidar-Based Digital Elevation Models for Large-Scale Archaeological Landscape Analysis. Available at SSRN: <a href="https://ssrn.com/abstract=4063514">https://ssrn.com/abstract=4063514</a></p>
Data and code from: Geomorphological processes shape plant community traits in the Arctic
<p><strong>Aim</strong></p> <p>Geomorphological processes profoundly affect plant establishment and distributions, but their influence on functional traits is insufficiently understood. Here, we unveil trait-geomorphology relationships in Arctic plant communities.</p> <p> </p> <p><strong>Location</strong></p> <p>High-Arctic Svalbard, low-Arctic Greenland, and sub-Arctic Fennoscandia.</p> <p> </p> <p><strong>Time period</strong></p> <p>2011-2018</p> <p> </p> <p><strong>Major taxa studied</strong></p> <p>Vascular plants</p> <p> </p> <p><strong>Methods</strong></p> <p>We collected field-quantified data on vegetation, geomorphological processes, microclimate, and soil properties from 5280 plots and 200 species across the three Arctic regions. We combined these data with database trait records to relate local plant community trait composition to dominant geomorphological processes of the Arctic, namely cryoturbation, deflation, fluvial processes, and solifluction. We investigated the relationship between plant functional traits and geomorphological processes using hierarchical generalised additive modelling.</p> <p> </p> <p><strong>Results</strong></p> <p>Our results demonstrate that community-level traits are related to geomorphological processes, with cryoturbation most strongly influencing both structural and leaf economic traits. These results were consistent across regions, suggesting a coherent biome-level trait response to geomorphological processes.</p> <p> </p> <p><strong>Main conclusions</strong></p> <p>The results indicate that geomorphological processes shape plant community traits in the Arctic. We provide empirical evidence for the existence of generalisable relationships between plant functional traits and geomorphological processes. The results indicate that the relationships are consistent across these three distinct tundra regions and that geomorphological processes should be considered in future investigations of functional traits.</p> <p> </p> <p>Kemppinen, Niittynen, Happonen, le Roux, Aalto, Hjort, Maliniemi, Karjalainen, Rautakoski & Luoto. Provisionally accepted. Geomorphological processes shape plant community traits in the Arctic. Global Ecology and Biogeography.</p> <p>These are the data and code from Kemppinen et al. (Provisionally accepted).</p>
Input data for ERIC values calculation for RAA process
<p>For each area (in total 9 areas of SEE) and for each week (in total 2 weeks of 2021) used in final demonstration of CROSSBOW TC1.1.1 one input file is prepared – in total 18 excel files (*.xlsx).</p> <p>Each file is consisted of 4 sheets:</p> <ul> <li>LOAD – load forecast for 168 timestamps of the given week;</li> <li>DISP_GEN_THERMAL – unit capacity and Forced Outage Rate of production for dispatchable thermal units;</li> <li>DISP_GEN_HYDRO – unit capacity and Forced Outage Rate of production for dispatchable hydro units;</li> <li>NONDISP_GEN – forecasted values for Photovoltaic, Wind, Run of River, Combined Heat and Power and Biomass non-dispatchable units for 168 timestamps of the given week.</li> </ul> <p>Sampling period: Week 9 and Week 10 of 2021</p>
Compound Data for Robust Processes for Polymer Modification and Pharmaceutical Synthesis
<p>Compound structural (IUPAC name, InChI, InChI Key, SMILES, .mol, .sdf) and spectral (NMR, MS) data included for compounds reported in the associated doctoral thesis. NMR data collected on Bruker Avance 400, 500, or 600 MHz spectrometers. Compound structure data were generated by ChemDraw v.20 (PerkinElmer). More details about the preparation and characterization of these compounds can be found in the associated thesis.</p>
In-network data collection and data processing - Supplementary materials for deliverable D5.1 - EU-H2020 FET project 'Watchplant'
<p>Supplementary material for D5.1 - Watchplant. Contains collected dataset from plant experiments with blue and red light stimuli, classification results based on statistical methods, and plots of the recorded electropotentials.</p>
Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models
<p>This data set contains the data, JMP scripts, and figures of the article titled "Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 3. Validation of prediction models" to be published in the journal Animal - Open Space.</p>
Data from: Experimental Investigation of Efficiency and Deposit Process Temperature during Multi-Layer Friction Surfacing
<p>This dataset contains the data for the publication "Experimental Investigation of Efficiency and Deposit Process Temperature during Multi-Layer Friction Surfacing"</p>
Data and scripts for 'Sub-seasonal variability of supraglacial ice cliff melt rates and associated processes from time-lapse photogrammetry'
<p>This repository contains three zipped elements used in the study <em>Sub-seasonal variability of supraglacial ice cliff melt rates and associated processes from time-lapse photogrammetry (</em>https://doi.org/10.5194/tc-2022-81):</p> <p>1. The time-lapse DEMs (original and flow-corrected), orthomosaics (not flow-corrected) and cliff outlines (original and flow-corrected) of the 24K and Langtang survey areas used in this study. The spatial resolution is the same as used in the analysis. These zipped files also contain a .csv file (time_selection.csv) indicating for each index (indicated in the file name) the serial date number in days (date origin January 0, 0000).</p> <p>2. The R and Python scripts (Scripts_final.zip) used to process the DEMs and orthomosaics from the time-lapse images as well as the script to calculate the slope-perpendicular melt. These scripts come with .csv and .txt files that serve as template for the required input data format.</p>
Ancillary data for the wrf_to_tell processing chain
<p>This data supports the sequence of processing scripts that convert the meteorology from IM3's climate simulations using the Weather Research and Forecasting (WRF) model into input files ready for use in the Total ELectricity Load (TELL) model. More details about the processing chain and the associated scripts can be found in the IM3 components library: https://github.com/IMMM-SFA/im3components/tree/main/im3components/wrf_to_tell.</p>
Raw and processed cell lines (melanomaC818/melanomaMUM-2B/SK-MEL-28) data for detecting DMKN's mutations in melanoma cancer
<p>Raw and processed cell lines (melanomaC818/melanomaMUM-2B/SK-MEL-28) data for detecting DMKN's mutations in melanoma cancer. This research was concluded that DMKN is a trigger of epithelial-mesenchymal transition-driven melanoma.</p>
Pre-processed ex vivo MRI data for manuscript titled "Neuroanatomical and cognitive biomarkers of alpha-synuclein propagation in a mouse model of synucleinopathy prior to onset of motor symptoms""
<p>Repository for <em>ex vivo</em> magnetic resonance imaging data from the project titled "Presymptomatic neuroanatomical and cognitive biomarkers of alpha-synuclein propagation in a mouse model of synucleinopathy"</p> <p>Contains the pre-processed <em>ex vivo</em> T1-weighted images (Bruker 7T; 70 micron isotropic voxel resolution) for M83 alpha-synuclein A53T hemizygous mice that received either a phosphate buffered saline (PBS) or alpha-synuclein pre-formed fibrils (PFF) injection in the right dorsal striatum. Full subject list can be viewed with the "subject_list.csv" file. More details are available in the manuscript. </p>
Processed data for FOXA2 analysis in TCGA KIRP and KIRC patients
<pre># Data and code to test whether FOXA2 is changed in KIRP patients with low FH ## Code https://github.com/ArianeMora/foxa2_kirp_kirc ## Datasets RNA count data were downloaded from TCGA (https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga) using scidat (https://github.com/ArianeMora/scidat) for patients with kidney cancers (rna_df.csv). ## Processing The kidney cancer patient count data were split into KIRC and KIRP, and only the tumour data was used for this analysis, see notebook FOXA2.ipynb in the code folder. Samples were split by their expression of FH in their tumour samples, with several annotations used to separate patients for completeness: 1. Low-High: Comparing the bottom 25% (< Q1) of patients by FH vs “high” FH (i.e. top 25%, > Q3): p.adj 0.00004 2. Low-Normal: Comparing the bottom 25% of patients by FH to the patients with “normal” range FH (between Q1 and Q3): p.adj 0.053 3. Outlier-High: Comparing outlier FH to “high” (i.e. top 25%): p.adj 0.067 4. Outlier-Normal: Comparing the outlier FH (Q1 – 1.5*IQR) to all “normal” FH patients: p.adj 0.169 We did the same for KIRC patients – we don’t see FOXA2 as expected 1. Comparing the bottom 25% of patients by FH vs the top 25% of patients with FH: 0.14 2. Comparing the bottom 25% of patients by FH to the patients with “normal” range FH: 0.25 3. Comparing the outlier FH to all “normal” FH patients: 0.31 4. Comparing outlier FH to “high” (i.e. top 25%): 0.32 Each of these groups were used to also perform DE analysis between the two groups, see respective RMD files in code for details. ### References If you use this work please cite TCGA: ``` Creighton, C. J., Morgan, M., Gunaratne, P. H., Wheeler, D. A., Gibbs, R. A., Gordon Robertson, A., Chu, A., Beroukhim, R., Cibulskis, K., Signoretti, S., Vandin Hsin-Ta Wu, F., Raphael, B. J., Verhaak, R. G. W., Tamboli, P., Torres-Garcia, W., Akbani, R., Weinstein, J. N., Reuter, V., Hsieh, J. J., … University of North Carolina at Chapel Hill. (2013). Comprehensive molecular characterization of clear cell renal cell carcinoma. Nature, 499(7456), Article 7456. https://doi.org/10.1038/nature12222 ``` </pre>
Data set for "Cortical sensory processing across motivational states during goal-directed behavior"
<p>Data set for: Matteucci G, Guyoton M, Mayrhofer JM, Auffret M, Foustoukos G, Petersen CCH, El-Boustani S, Cortical sensory processing across motivational states during goal-directed behavior (2022).</p> <p>Neuron https://doi.org/10.1016/j.neuron.2022.09.032</p> <p>There are 2 files in this upload:</p> <p>1. The file named "Matteucci2022.pdf" is the Open Access pdf file of the manuscript published in Neuron.</p> <p>2. The file named "Matteucci_data_code.zip" (~26.5 GB) is a zipped version of a folder "Matteucci_data_code" (~33 GB), which contains the data analysed in the study along with Matlab code used to generate all main figures of the paper. The analysis code is in a subfolder named "code". This subfolder in turn has three subfolders "analysis_scripts", “analysis_functions” (containing the original code for intermediate data processing) and “paper_figures_scripts” (containing the code for generating each figure panel from pre-processed data). The main script “reproduce_figures.m” will call the subscripts contained in the “paper_figures_scripts” folder to reproduce the plots contained in all main figures of the paper (and take care of adding the relevant code and data folders and subfolders to Matlab file path). The raw and pre-processed data analysed in the study can be found in the folder named "data". A “README.txt” file provides further details on the content of each subfolder.</p>
Meta-analysis of diurnal transcriptomics reveals strong patterns of concordance and discordance in mouse liver: processed data
<p>The accumulation of public transcriptomic timeseries data enables robust meta-analyses that were not possible until recently. To assess the consistency of biological rhythms across studies, 43 public mouse liver tissue timeseries totaling 805 RNA-seq samples were obtained and analyzed. Only the control groups of each study were included, to create comparable data. Technical factors in RNA-seq library preparation were the largest contributors to transcriptome-level differences, beyond biological or experiment-specific factors such as lighting conditions. Core clock genes were remarkably consistent in phase across all studies, while phase distributions of other periodic genes were generally less consistent. Overlap of genes identified as rhythmic across studies was generally low, with around 50% between some of the highest sample count studies. Distributions of phases of significant genes were remarkably inconsistent across studies, but genes consistently identified as rhythmic clustered near ZT0 and ZT12 in acrophase. Data was integrated across studies in a JIVE analysis, which showed that the top two components of joint within-study variation are determined by time of day. A shape-invariant model with random effects was fit to the genes to identify the underlying shape of the rhythms, consistent across all studies. This revealed the extent of asymmetric and multimodal genes.<br> <br> This supplemental file provides preprocessed RNA-seq quantifications of all reviewed datasets, as well as results of multiple analyses.</p>
Data and code for article "Nature reserve customized method of photo and video camera traps materials processing using two-stage neural network approach"
<p><strong>DESCRIPTION</strong> 📓</p> <p>"data" folder directory contains the datasets for classification and detection. </p> <ol> <li>The detection dataset has <strong>YOLOv5 format</strong> and contains three classes <strong>[tigers, leopards, empty]</strong>. The class empty is about <strong>10%</strong> of the total data. The leopard and tiger classes contain <strong>3500</strong> images each. The entire amount of data for the detection task is <strong>7600</strong> images.</li> <li>The classification dataset contains two classes <strong>[tigers, leopards]</strong>. Images for classification are cropped images from the detection task using bounding boxes. Each class has <strong>3500</strong> images</li> </ol> <p> </p> <p>The "weights" folder contains pretrained models for classification and detection tasks. </p> <ul> <li>The detector weights were pre-trained on <strong>231k</strong> images from camera traps located throughout Russia.</li> <li>The classifier weights were pre-trained on <strong>416k</strong> images that were cropped with <strong>bounding boxes</strong> from photographs for the detection task. Some of the images for the classification task were taken from the <strong>Internet</strong>. The classifiers were trained for <strong>29 classes</strong>.</li> <li>You can also find folder <strong>tigers_vs_leopards</strong> in both the detection and classification directory, where there are weights that have been trained on a part of the camera trap images available at the link below.</li> </ul> <p><em>Classification weights</em></p> <ol> <li>EfficientNetv2-M</li> <li><strong>ResNeSt-101e</strong> (🚀 RECOMMENDED)</li> <li>ResNet-101d</li> <li>ReXnet-100</li> <li>SeResNet-152d</li> </ol> <p><em>Detection weights</em></p> <ol> <li>YOLOR-W6-1280</li> <li>YOLOX-X-640</li> <li>YOLOv5-X-640</li> <li>YOLOv5-X-1280</li> <li>YOLOv5-M6-1280</li> <li><strong>YOLOv5-L6-1280</strong> (🚀 RECOMMENDED)</li> </ol> <p>Read README.md file for more details</p>
Supporting data sets for "Estimating Carbon Fixation of Plant Organs for Afforestation Monitoring using a Process-based Ecosystem Model and Ecophysiological Parameter Optimization". (the survey of tree breast diameter and tree height in 11-year old Eucommia ulmoides plantation, values of simulation results used in figures and tables.)
<p>Supporting data sets for Miyauchi et al., Ecology and Evolution, 2019 (accepted).</p> <p>The files store: </p> <p>(1) The survey of tree breast diameter and tree height in <em>Eucommia ulmoides</em> plantation<em>.</em> The ring and stem analysis and dry weight of seven harvested sample trees in the plantation.</p> <p>(2) Values of optimization result used fig.7.</p> <p>(3) Values of prediction result used fig.8. and table 4.</p> <p>(4) Values of optimized parameters by optimization methods, parameter range and constrain.</p>
Processed and annotated yeast gene expression data from yeast2 and ygs98 platforms
<p>This dataset contains the following files:</p> <ul> <li><em>yeast2_processed_rds.tar.gz -</em> processed gene expression matrices from the yeast2 platform. The data is stored in binary R format (.rds).</li> <li><em>ygs98_processed_rds.tar.gz </em>- processed gene expression matrices from the yeast2 platform. The data is stored in binary R format (.rds).</li> <li><em>yeast2-curated-annotations.txt</em> - metadata for the yeast2 platform.</li> <li><em>ygs98-curated-annotations.txt</em> - metadata for the ygs98 platform.</li> </ul> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.