Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
Galaxy Training Material for Mass spectrometry: GC-MS data processing (with XCMS, RAMClustR, RIAssigner, and matchms)
<p>This dataset contains the training data for the <strong>Mass spectrometry: GC-MS data processing (with XCMS, RAMClustR, RIAssigner, and matchms)</strong> GTN tutorial. It includes 3 GC-[EI+]-HRMS files from seminal plasma samples, the RECETOX Metabolome HR-[EI+]-MS library collected from mostly endogoenous compounds from MetaSci Human Metabolite Library, reference alkanes, sample metadata table, and preprocessed XCMS object.</p>
Input and output data for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 2)
<p>The dataset contains:</p> <p>i. the meteorological forcing, hydrological boundary condition and chlorophyll-a files used as an input</p> <p>ii. the model output and skill produced</p> <p>for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 2)</p>
Input and output data for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 1)
<p>The dataset contains:</p> <p>i. the meteorological forcing, hydrological boundary condition and chlorophyll-a files used as an input</p> <p>ii. the model output and skill produced</p> <p>for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 1)</p>
Input and output data for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 3)
<p>The dataset contains:</p> <p>i. the meteorological forcing, hydrological boundary condition and chlorophyll-a files used as an input</p> <p>ii. the model output and skill produced</p> <p>for Experiment B2 (Deliverable 3.3 - Data assimilation in process-based models for algae bloom forecasting - Section 3)</p>
Simulation data for Doubly Robust Estimation of Business Process Intervention
<p>Event logs of simulated execution of two variants of the same process.</p> <p>A case matrix that contain the outcome, the intervention, and the confounders.</p>
CMIP5 and CMIP6 post-processed AMOC and MLD data supporting Jesse et al. 2023
<p>These are the datasets supporting the paper "Why is CMIP6 projecting larger ocean dynamic sea level in the North Sea than CMIP5?" from Jesse et al. submitted to ERL.</p> <p>See the paper for more information about the data and GitHub for the code that generated the data: https://github.com/dlebars/CMIP_SeaLevel</p> <p>The Atlantic meridional overturning circulation (AMOC) is computed at two latitudes 26N and 35N from different CMIP variables:</p> <p>cmip5_amoc is computed from the variable "msftmyz".</p> <p>cmip6_amoc is computed from the variable "msftmz" or "msftyz" as indicated in the name of the file.</p> <p>cmip5_amoc_vo and cmip6_amoc_vo are computed from the meridional ocean velocity ("vo").</p> <p>For mixed layer depth the variable "mlotst" is used.</p> <p> </p>
Processed data supporting figures in Yang et al. 2023: Oceanic eddies induce a rapid formation of an internal wave continuum
<p>This data repository supports a manuscript by Luwei Yang, Roy Barkan, Kaushik Srinivasan, James C. McWilliams, Callum J. Shakespeare, and Angus H. Gibson, submitted to <em>Communications Earth & Environment</em>. This repository contains the processed data that support the figures in the manuscript. </p>
Spatial data sets of the paper Heinrich Stadial 1 continental sand dunes and Middle to Late Holocene paleosol sequences in SE Iberia: implications for human occupation and site formation processes
<p>Spatial data sets of the paper Heinrich Stadial 1 continental sand dunes and Middle to Late Holocene paleosol sequences in SE Iberia: implications for human occupation and site formation processes. This data set is composed by 3 shapefiles:</p> <ol> <li>Dune_field: Feature class polygon shapefile geometry representing the individual dunes identified in the Villena dune field.</li> <li>Sampled dunes: Shapefile of point geometry representing the location of the stratigraphic sequences of CC1, CC2 and CC3 sampled for texture, soil chemistry, OSL and radiocarbon dating. </li> <li>Sediment sourcing samples: Shapefile of point geometry representing the location of the reference samples of El Moron, El Arenal de la Virgen and Sierra del Castellar. </li> </ol> <p>The spatial reference system is EPSG 25830.</p>
Processed snRNA-seq data from "Divergent single cell transcriptome and epigenome alterations in ALS and FTD patients with C9orf72 mutation"
<p>Processed snRNA-seq data from "Divergent single cell transcriptome and epigenome alterations in ALS and FTD patients with C9orf72 mutation". All nuclei passed QC and were corrected for background noise using cellBender. Files are in R objects saved in RDS (R Data Serialization) format. This repo contains one Seurat v4 object and one gene-by-cell raw RNA count matrix in sparse matrix format (dgCMatrix).</p>
Processed metabolomic data from the EXPOsOMICS Personal Exposure Monitoring study
<p>Metabolomic data from the 'Variability of the Human Serum Metabolome over 3 Months in the EXPOsOMICS Personal Exposure Monitoring Study' paper <a href="https://doi.org/10.1021/acs.est.3c03233">DOI: 10.1021/acs.est.3c03233</a> . </p> <p>The data was originally collected and generated by the multicenter EXPOsOMICS Personal Exposure Monitoring study. Details on data collection and processing are described in the aforementioned paper. The statistical analysis from that paper is available at <a href="https://github.com/moosterwegel/variability-metabolites-paper">https://github.com/moosterwegel/variability-metabolites-paper</a> and may contain useful information/code to work with this data.</p> <p>`processed_covariate_data.csv`:<br> ```<br> Rows: 298<br> Columns: 7<br> $ subjectid: hashed identifier subject<br> $ sample_code: indicates if it's the first (A) or second (B) blood sample<br> $ centre: indicates in which centre the data was collected<br> $ age_cat: indicates age category at the time of a PEM session<br> $ sq_sex: indicates the sex of the participant (male, female) as filled in during the screening questionaire<br> $ traf: indicates the exposure to traffic (PM2.5 and UFP) as measured during the PEM sessions. <br> $ bmi_cat: indicates BMI category at the time of a PEM session<br> ```</p> <p>`processed_lcms_data data.csv` contains the processed LCMS data:<br> ```<br> Rows: 298<br> Columns: 4297<br> $ subjectid: hashed identifier subject<br> $ sample_code: indicates if it's the first (A) or second (B) blood sample<br> $ centre: indicates in which centre the data was collected<br> $ compounds: measured features (compounds) are prefixed by the letter X. The name contains information on the measured monoisotopicmass_retentiontime.<br> Non-detects (below limit of detection (LOD) are coded as 1 for the compounds.<br> ....<br> ```<br> In the datasets each row indicates a measurement on a day (`sample_code`) and person (`subjectid`). The datasets can be joined on these variables.</p> <p>The other data files (`annotations.xslx`, `ancestors_annotations.xlsx`, `annotations_plus_kegg_pathways.csv`) contain the annotations, ancestors of the annotations (to assign a class to a compound based on ChEBI ontology, see our paper for details), annotations plus KEGG pathways respectively. </p>
Raw data for the article: High-throughput computational solvent screening for lignocellulosic biomass processing
<p>This data set contains the raw data for the article "High-throughput computational solvent screening for lignocellulosic biomass processing" published in <em>Chemical Engineering Journal</em>, DOI: <a href="https://doi.org/10.1016/j.cej.2022.139476">https://doi.org/10.1016/j.cej.2022.139476</a></p>
Data for: Processability of Mg-Gd powder via friction extrusion
<p>This dataset contains measurement data, machine logs as well as microstructure and overview images for the publication "Processability of Mg-Gd powder via friction extrusion".</p>
CRISPR/dCas9-mediated DNA demethylation screen identifies driver epigenetic determinants of colorectal cancer (Processed data)
<p><strong>Background:</strong> Promoter hypermethylation of tumour suppressor genes is frequently observed during the malignant transformation of colorectal cancer (CRC). However, whether this epigenetic mechanism is an actual driver of cancer or is a mere consequence of the carcinogenic process remains to be elucidated.</p> <p><strong>Results: </strong>In this work we performed an integrative multi -omic approach to identify gene candidates with strong correlations between DNA methylation and gene expression in human CRC samples and a set of 8 colon cancer cell lines. As a proof of concept, we combined recent CRISPR-Cas9 epigenome editing tools (dCas9-TET1, dCas9-TET-IM) with a custom arrayed gRNA library to modulate the DNA methylation status of 56 promoters previously linked with strong epigenetic repression in CRC, and we monitored the potential functional consequences of such DNA methylation loss by means of a high-content cell proliferation screen. Overall, the epigenetic modulation of most of these DNA methylated regions had a mild impact in the reactivation of gene expression and in the viability of cancer cells. Interestingly, we found that epigenetic reactivation of RSPO2 in the tumour context was associated with a significant impairment in cell proliferation in p53-/- cancer cell lines and further validation with human samples demonstrated that the epigenetic silencing of RSPO2 is a mid-late event in the adenoma to carcinoma sequence.</p> <p><strong>Conclusions: </strong>These results highlight the potential role of DNA methylation as a driver mechanism of CRC and open up the venue for the identification of novel therapeutic windows based on the epigenetic reactivation of certain tumour suppressor genes.</p>
Root biomass data: Biodiversity II: Effects of Plant Biodiversity on Population and Ecosystem Processes
Biodiversity II (E120) is designed to determine how the number of plant species affects the dynamics of ecological processes at the population, community, and ecosystem levels. By experimentally manipulating the number of species and the kinds of species, the amount of plant growth and the change from year to year, that result can be examined. Plots are large (9m x 9m actively maintained) and well-replicated, allowing responses of plant pathogens, insect herbivores, seed predators, soil parameters, invasive plant species and other variables to also be studied. Plots were seeded in May 1994 to have 1, 2, 4, 8, or 16 species, with roughly 30 replicates of each diversity level. The species composition of each plot was chosen by random draw from a pool of 18 grassland perennials that included four warm-season (C4) grasses, four cool-season (C3) grasses, four legumes, four non-legume forbs, and two woody species. All species occur in monoculture allowing comparison of responses of each species in monoculture to combinations of these same species. The experiment was established in 1994 by the lead investigators David Tilman, Peter Reich, Johannes Knops, and David Wedin. Experiment 120 is similar to Experiment 123, but it uses larger plots to provide a large capacity for long-term subexperiments.
Plant aboveground biomass data: Biodiversity II: Effects of Plant Biodiversity on Population and Ecosystem Processes
Biodiversity II (E120) is designed to determine how the number of plant species affects the dynamics of ecological processes at the population, community, and ecosystem levels. By experimentally manipulating the number of species and the kinds of species, the amount of plant growth and the change from year to year, that result can be examined. Plots are large (9m x 9m actively maintained) and well-replicated, allowing responses of plant pathogens, insect herbivores, seed predators, soil parameters, invasive plant species and other variables to also be studied. Plots were seeded in May 1994 to have 1, 2, 4, 8, or 16 species, with roughly 30 replicates of each diversity level. The species composition of each plot was chosen by random draw from a pool of 18 grassland perennials that included four warm-season (C4) grasses, four cool-season (C3) grasses, four legumes, four non-legume forbs, and two woody species. All species occur in monoculture allowing comparison of responses of each species in monoculture to combinations of these same species. The experiment was established in 1994 by the lead investigators David Tilman, Peter Reich, Johannes Knops, and David Wedin. Experiment 120 is similar to Experiment 123, but it uses larger plots to provide a large capacity for long-term subexperiments.
Chlorophyll and phytoplankton composition climatological data on the Northwest Atlantic Shelf from 1978 to 2014: post-processed model data
This dataset includes 8-day composite of surface chlorophyll and bimonthly phytoplankton size composition climatological results on the Northwest Atlantic Shelf from the Gulf of Maine to the Mid-Atlantic Bight based on the physical-biological coupled model results from 1978 to 2014. Two size classes, small phytoplankton (SP) and large phytoplankton (LP), are provided. For more details please see: Zhengchen Zang, Rubao Ji, Zhixuan Feng, Changsheng Chen, Siqi Li, and Cabell S Davis (2021) Spatially varying phytoplankton seasonality on the Northwest Atlantic Shelf: a model-based assessment of patterns, drivers, and implications. ICES Journal of Marine Science, Volume 78, Issue 5, 1920-1934, https://doi.org/10.1093/icesjms/fsab102.
Processed data of single cell RNAseq of 8 human cell lines
<p>Cell annotation and UMI count matrix for 8 human cell lines:</p> <ul> <li>HCT116</li> <li>IMR90</li> <li>A549</li> <li>Ramos</li> <li>H1437</li> <li>HEK293</li> <li>K562</li> <li>Jurkat</li> </ul> <p>The data set is part of the publication in https://rdcu.be/bYEsu</p> <p> </p>
Raw LFP data from Gothard Lab multisensory processing project
<p>Data set contains LFPs collected from two adult male rhesus macaques (foz and blu). There are 41 total sessions. Data were collected with a 16-channel linear electrode array that was acutely lowered into the amygdala and surrounding tissue.</p>
Fast Pixelated Detectors in Scanning Transmission Electron Microscopy. Part II: Post Acquisition Data Processing, Visualisation, and Structural Characterisation
<p>Scanning transmission electron microscopy data related to paper "Scanning transmission electron microscopy data related to paper "Fast Pixelated Detectors in Scanning Transmission Electron Microscopy. Part II: Post Acquisition Data Processing, Visualisation, and Structural Characterisation", <a href="https://doi.org/10.1017/S1431927620024307">https://doi.org/10.1017/S1431927620024307</a>.</p>
Visualization of Dissolution-Precipitation Processes in Lithium-Sulfur Batteries: Supporting Data
<ul> <li>Contours_1.gif: 0 mA/g - pristine state</li> <li>Contours_2.gif: 30 mA/g</li> <li>Contours_3.gif: 80 mA/g</li> <li>Contours_4.gif: 130 mA/g</li> <li>Contours_5.gif: 180 mA/g</li> <li>Contours_6.gif: 230 mA/g</li> <li>Contours_7.gif: 330 mA/g - no remaining solid sulphur</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.