Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,481
datasets available to search
ShareScore release 0.9.0
Dataset results
1,481 results for “data processing”
HarmonEPS modified routines, post-processing scripts and example data used in Tsiringakis, A., Frogner, I.L., de Rooy, W., Andrae, U., Hally, A., Contreras Osorio, S., van der Veen, S. and Barkmeijer, J., An Update to the Stochastically Perturbed Parametrizations Scheme of HarmonEPS. Monthly Weather Review
<p>This dataset contains:</p> <p>- Modified code routines/configurations files used in the EPS of Harmonie-Arome (HarmonEPS) CY43H2.2 version of the model.</p> <p>- Verification scripts from the HARP verification tool, used to verify model output against SYNOP observations.</p> <p>- Post-processing and plotting scripts in python, used in the manuscript.</p> <p>- A subset of the data produced by this study as example input in the verification and post-processing/plotting scripts.</p> <p>This dataset is used in:</p> <p>~Tsiringakis, A., Frogner, I.L., de Rooy, W., Andrae, U., Hally, A., Contreras Osorio, S., van Der Veen, S. and Barkmeijer, J. An Update to the Stochastically Perturbed Parametrizations Scheme of HarmonEPS. Monthly Weather Review</p>
Data on quantum mechanics and electro-magnetic excitation of particles involved in combustion process
<p><span>Distribution of average free electron energy in the body and the closest vicinity of the high frequency spark, arc and corona gaseous discharge</span></p>
Data for "Numerical simulation study of the evolution of lightning channel decay and reactivation processes" by Zheng et al.
<p>All data of the manuscript "Numerical simulation study of the evolution of lightning channel decay and reactivation processes" submitted to Journal of Geophysical Research: Atmospheres.</p> <p>The data supports the manuscript entitled "Numerical simulation study of the evolution of lightning channel decay and reactivation processes”. Microsoft Notepad can open the *.txt files and the *.DAT files, they contain the channel information of two intracloud flashes (IC1 and IC2) and the channel elctrical parameters at different channel segments. </p> <p>The data can be used freely for scientific purposes with the appropriate citation. </p>
1H NMR spectra of commercial honey from 400 and 700 MHz spectrometers and tables with data after processing and binning
<p>Datasets contain the 1H NMR original raw spectral data (Bruker format) of commercial honey from 400 MHz and 700 MHz NMR spectrometers and Tables (.xlsx) with data after processing and binning.</p> <p> </p> <p> </p>
Dataset: Automatic Data Processing, Inc. (ADP) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Automatic Data Processing, Inc. (ADP) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Data for: "A test of the mechanistic process behind the convergent agonistic character displacement hypothesis"
<p><strong>Abstract</strong></p> <p><span>In this era of rapid global change, understanding the mechanisms that enable or prevent species from co-occurring has assumed new urgency. The convergent agonistic character displacement (CACD) hypothesis posits that signal similarity enables co-occurrence of ecological competitors by promoting aggressive interactions that reduce interspecific territory overlap and hence, exploitative competition. In northwestern Switzerland, ca. 10% of <em>Phylloscopus sibilatrix</em> produce songs containing syllables that are typical of their co-occurring sister species, P. bonelli (“mixed singers”). To examine whether the consequences of P. sibilatrix mixed singing are consistent with CACD, we combined a playback experiment and an analysis of interspecific territory overlap. Although P. bonelli reacted more aggressively to playback of mixed P. sibilatrix song than to playback of typical P. sibilatrix song, interspecific territory overlap was not reduced for mixed singers. Thus, the CACD hypothesis was not supported, which stresses the importance of distinguishing between interspecific aggressive interactions and their presumed spatial consequences. </span></p> <p><span> </span></p>
Processed data and summary tables for 'Context TFs establish cooperative environments and mediate enhancer communication'
<p>The linked datasets contain processed data files and summary tables for 'Context transcription factors establish cooperative environments and mediate enhancer communication'. Raw sequenceing data for STARR-seq expeeriments can be found at GEO (Accession ID: GSE229646). For a detailed description of files please refer to the README.</p>
Data of the publication: Recent Advances in Rare Earth Doped Inorganic Crystalline Materials for Quantum Information Processing
<p>Data corresponding to the figures of the publication "Recent Advances in Rare Earth Doped Inorganic Crystalline Materials for Quantum Information Processing" by N. Kunkel and Ph. Goldner (https://doi.org/10.1088/1361-648X/aa529a). A text file describes data in each compressed folder, please refer to the caption in the publication for more details. </p>
Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models
<p>Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models This repository contains the supplemental material for the <a href="https://pqdtopen.proquest.com/pubnum/10759956.html">thesis "Exploring Complexity Metrics for Artifact-Centric Business Process Models" by Marin, Mike A., Ph.D., University of South Africa (South Africa), 2017.</a></p>
A Genetic Algorithm Approach to Regenerate Image from a Reduce Scaled Image Using Bit Data Count-Figure 3. Data compression and image reconstruction (55:148 Digital Image Processing, 2017)
<p>There are several techniques which are normally divided into two categories lossy and lossless image compressions. In lossy compression, after recovery there are negligible difference present where lossless gives accurate image. Huffman encoding is very well known, which can provide optimal compression and decompression without error (55:148 Digital Image Processing, 2017). The basic idea of Huffman coding is to represent data by number of variable size, where more frequent info being represented by shorter number (55:148 Digital Image Processing, 2017). Currently the Lempel-Ziv (or Lempel-Ziv-Welch, LZW) algorithm for dictionary-based coding has got attention as a better compression algorithm (55:148 Digital Image Processing, 2017).</p>
Processed TCGA pan-cancer data set used in the I-Boost paper (Wong et al. 2019)
<p>This data set contains the clinical and genomics data for 1,420 subjects analyzed in the paper: Wong KY, Fan C, Tanioka M, Parker JS, Nobel AB, Zeng D, Lin DY, Perou CM. I-Boost: an integrative boosting approach for predicting survival time with multiple genomics platforms. <em>Genome Biology.</em> 2019. It contains data on time to death, cancer type, 4 clinical variables, expression of 12,434 genes, somatic mutation of 130 genes, expression of 305 miRNA, expression of 136 proteins or phospho-proteins, copy number of 216 DNA segments, and 497 gene expression modules. Data on time to death, clinical variables, somatic mutation, copy number variation, mRNA expression, and miRNA expression were derived from the pan-cancer data set at Synapse (syn2468297 at <a href="https://www.synapse.org/#!Synapse:syn2468297">https://www.synapse.org/#!Synapse:syn2468297</a>). The protein expression data were obtained from Broad GDAC Firehose (<a href="https://gdac.broadinstitute.org/runs/stddata__2016_01_28/">https://gdac.broadinstitute.org/runs/stddata__2016_01_28/</a>).</p>
ViSAPy-generated test data from Lee JH., et al. Advances in Neural Information Processing Systems 30 (NIPS 2017), pp4002--4012
<p>This dataset corresponds to the simulated test data for spike-sorting algorithms in Figure 3 of:</p> <p>Lee, Jin Hyung and Carlson, David E and Shokri Razaghi, Hooshmand and Yao, Weichi and Goetz, Georges A and Hagen, Espen and Batty, Eleanor and Chichilnisky, E.J. and Einevoll, Gaute T. and Paninski, Liam. YASS: Yet Another Spike Sorter. Advances in Neural Information Processing Systems 30 (NIPS 2017). Editors I. Guyon and U. V. Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett, year 2017, pp4002-4012.<br> publisher: Curran Associates, Inc. URL http://papers.nips.cc/paper/6989-yass-yet-another-spike-sorter.pdf</p>
Processed data for "Dissociation of solid tumour tissues with cold active protease for single-cell RNA-seq minimizes conserved collagenase-associated stress responses"
<p>tar.gz of processed data in the form of compressed R files (rds) of SingleCellExperiment (<a href="https://bioconductor.org/packages/release/bioc/html/SingleCellExperiment.html">https://bioconductor.org/packages/release/bioc/html/SingleCellExperiment.html</a>) objects and a metadata csv for the data in the publication <em>Dissociation of solid tumour tissues with cold active protease for single-cell RNA-seq minimizes conserved collagenase-associated stress responses </em>(O'Flanagan et al. 2019).</p>
Data for the publication "Impact of Isolated Atmospheric Aging processes on the Cloud Condensation Nucleiactivation of Soot Particles"
<p>The repository contains the data for the paper:</p> <p>Friebel, F., Lobo, P., Neubauer, D., Lohmann, U., Drossaart van Dusseldorp, S., Mühlhofer, E., and Mensah, A. A.: Impact of isolated atmospheric aging processes on the cloud condensation nuclei activation of soot particles, Atmos. Chem. Phys., 19, 15545–15567, https://doi.org/10.5194/acp-19-15545-2019, 2019.</p> <p>Note that the scripts to plot this data are to be found in the accompanying package (http://dx.doi.org/10.5281/zenodo.3452036)</p>
Processed data from "Chromatin information content landscapes inform transcription factor and DNA interactions"
<p><strong>Chromatin information content landscapes inform transcription factor and DNA interactions</strong></p> <p>Authors: Ricardo D’Oliveira Albanus, Yasuhiro Kyono, John Hensley, Arushi Varshney, Peter Orchard, Jacob O. Kitzman, Stephen C. J. Parker</p> <p><a href="https://doi.org/10.1101/777532">https://doi.org/10.1101/777532</a></p> <p> </p> <p>This record contains the processed data used in our manuscript. For instructions on how to use or regenerate this data, please refer to <a href="https://github.com/ParkerLab/chromatin_information">https://github.com/ParkerLab/chromatin_information</a>.</p>
Processed RNA expression count data from Groen et al.: The strength and pattern of natural selection on rice gene expression
<p>We assessed transcriptome variation in populations of 216 accessions of rice, <em>Oryza sativa</em>, which represented all major varietal groups including indica and japonica. During the 2016 Philippines dry season the accessions were planted in triplicate (with two accessions planted in triplicate three times as replicated checks) in identical alpha-lattice layouts of 660 plots in two fields: a continuously wet paddy, and a field where plants were exposed to intermittent drought in the vegetative and reproductive stages. We measured transcript levels in leaf blades of 50-day-old plants at 33 days after seedling transplant, and 17 days after withholding water in the dry field, using a liquid automation-based 3’ mRNA-seq quantification approach. Samples were multiplexed in batches of 96 per library. Raw sequencing data are available at the SRA in BioProject accession number PRJNA588478. A key to the raw sequencing data in this BioProject can be found in the metadata of the processed RNA expression count data here.</p>
Data and Code for "A lasting impact of serotonergic psychedelics on visual processing and behavior"
<p>Processed data and analysis code for "<span><span>A lasting impact of serotonergic psychedelics on visual processing and behavior". <span>https://doi.org/10.1101/2024.07.03.601959 </span></span></span></p>
Test dataset for Signature 500 data processing
<p>Subset of dataset from a deployment of Nortek Signature500 on a ocean mooring (M1-1) in the Northern Barents Sea during 2018-2019. These files are intended for testing post-processing software. </p> <p>The files have been converted from the native .ad2cp format to .mat using Nortek <a href="https://www.nortekgroup.com/software" rel="nofollow">SignatureDeployment</a> software.</p> <p>Data in these files were collected during this period:</p> <pre>18 May 2019 12:15 --> 15 Jul 2019 08:00 (57.8 days)</pre> <p> </p>
Processed snRNAseq data from female Aedes aegypti antennal neurons
<p>Single-nucleus RNA sequencing data accompanying Adavi et al. 2024 <em>bioRxiv </em>preprint: https://doi.org/10.1101/2024.08.21.608847</p> <p>For analysis scripts see: https://github.com/mcbridelab/Adavi_2024_snRNAseqAaegAntennae</p> <p>For raw sequencing files see NCBI BioProject: PRJNA1138769</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.