Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “nonparametric”
Astrophysical constraints on neutron star f -modes with a nonparametric equation of state representation
<p>Data release for Mohanty et al. "<em>Astrophysical constraints on neutron star f-modes with a nonparametric equation of state representation"</em></p> <p>The data release consists of three files: </p> <ol> <li><a href="https://zenodo.org/api/records/13952437/draft/files/EoS_posterior_samples_PSR.h5/content" target="_blank" rel="noopener noreferrer">EoS_posterior_samples_PSR.h5</a> </li> <li><a href="https://zenodo.org/api/records/13952437/draft/files/EoS_posterior_samples_PSR+GW.h5/content" target="_blank" rel="noopener noreferrer">EoS_posterior_samples_PSR+GW.h5</a> </li> <li><a href="https://zenodo.org/api/records/13952437/draft/files/EoS_posterior_samples_PSR+GW+NICER.h5/content" target="_blank" rel="noopener noreferrer">EoS_posterior_samples_PSR+GW+NICER.h5</a> </li> </ol> <p>Each file contains 9,835 samples of EOS draws. The equation of state id's matches those of Legred et. al. 2022</p> <p>The data structure follows Legred, I. (2022) “<em>Impact of the PSR J0740+6620 radius constraint on the properties of high-density matter: Neutron star equation of state posterior samples</em>”. Zenodo. doi: 10.5281/zenodo.6502467.</p> <p>Samples were generated using stanspy, a general relativistic neutron star code written by Sailesh Ranjan Mohanty. </p> <p>Please see the readme (adapted from Legred et. al. 2022 Zenodo. doi: 10.5281/zenodo.6502467) </p>
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: Calibration Data
<p>Calibration data accompanying our work, "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse.</p>
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data A
<p>This is the original data for the manuscript "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part A</p>
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data C
<p>This is the original data for the manuscript "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part C.</p>
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data B
<p>This is the original data for the manuscript "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part B.</p>
Figure. Observed (S obs) and estimated species richness for Chao 2, Jackknife 2, and Bootstrap, calculated for Lumbricidae in East Serbia. Vertical dashed lines represent 50%, 75%, and 100% of the sampling effort, respectively. in A nonparametric approach in quantifying species richness of Lumbricidae in East Serbia, Balkan Peninsula
Figure. Observed (S obs) and estimated species richness for Chao 2, Jackknife 2, and Bootstrap, calculated for Lumbricidae in East Serbia. Vertical dashed lines represent 50%, 75%, and 100% of the sampling effort, respectively.
Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 35 Binding Site Data
<p>This is the original data for the manuscript "Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics" by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 35 binding sites.</p>
Replication package for Nonparametric Analysis of Time-Inconsistent Preferences
<p>Stata datasets, R and Julia files.</p>
PixelPop: Nonparametric analysis of correlations in the binary black hole population with LIGO–Virgo–KAGRA data
<p>Data release accompanying the PixelPop papers, analyzing gravitational wave populations.</p> <p>The first dataset (in gwtc3_result_files) is the posterior samples for the runs presented in analysis of LIGO--Virgo--KAGRA data, following the third gravitational wave catalog, see https://arxiv.org/abs/2406.16844. We include a python notebook (example_plot.ipynb) showing how to create the plots presented in this paper.</p> <p>In v2, we also include samples from the predictive distributions. Due to the large uncertainties, marginalizing over the hyperposterior may be a poor representation of the inferred distribution, and so instead we provide samples from the <em>median</em> predictive distribution. That is, samples from the distribution shown in the central panels of the figures. </p> <p>The second dataset (in o4inj_result_files) is the posterior samples accompanying the runs presented in the technical background paper, see https://arxiv.org/abs/2406.16813. </p>
Data from: Evaluation of parametric and nonparametric machine-learning techniques for prediction of saturated and near-saturated hydraulic conductivity
Parametric and nonparametric supervised machine learning techniques were used to estimate saturated and near saturated hydraulic conductivities (Ks, K10) from easily measurable soil properties including name of pedological horizon (HOR), soil texture (sand, silt & clay), organic matter (OM), bulk density (BD) and water contents (θpF1, θpF2, θpF3 and, θpF4.2) measured at four different matric heads (-10, -100, -1000, and -15848 cm). Using a stepwise linear model (SWLM) and the Lasso regression as parametric methods with 316 data in training and 135 data in testing phase, four pedotransfer functions (PTFs) were obtained in which water contents for both methods play an important role compared to other variables. SWLM showed better performance than Lasso in the testing phase for log(Ks) and log(K10) prediction with RMSE of 0.666 and 0.551 cm d-1 and R2 of 0.26 and 0.65. Nonparametric supervised machine learning methods trained and tested with similar data set significantly improved the accuracy of Ks prediction with R2 of 0.52, 0.36 and 0.53 for Gaussian regression process (GPR), support vector machine (SVM) and Ensemble (ENS) method in the testing stage. These methods also described 74.9, 66.7 and 72.5% of the variation of log(K10). Bootstrapping method validated the strong performance of nonparametric techniques. Feature selection capability of GPR determined that instead of using a model with all predictors, HOR, silt, θpF1 and θpF3 are sufficient for the prediction of log(Ks) or log(K10), HOR, silt, and OM can predict as accurate as the comprehensive model with all variables.
Data from: A novel nonparametric measure of explained variation for survival data with an easy graphical interpretation
Introduction: For survival data the coefficient of determination cannot be used to describe how good a model fits to the data. Therefore, several measures of explained variation for survival data have been proposed in recent years. Methods: We analyse an existing measure of explained variation with regard to minimisation aspects and demonstrate that these are not fulfilled for the measure. Results: In analogy to the least squares method from linear regression analysis we develop a novel measure for categorical covariates which is based only on the Kaplan-Meier estimator. Hence, the novel measure is a completely nonparametric measure with an easy graphical interpretation. For the novel measure different weighting possibilities are available and a statistical test of significance can be performed. Eventually, we apply the novel measure and further measures of explained variation to a dataset comprising persons with a histopathological papillary thyroid carcinoma. Conclusion: We propose a novel measure of explained variation with a comprehensible derivation as well as a graphical interpretation, which may be used in further analyses with survival data.
Supplementary Material for the paper entitled "Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data"
<p>This repo contain supplementary tables from the manuscript entitled: "Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data". Clustering is a common way to identify cell types in single-cell RNA-sequencing (scRNA-seq) data. Unfortunately, current methods (i) require users to make human-in-the-loop decisions, which adds significant runtime to bioinformatic analyses, and (ii) reuse the same data twice when testing for differentially expressed genes, which can lead to an increased number of false discoveries. In this work, we overcome these limitations with NCLUSION: a Bayesian nonparametric method that simultaneously clusters cells and selects marker genes. NCLUSION operates without user-defined heuristics to set model parameters and leverages variational expectation-maximization (EM) for posterior inference which allows it to scale well up to 1 million cells. By analyzing publicly available datasets, we illustrate that NCLUSION matches the state-of-the-art clustering performance of competing approaches, achieves improved computational efficiency, and directly enables identification of biologically relevant gene sets driving cluster definitions.</p>
Data from: A novel nonparametric measure of explained variation for survival data with an easy graphical interpretation
Open the record for dataset details and reuse information.
Data from: Evaluation of parametric and nonparametric machine-learning techniques for prediction of saturated and near-saturated hydraulic conductivity
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.