Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

79

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

79 results for “high-dimensional”

Learn how ShareScore rates datasets ↗
zenodo44/100

Data and Code for Publication "Estimating inter-individual Mahalanobis distances from mixed incomplete high-dimensional data: Application to human skeletal remains from 3rd to 1st millennia BC Southwest Germany"

<p>Data and code for publication: H. Rathmann, S. Lismann, M. Francken, A. Spatzier, Estimating inter-individual Mahalanobis distances from mixed incomplete high-dimensional data: Application to human skeletal remains from 3<sup>rd</sup> to 1<sup>st</sup> millennia BC Southwest Germany.&nbsp;<em>Journal of Archaeological Science</em> 156: 105802. <a href="https://doi.org/10.1016/j.jas.2023.105802">https://doi.org/10.1016/j.jas.2023.105802</a></p> <p>The repository contains:</p> <ul> <li>&ldquo;R code for FLEXDIST.txt&rdquo;: R code for executing FLEXDIST, a tool to estimate inter-individual Mahalanobis-type distances, taking correlations among variables into account, applicable to multiple variable scales (nominal, ordinal, continuous, or any mixture thereof), accommodating missing values, and handling high-dimensional data. <strong>Please refer to the latest version of this repository for the most up-to-date R code</strong>.</li> <li>&ldquo;data.csv&rdquo;: Pre-processed dataset comprising 85 dental morphological features collected from 64 archaeological human remains from Final Neolithic to Early Iron Age Southwest Germany used for analysis.</li> <li>&ldquo;complete dataset.xlsx&rdquo;: Complete dataset comprising 199 dental morphological features collected from 144 archaeological human remains from Final Neolithic to Early Iron Age Southwest Germany.</li> </ul>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Datasets used in the thesis "Contributions to High-Dimensional Pattern Recognition"

<p>Datasets used in the thesis &quot;Contributions to High-Dimensional Pattern Recognition&quot; and related publications.</p>

opencc-by-4.0May 2011View details →
zenodo40/100

Materials Science Optimization Benchmark Dataset for High-dimensional, Multi-objective, Multi-fidelity Optimization of CrabNet Hyperparameters

Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a "Turing test" of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 173219 quasi-random hyperparameter combinations were generated across 23 hyperparameters and used to train CrabNet on the Matbench experimental band gap dataset. The results were logged to a free-tier shared MongoDB Atlas dataset. This study resulted in a regression dataset mapping hyperparameter combinations (including repeats) to MAE, RMSE, computational runtime, and model size for CrabNet model trained on the Matbench experimental band gap benchmark task1. This dataset is used to create a surrogate model as close as possible to running the actual simulations by incorporating heteroskedastic noise. Failure cases for bad hyperparameter combinations were excluded via careful construction of the hyperparameter search space, and so were not considered as was done in prior work. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This contrasts with a more traditional approach that imposes a-priori assumptions such as Gaussian noise, e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.

opencc-zeroMar 2023View details →
zenodo40/100

Data of "Bayesian inference of high-dimensional finite-strain visco-elastic-visco-plastic model parameters for additive manufactured polymers and neural network based material parameters generator"

<p><strong>General</strong></p> <p>Data of <a href="http://doi.org/10.1016/j.ijsolstr.2023.112470">https://doi.org/10.1016/j.ijsolstr.2023.112470</a> related to MOAMMM project.</p> <p>Data related to the publication (we would be grateful if you could cite the paper in the case in which you are using the data):</p> <p>title = &quot;Bayesian inference of high-dimensional finite-strain visco-elastic-visco-plastic model parameters for additive manufactured polymers and neural network based material parameters generator.&quot;,<br> journal = &quot;International Journal of Solids and Structures&quot;,<br> year = &quot;2023&quot;,<br> volume = &quot;283&quot;,<br> pages = &quot;112470&quot;,<br> doi = &quot;10.1016/j.ijsolstr.2023.112470&quot;,<br> author = &quot;Ling Wu, Cyrielle Anglade, Lucia Cobian, Miguel Monclus, Javier Segurado, Fatma Karayagiz, Ubiratan Freitas, and Ludovic Noels&quot;</p> <p>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No 862015. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them.</p> <p><strong>Description</strong></p> <p>BI code and results of the inference of a pressure-dependent visco-elastic visco-plastic model developed in [NGU16] with a umat implementation in <a href="https://gitlab.uliege.be/moammm/moammmPublic/code/-/tree/main/MaterialModels/FiniteStrain/Finite_VEVP">https://gitlab.uliege.be/moammm/moammmPublic/code/-/tree/main/MaterialModels/FiniteStrain/Finite_VEVP</a>. The BI is described in [WU23] .The experimental results used in the BI are reported in [COB22,COB22b]. To run the BI you need the open source code <a href="https://gitlab.onelab.info/cm3/cm3Libraries">https://gitlab.onelab.info/cm3/cm3Libraries</a> If you use these data or model, we would be grateful if you could cite the related papers.</p> <p><strong>Bibliography</strong></p> <ul> <li>[WU23] L. Wu, C. Anglade, L. Cobian, M. Monclus, J. Segurado, F. Karayagiz, U. Santos Freitas, L. Noels, Bayesian inference of high-dimensional finite-strain visco-elastic-visco-plastic model parameters for additive manufactured polymers and neural network based material parameters generator, International Journal of Solids and Structures (2023) 112470: https://doi.org/10.1016/j.ijsolstr.2023.112470</li> <li>[COB22] L. Cobian, M. Rueda-Ruiz, J.P. Fernandez-Blazquez, V. Martinez, F. Galvez, F. Karayagiz, T. L&uuml;ck, J. Segurado, M.A. Monclus, Micromechanical characterization of the material response in a PA12-SLS fabricated lattice structure and its correlation with bulk behaviour, Polymer Testing 110 (2022) 107556: https://doi.org/10.1016/j.polymertesting.2022.107556 (in Open access)</li> <li>[COB22b] Data of &ldquo;. Cobian, M. Rueda-Ruiz, J.P. Fernandez-Blazquez, V. Martinez, F. Galvez, F. Karayagiz, T. L&uuml;ck, J. Segurado, M.A. Monclus, Micromechanical characterization of the material response in a PA12-SLS fabricated lattice structure and its correlation with bulk behaviour, Polymer Testing 110 (2022) 107556&rdquo; http://dx.doi.org/10.5281/zenodo.6136935 (in Open access)</li> <li>[NGU16] V. D. Nguyen, F. Lani, T. Pardoen, X. Morelle, L. Noels, A large strain hyperelastic viscoelastic-viscoplastic-damage constitutive model based on a multi-mechanism non-local damage continuum for amorphous glassy polymers. International Journal of Solids and Structures 96 (2016): 192-216; https://dx.doi.org/10.1016/j.ijsolstr.2016.06.008, Open access: https://orbi.uliege.be/handle/2268/197898</li> </ul> <p><strong>Directories</strong></p> <p>All the codes and experimental results are in five directories:</p> <ol> <li>experimentalTests: experimental data, see the README.txt in each subdirectory for details</li> <li>BayesianVE: BI of the visco-elastic parameters <ol> <li>PlotExperimentalCurves: to vizualize the experimental curves and prepare the observations for the BI in the VE range <ol> <li>Loadcase_H.py and Loadcase_V.py read experimental results and create Load_ExpVE_H.dat and Load_ExpVE_V.dat, which keep the experimental observations and loading conditions to perform the BI.</li> <li>PrintDir_H &amp; PrintDir_V subdirectories with the functions called by Loadcase_H.py and Loadcase_V.py</li> <li>Load_ExpVE_H.dat and Load_ExpVE_V.dat created files with the observations and loading conditions to perform the BI</li> </ol> </li> <li>VE_V2Step and VE_H: BI for viscoelastic properties of &quot;V&quot; specimen (VE_V2Step) and &quot;H&quot; specimen (VE_H) <ol> <li>BI_allpos_sequence.py runs the BI using Predict_VETest.py and creates the MCMC_VE_....dat</li> <li>WarmStart = True is used to restart an inference</li> <li>MCMC_VE_....dat in the VE_V2Step and VE_H directories are the BI results</li> <li>When proceeding in two steps in VE_V2Step, a first step generates MCMC_VE_VN8_1st.dat whose posterior is used as prior in the second step to generate MCMC_VE_VN8_2nd.dat</li> </ol> </li> <li>CheckBayRes: to visualize predictions of a BI sample and experimental curves <ol> <li>MCMCRes.py is used to check the numerical predictions of a BI parameter sample (read last sample by default, V or H direction can be selected at line</li> <li>ResKGEmu.py plots the evolution of elastic properties with time</li> <li>uses as input VE_V2Step/MCMC_VE_....dat or VE_H/MCMC_VE_....dat</li> <li>uses local ViscoElasticTest.py, line.geo, line. msh as interface with https://gitlab.onelab.info/cm3/cm3Libraries code</li> <li>uses local functions plotExpLoad_Unload.py, plotExp.py</li> </ol> </li> <li>ViscoElasticTest.py, line.geo, line.msh: interface with <a href="https://gitlab.onelab.info/cm3/cm3Libraries">https://gitlab.onelab.info/cm3/cm3Libraries</a> code used by VE_V2Step and VE_H to call the VEVP model</li> </ol> </li> <li>BayesianVEVP: BI of the visco-elastic and visco-plastic parameters <ol> <li>PlotExperimentalCurves: to vizualize the experimental curves and prepare the observations for the BI in the VE-VP ranges <ol> <li>Loadcase_H.py and Loadcase_V.py read experimental results and create Load_ExpVEVP_H.dat and Load_ExpVEVP_V.dat, which keep the experimental observations and loading conditions to perform BI at the viscoplastic stage.</li> <li>PrintDir_H &amp; PrintDir_V subdirectories with the functions called by Loadcase_H.py and Loadcase_V.py</li> <li>Load_ExpVEVP_H.dat and Load_ExpVEVP_V.dat created files with the observations and loading conditions to perform the BI</li> </ol> </li> <li>VP_V2step and VP_H2step: BI for viscoelastic-viscoplastic properties of &quot;V&quot; specimen (VP_V2Step) and &quot;H&quot; specimen (VP_H2Step) <ol> <li>BI_allpos_sequence.py runs the BI using Predict_VETest.py and creates the MCMC_VP_....dat</li> <li>WarmStart = True is used to restart an inference</li> <li>It starts from the VE prosterior as prior, see point 2, and generates a MCMC_VP_?_1of2Steps.dat (? being H or V)</li> <li>Then using MCMC_VP_?_1of2Steps.dat posterior to get a new prior, it generates MCMC_VP_?_2of2Steps.dat (? being H or V)</li> </ol> </li> <li>CheckBayRes: visualize predictions of a BI sample and experimental curves <ol> <li>MCMCRes.py is used to check the numerical predictions with 3 BI parameter samples ([28000, 45000,70000] by default, V or H direction can be selected at line 12) using the samples of BayesianVEVP/VP_?2step/MCMC_VP_?_2of2Steps.dat (? being H or V)</li> <li>plot_hist.py is used to plot histograms of all the inferred parameters using the samples of BayesianVEVP/VP_?2step/MCMC_VP_?_2of2Steps.dat (? being H or V)</li> <li>Plot_Prop.py plots joints histograms of the inferred parameters using the samples of BayesianVEVP/VP_?2step/MCMC_VP_?_2of2Steps.dat (? being H or V)</li> <li>ResKGEmu.py plots the evolution of elastic properties with time</li> </ol> </li> <li>VEVPTest.py: interface with https://gitlab.onelab.info/cm3/cm3Libraries code used by VP_V2Step and VP_H2Step to call the VEVP model</li> </ol> </li> <li>RandomParametersGenerator: used to generate the parameters from the BI samples, with the same statistical content <ol> <li>Generator <ol> <li>DataProcess.py: creates normalized data for training from final inferred parameters in ../MCMC_ResData and creates ?_dirNormData (? being H or V)</li> <li>KmeanDataProcess.py: performs clustering for the data of H_dirNormDat and creates H_dirNormData_2cluster (no need for V direction because not bimodal)</li> <li>Gan_V.py and Gan_H.py are used to train the random material parameter generators and create the VDir_Gan or HDir_Gan200_0/HDir_Gan200_1</li> <li>GenerateParameters.py generates random parameters using the Gan files VDir_Gan or HDir_Gan200_0/HDir_Gan200_1 and checks the joint histograms of generated parameters, generated parameters are in V_GenData and H_GenData</li> <li>Ganlib.py is used by the generator</li> </ol> </li> <li>CheckRes <ol> <li>GenDataRes.py is used to check the numerical predictions with the generated parameter samples, see point 4) (using V_GenData and H_GenData).</li> <li>Plot_PropGen.py plots joints histograms of the generated parameters using the samples of V_GenData or H_GenData</li> </ol> </li> </ol> </li> <li>MCMC_ResData:All final data used in the paper (they can substitute the ones used here above) <ol> <li>H_direction and V_direction keep the MCMC random walk results of BI.</li> <li>RandomParameterGenerator keeps results of the generator Paper</li> </ol> </li> </ol> <p><strong>Figures of [WU23]</strong></p> <ul> <li>Fig. 5: From directory BayesianVE/PlotExperimentalCurves, run python3 ./PrintDir_V/plotExp_T.py or ./PrintDir_V/plotExp_C.py or ./PrintDir_V/plotExp_R.py</li> <li>Fig. 7: BayesianVEVP/CheckBayRes/Plot_Prop.py with direct = &quot;V&quot; and then with direct = &quot;H&quot; and with Var = [0,1,20,24,28,29,30,31]</li> <li>Fig. 8: BayesianVEVP/CheckBayRes/MCMCRes.py with direct = &quot;V&quot; (requires <a href="https://gitlab.onelab.info/cm3/cm3Libraries">https://gitlab.onelab.info/cm3/cm3Libraries</a> code)</li> <li>Fig. 9: BayesianVEVP/CheckBayRes/MCMCRes.py with direct = &quot;H&quot; (requires<a href="https://gitlab.onelab.info/cm3/cm3Libraries"> https://gitlab.onelab.info/cm3/cm3Libraries</a> code)</li> <li>Fig. 11: RandomParametersGenerator/CheckRes/Plot_PropGen.py with direct = &quot;V&quot; and then with direct = &quot;H&quot; and with Var = [0,1,20,24,28,29,30,31]</li> <li>Fig. 12: RandomParametersGenerator/CheckRes/GenDataRes.py with direct = &quot;V&quot; (requires <a href="https://gitlab.onelab.info/cm3/cm3Libraries">https://gitlab.onelab.info/cm3/cm3Libraries</a> code)</li> <li>Fig. 13: RandomParametersGenerator/CheckRes/GenDataRes.py with direct = &quot;H&quot; (requires <a href="https://gitlab.onelab.info/cm3/cm3Libraries">https://gitlab.onelab.info/cm3/cm3Libraries</a> code)</li> <li>Fig. 14A: From directory BayesianVE/PlotExperimentalCurves, run python3 ./PrintDir_H/plotExp_T.py or ./PrintDir_H/plotExp_C.py or ./PrintDir_H/plotExp_R.py</li> <li>Fig. 15C: BayesianVEVP/CheckBayRes/Plot_Prop.py with direct = &quot;V&quot;, Var = [2,3,8,9,14,15,18,19] and [20,21,22,23,24,25,26,27]</li> <li>Fig. 16C: BayesainVEVP/CheckBayRes/plot_hist.py with direct = &quot;V&quot;</li> <li>Fig. 17C: BayesianVEVP/CheckBayRes/plot_hist.py with direct = &quot;V&quot;</li> <li>Fig. 18C: BayesianVEVP/CheckBayRes/Plot_Prop.py with direct = &quot;H&quot;, Var = [2,3,8,9,14,15,18,19] and [20,21,22,23,24,25,26,27]</li> <li>Fig. 19C: BayesianVEVP/CheckBayRes/plot_hist.py with direct = &quot;H&quot;</li> <li>Fig. 20C: BayesianVEVP/CheckBayRes/plot_hist.py with direct = &quot;H&quot;</li> <li>Fig. 21D: RandomParametersGenerator/CheckRes/Plot_PropGen.py with direct = &quot;V&quot; , Var = [2,3,8,9,14,15,18,19] and [20,21,22,23,24,25,26,27]</li> <li>Fig. 22D: RandomParametersGenerator/CheckRes/Plot_PropGen.py with direct = &quot;H&quot; , Var = [2,3,8,9,14,15,18,19] and [20,21,22,23,24,25,26,27]</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Research data supporting: "TimeSOAP: Tracking high-dimensional fluctuations in complex molecular systems via time variations of SOAP spectra"

<p>This repository contains the set of data shown in the paper&nbsp;<strong>&quot;<em>Time</em>SOAP: Tracking high-dimensional fluctuations in complex molecular systems via time variations of SOAP spectra&quot;</strong>, published on The Journal of Chemical Physics&nbsp;(DOI: 10.1063/5.0147025).</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Projection of high-dimensional genome-wide expression on SOM transcriptome landscapes: supplementary datasets

<p>This is a supplementary dataset with raw data, scripts, and complete analysis results for the paper &quot;Supervised projection of high-dimensional genome-wide expression on SOM transcriptome landscapes&quot;.&nbsp;&nbsp;</p> <p>The archive contains three folders:</p> <p>1. &quot;Simdata&quot; folder contains data, scripts, and results of performance evaluation of extension SOM and supervised SOM with simulated data.&nbsp;</p> <p>2. &quot;IBD&quot; - folder contains data, scripts, and results of analysis of Inflammatory bowel disease datasets.</p> <p>3. &quot;BC&quot; - folder&nbsp;contains data, scripts, and results of analysis of breast cancer&nbsp;datasets.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

High-dimensional Anomaly Detection with Radiative Return in e+e- Collisions

<p>Numpy files of Pythia + Delphes simulated e+e- collisions used as DNN/PFN training inputs.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Imaging Mass Cytometry for high-dimensional tissue profiling in the eye

<p>Imaging mass cytometry data (folders containing single tiffs + cell masks) generated for the analysis of healthy conjunctiva and conjunctival melanoma.</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Raw data of high-dimensional aptamer selection generated by ProSELEX

<p>This is the raw selection dataset generated by ProSELEX pipeline (https://www.nature.com/articles/s41557-023-01207-z). The target of selection is human myeloperoxidase (MPO). See&nbsp;https://github.com/dwangnu/AptaZ for instructions.</p>

opencc-by-4.0Jul 2023View details →
dryad36/100

Data from: High-dimensional imaging of vestibular schwannoma reveals distinctive immunological networks across histomorphic niches in NF2-related schwannomatosis

Open the record for dataset details and reuse information.

publicFeb 2025View details →
dryad32/100

Data from: A method for analysis of phenotypic change for phenotypes described by high-dimensional data

The analysis of phenotypic change is important for several evolutionary biology disciplines, including phenotypic plasticity, evolutionary developmental biology, morphological evolution, physiological evolution, evolutionary ecology and behavioral evolution. It is common for researchers in these disciplines to work with multivariate phenotypic data. When phenotypic variables exceed the number of research subjects—data called 'high-dimensional data'—researchers are confronted with analytical challenges. Parametric tests that require high observation to variable ratios present a paradox for researchers, as eliminating variables potentially reduces effect sizes for comparative analyses, yet test statistics require more observations than variables. This problem is exacerbated with data that describe 'multidimensional' phenotypes, whereby a description of phenotype requires high-dimensional data. For example, landmark-based geometric morphometric data use the Cartesian coordinates of (potentially) many anatomical landmarks to describe organismal shape. Collectively such shape variables describe organism shape, although the analysis of each variable, independently, offers little benefit for addressing biological questions. Here we present a nonparametric method of evaluating effect size that is not constrained by the number of phenotypic variables, and motivate its use with example analyses of phenotypic change using geometric morphometric data. Our examples contrast different characterizations of body shape for a desert fish species, associated with measuring and comparing sexual dimorphism between two populations. We demonstrate that using more phenotypic variables can increase effect sizes, and allow for stronger inferences.

opencc-zeroDec 2013View details →
zenodo32/100

Mass cytometry data for "High-dimensional mass cytometry reveals stemness state heterogeneity in pancreatic ductal adenocarcinoma"

<p>Raw suspension mass cytometry data supporting the publication "High-dimensional mass cytometry reveals stemness state heterogeneity in pancreatic ductal adenocarcinoma".</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Reproducible, high-dimensional imaging in archival human tissue by Multiplexed Ion Beam Imaging by Time-of-Flight (MIBI-TOF)

<p>1. SingleChannelMIBI.zip: Single-channel MIBI-TOF images</p> <p>All folders are labeled as Slide[Number]Stain[Number]_Point[Number]_[TMACoreIndex], where the slide number and stain number correspond to the slide and day of staining, the point number corresponds to the order in which the images were collected for each slide, and the TMA core index corresponds to the ID of the tissue microarray core. Each folder contains single-channel TIFFs for each marker. See paper for details.</p> <p>2. SegmentationOutput.zip: Segmentation output of MIBI-TOF images</p> <p>Cell segmentation was performed using Mesmer (Greenwald NF, Nature Biotechnology 2021,&nbsp;https://www.deepcell.org/predict). Output of Mesmer that delineates the single cells in each of the images is included here. Naming convention is the same as above.</p> <p>3. DataTables.zip: Data tables that are needed to run&nbsp;mpi_ppp_ihc_regression.ipynb</p> <p>Contains MIBI-TOF data (ionpath_processed_data.csv), MIBI-TOF calibration data (calibration_data.csv), IHC data (ihc_data.csv), and a map of each sample to its tissue type (tissue_data.csv). Also includes cell table output from Mesmer with the cell clusters appended to the table (cell_table_size_normalized_clusters.csv).</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Evaluating Feature Attribution Methods in the Image Domain: High-Dimensional Datasets

<p>Here you can find the versions of the Places-365 and Caltech-256 datasets used in the paper Evaluating Feature Attribution Methods in the Image Domain.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Subsystem Discovery in High-Dimensional Time-Series Using Masked Autoencoders

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad32/100

Data from: Mapping beta diversity from space: Sparse Generalized Dissimilarity Modelling (SGDM) for analysing high-dimensional data

1. Spatial patterns of community composition turnover (beta diversity) may be mapped through Generalised Dissimilarity Modelling (GDM). While remote sensing data are adequate to describe these patterns, the often high-dimensional nature of these data poses some analytical challenges, potentially resulting in loss of generality. This may hinder the use of such data for mapping and monitoring beta-diversity patterns. 2. This study presents Sparse Generalised Dissimilarity Modelling (SGDM), a methodological framework designed to improve the use of high-dimensional data to predict community turnover with GDM. SGDM consists of a two-stage approach, by first transforming the environmental data with a sparse canonical correlation analysis (SCCA), aimed at dealing with high-dimensional datasets, and secondly fitting the transformed data with GDM. The SCCA penalisation parameters are chosen according to a grid search procedure in order to optimise the predictive performance of a GDM fit on the resulting components. The proposed method was illustrated on a case study with a clear environmental gradient of shrub encroachment following cropland abandonment, and subsequent turnover in the bird communities. Bird community data, collected on 115 plots located along the described gradient, were used to fit composition dissimilarity as a function of several remote sensing datasets, including a time series of Landsat data as well as simulated EnMAP hyperspectral data. 3. The proposed approach always outperformed GDM models when fit on high-dimensional datasets. Its usage on low-dimensional data was not consistently advantageous. Models using high-dimensional data, on the other hand, always outperformed those using low-dimensional data, such as single date multispectral imagery. 4. This approach improved the direct use of high-dimensional remote sensing data, such as time series or hyperspectral imagery, for community dissimilarity modelling, resulting in better performing models. The good performance of models using high-dimensional datasets further highlights the relevance of dense time series and data coming from new and forthcoming satellite sensors for ecological applications such as mapping species beta diversity.

opencc-zeroDec 2014View details →
zenodo32/100

First-passage probability estimation of high-dimensional nonlinear stochastic dynamic systems by a fractional moments-based mixture distribution approach

<p>First-passage probability estimation of high-dimensional nonlinear stochastic dynamic systems is a significant task to be solved in many science and engineering fields, but remains still an open challenge. The present paper develops a novel approach, termed &lsquo;fractional moments-based mixture distribution&rsquo;, to address such challenge. This approach is implemented by capturing the extreme value distribution (EVD) of the system response with the concepts of fractional<br> moment and mixture distribution. In our context, the fractional moment itself is by definition a high-dimensional integral with a complicated integrand. To efficiently compute the fractional moments, a parallel adaptive sampling scheme that allows for sample size extension is developed using the refined Latinized stratified sampling (RLSS). In this manner, both variance reduction and parallel computing are possible for evaluating the fractional moments. From the knowledge<br> of low-order fractional moments, the EVD of interest is then expected to be reconstructed. Based on introducing an extended inverse Gaussian distribution and a log extended skew-normal distribution, one flexible mixture distribution model is proposed, where its fractional moments are derived in analytic form. By fitting a set of fractional moments, the EVD can be recovered via the proposed mixture model. Accordingly, the first-passage probabilities under different<br> thresholds can be obtained from the recovered EVD straightforwardly. The performance of the proposed method is verified by three examples consisting of two test examples and one engineering problem.</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

High-dimensional multivariate autoregressive model estimation of human electrophysiological data using fMRI priors

<p>Data to reproduce figures in&nbsp;submitted&nbsp;manuscript &quot;High-dimensional multivariate autoregressive model estimation of<br> human electrophysiological data using fMRI priors&quot;</p> <p>https://www.biorxiv.org/content/10.1101/2022.11.18.516669v1</p>

opencc-by-sa-4.0Apr 2023View details →
zenodo32/100

Raw results for medium- and high-dimensional settings

<p>Raw results of the Best Subset Selection, Forward Stepwise, Lasso and Elastic-Net methods for medium- and high-dimensional settings.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

High-dimensional deconstruction of pancreatic cancer identifies tumor microenvironmental and developmental stemness features that predict survival

<p>Annotated single cell count data for &quot;High-dimensional deconstruction of pancreatic cancer identifies tumor microenvironmental and developmental stemness features that predict survival&quot;</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record