Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
608
datasets available to search
ShareScore release 0.9.0
Dataset results
608 results for “ensembles”
Video Figure: Intelligent Agents and Networked Buttons Improve Free-Improvised Ensemble Music-Making on Touch-Screens
<p>This video figure is an overview of our study comparing two designs for network communications between touch-screen musical instruments played in free-improvised ensemble performances.</p> <p>The video shows an overview of the touch-screen app (PhaseRings) used in the study and each of the interface conditions.</p> <p>The abstract of the paper relating to this figure is as follows:</p> <p>We present the results of two controlled studies of free-improvised ensemble music-making on touch-screens. In our system, updates to an interface of harmonically-selected pitches are broadcast to every touch-screen in response to either a performer pressing a GUI button, or to interventions from an intelligent agent. In our first study, analysis of survey results and performance data indicated significant effects of the button on performer preference, but of the agent on performance length. In the second follow-up study, a mixed-initiative interface, where the presence of the button was interlaced with agent interventions, was developed to leverage both approaches. Comparison of this mixed-initiative interface with the always-on button-plus-agent condition of the first study demonstrated significant preferences for the former. The different approaches were found to shape the creative interactions that take place. Overall, this research offers evidence that an intelligent agent and a networked GUI both improve aspects of improvised ensemble music-making.</p>
Arctic Atmoopheric Rivers and Sea Ice Data based on CESM2 Large Ensemble
Open the record for dataset details and reuse information.
Eighty Member Ensemble of 1.0° POP and CICE restart files on the CESM gx1v7 grid
The restart files are from Who Kim's g210.G_JRA.v14 experiment. This experiment was run with repeating "normal-year" atmospheric forcing from the JRA-55 reanalysis. Since the atmospheric forcing was repeated throughout the length of the centuries-long ocean run, the eighty sets of restart files were collected on January 1st, 00Z from eighty different years of the run. These restart files are intended to serve to initialize an 80-member CESM ensemble for users who intend to run an ensemble data assimilation experiment.
CLM5 Perturbed Parameter Ensembles
<p>These data are the results of parameter sensitivity simulations and multiple perturbed parameter ensembles (PPE) with the Community Land Model, version 5 (CLM5). These simulations form the basis of a publication documenting a machine learning approach to studying parameter uncertainty in CLM5.</p> <p>The parameter sensitivity simulations include thirty-four CLM5 biophysical parameters perturbed one at a time to their minimum and maximum values, as well as a control simulation with default parameter values. The output files are labeled by the parameter name with "_min" or "_max" to denote simulation type, where each file includes 5 years of monthly output. The simulation with default parameter values includes 30 years of monthly output, labeled "CLM_control".</p> <p>The PPE simulations include six of the aforementioned parameters perturbed simultaneously. We use Latin Hypercube (LHC) sampling to generate 100 unique parameter sets for the top six parameters identified by the sensitivity simulations. There are two 100-member PPEs with present-day climate forcing conditions, but each uses a different LHC sampling. The output files are labeled "PPE_x" and "PPE_v2_x" with the ensemble member x from 1 to 100. There is also a simulation with default parameter values labeled "PPE_control". Again each file includes 5 years of monthly output.</p> <p>We test optimized parameter values with CLM, the results of which are labeled "CLM_test_calibration", and this file also includes 5 years of monthly output.</p>
Atomistic ensemble of active SHP2
<p>SHP2 phosphatase plays an important role in regulating several intracellular signaling pathways. Pathogenic mutations of SHP2 cause developmental disorders and are linked to hematological malignancies and cancer. This dataset contains the atomistic ensemble of the constitutively active E76K mutant of SHP2 generated by molecular dynamics simulations in solution and validated by accurate explicit-solvent SAXS curve predictions.</p>
FY-4A/AGRI Infrared Brightness Temperature Estimation of Precipitation based on Multi-model Ensemble Learning
Open the record for dataset details and reuse information.
Extended gtf based on a customized gtf file from Ensembl version 102 mm10 for Gastruloid
<p>This gtf has been generated based on https://doi.org/10.5281/zenodo.7510406 and extends 3' of genes using RefSeq and bulk RNA-seq from GSE106225, GSE113885 and time-course samples from GSE205781. All command lines can be found at <a href="https://github.com/lldelisle/extendMouseGTFUsingGastruloidData">https://github.com/lldelisle/extendMouseGTFUsingGastruloidData</a>.</p>
Input data and some models (all except multi-model ensembles) for JAMES paper "Machine-learned uncertainty quantification is not magic"
<p>The tar file contains two directories: data and models. Within "data," there are 4 subdirectories: "training" (the clean training data -- without perturbations), "training_all_perturbed_for_uq" (the lightly perturbed training data), "validation_all_perturbed_for_uq" (the moderately perturbed validation data), and "testing_all_perturbed_for_uq" (the heavily perturbed validation data). The data in these directories are unnormalized. The subdirectories "training" and "training_all_perturbed_for_uq" each contain a normalization file. These normalization files contain parameters used to normalize the data (from physical units to z-scores) for Experiment 1 and Experiment 2, respectively. To do the normalization, you can use the script normalize_examples.py in the code library (ml4rt) with the argument input_normalization_file_name set to one of these two file paths. The other arguments should be as follows:</p><p>--uniformize=1</p><p>--predictor_norm_type_string="z_score"</p><p>--vector_target_norm_type_string=""</p><p>--scalar_target_norm_type_string=""</p><p> </p><p>Within the directory "models," there are 6 subdirectories: for the BNN-only models trained with clean and lightly perturbed data, for the CRPS-only models trained with clean and lightly perturbed data, and for the BNN/CRPS models trained with clean and lightly perturbed data. To read the models into Python, you can use the method neural_net.read_model in the ml4rt library.</p>
Characterizing the complexity of subduction zone flow with an ensemble of multiscale global convection models
<p>Parameter files and model input .txt files for ASPECT mantle convection simulations.</p>
Landslide susceptibility maps using ensemble machine learning models on basin and regional level in Lombardy, Italy
<p>A selection of landslide susceptibility maps computed through ensemble machine learning models for the basins of Val Tartano, Upper Valtellina and Valchiavenna, and on a regional level for the Lombardy region in Italy.</p> <p>A list of the used base machine learning methods:</p> <ul> <li>Random Forest,</li> <li>AdaBoost,</li> <li>Neural Networks.</li> </ul> <p>A list of the used ensemble models:</p> <ul> <li>Stacking,</li> <li>Blending,</li> <li>Soft Voting.</li> </ul> <p>A full list of the model combinations can be found in the "Case Studies" document.</p> <p>The maps are in WGS 84/ UTM zone 32N (EPSG:32632).</p> <p>The map production process details are discussed in Xu et al. 2024. If you use the dataset, please, cite also the paper:</p> <p><em>Qiongjie Xu, Vasil Yordanov, Lorenzo Amici & Maria Antonia Brovelli (2024) Landslide susceptibility mapping using ensemble machine learning methods: a case</em><br><em>study in Lombardy, Northern Italy, International Journal of Digital Earth, 17:1, 2346263, DOI:10.1080/17538947.2024.2346263</em></p> <p>The maps are produced as part of the "Geoinformatics and Earth Observation for Landslide Monitoring" Italy-Vietnam.</p> <p>The work is partially funded by the Italian Ministry of Foreign Affairs and International Cooperation within the project “Geoinformatics and Earth Observation for Landslide Monitoring” CUP D19C21000480001.</p> <p> </p>
Data and code for manuscript ``Insights on the vulnerability of Antarctic glaciers from the ISMIP6 ice sheet model ensemble and associated uncertainty''
<p>Supporting data and code for manuscript:</p><p>Seroussi, H., Verjans, V., Nowicki, S., Payne, A. J., Goelzer, H., Lipscomb, W. H., Abe-Ouchi, A., Agosta, C., Albrecht, T., Asay-Davis, X., Barthel, A., Calov, R., Cullather, R., Dumas, C., Galton-Fenzi, B. K., Gladstone, R., Golledge, N. R., Gregory, J. M., Greve, R., Hattermann, T., Hoffman, M. J., Humbert, A., Huybrechts, P., Jourdain, N. C., Kleiner, T., Larour, E., Leguy, G. R., Lowry, D. P., Little, C. M., Morlighem, M., Pattyn, F., Pelle, T., Price, S. F., Quiquet, A., Reese, R., Schlegel, N.-J., Shepherd, A., Simon, E., Smith, R. S., Straneo, F., Sun, S., Trusel, L. D., Van Breedam, J., Van Katwyk, P., van de Wal, R. S. W., Winkelmann, R., Zhao, C., Zhang, T., and Zwinger, T.: Insights into the vulnerability of Antarctic glaciers from the ISMIP6 ice sheet model ensemble and associated uncertainty, The Cryosphere, 17, 5197–5217, https://doi.org/10.5194/tc-17-5197-2023, 2023.</p><p> </p><p>It contains the code to prepare the datasets, to create the figures and the data for the analysis, and the scalar values computed for the 198 Antarctic glaciers stored by ice flow models.</p><p>The files Glacier_XX contain the data to emulate the results for individual glaciers.</p><p>The files Antarctica and AntarcticaWithCtrl contain the data to emulate the results for the Antarctic runs without and with the ctrl_proj experiment.</p><p>The files GROUP_ICEFLOW contain the ice flow model data for all the experiments recomputed for the 198 glaciers in Antarctica.</p>
Reconciling ASPP-p53 Binding Mode Discrepancies through an Ensemble Binding Framework that Bridges Crystallography and NMR Data
<p>The trajectory of 'Reconciling ASPP-p53 Binding Mode Discrepancies through an Ensemble Binding Framework that Bridges Crystallography and NMR Data'</p>
A field-validated ensemble species distribution model of Eriogonum pelinophilum, an endangered subshrub in Colorado, USA
<p>Understanding the suitable habitat of endangered species is crucial for agencies such as the Bureau of Land Management to plan management and conservation. However, few species distribution models are directly validated, potentially limiting their application. In preparation for a Species Status Assessment of clay‐loving wild buckwheat (<em>Eriogonum pelinophilum</em>), an endangered subshrub found in southwest Colorado, we ran a series of species distribution models to estimate the species' potential occupied habitat and validated these models in the field. A 1‐meter resolution digital elevation model derived from LiDAR and a high‐resolution geology mapping helped identify biologically relevant characteristics of the species' habitat. We employed a weighted ensemble model based on two Random Forest and one Boosted Regression Tree model, and the discrimination performance of the ensemble model was high (AUC-PR = 0.793). We then conducted a systematic field survey of model habitat suitability predictions, during which we discovered 55 new subpopulations of the species and demonstrated that new species observations were strongly associated with model predictions (p < .0001, Cliff's delta = 0.575). We then further refined our original models by incorporating the additional species occurrences collected in the field survey, a new explanatory variable, and a more diverse set of models. These iterative changes to the model marginally improved performance (AUC‐PR = 0.825). Direct validation of species distribution models is extremely rare, and our field survey provides strong validation of our model results. This helps increase confidence in utilizing predictions in planning. The final model predictions greatly improve the Bureau of Land Management's understanding of the species' habitat and increase our ability to consider potential habitat in planning land use activities such as road development and travel management.</p>
Dataset for "Enhanced Regional Ocean Ensemble Data Assimilation Through Atmospheric Coupling in the SKRIPS Model"
Open the record for dataset details and reuse information.
Programs and data used for Extracting latent variables from forecast ensembles and advancements in similarity metric utilizing optimal transport
<p>This is compiled from the program and output data using in Nishizawa (2024).</p> <p> </p> <p>Nishizawa, 2024: Extracting latent variables from forecast ensembles and advancements in similarity metric utilizing optimal transport. submitted to JGR: Machine Learning and Computation.</p>
Dataset and code for "Evident decrease in future European soil moisture in the Kiel Climate Model grand ensemble"
<p>Here, you only have the processed data. Most calculations have been made with CDO (version 2.0.6 or 1.9.9). The history of commands made can be seen in the file with a simple ncdump -h file. If you need to see more data, please contact the corresponding author.</p>
A companion dataset to the paper Scenarios of future climate zone changes in Europe based on EURO-CORDEX regional model ensemble by Holtanová et al., to be submitted to Regional Environmental Change
<p>The content of the dataset is described in the metadata.txt file. </p>
Data for Deriving WMO cloud classes from ground-based RGB pictures with a residual neural network ensemble
Open the record for dataset details and reuse information.
Future wave climate in the Mediterranean Sea and associated uncertainty from an ensemble of 31 GCM-RCM wave simulations
<p>The data provided is used in a study aimed at assessing future changes in the Mediterranean wave climate. A total of 31 GCM-RCM simulations were used to characterize the wave climate during the historical (1979-2005), mid-century (2034-2060), and end-century (2074-2100) periods. Changes in seasonal significant wave height (Hs) and peak period (Tp) wave parameters are evaluated for both the mean and intense (quantile 0.95) wave climate, along with the shift in wave direction (wave peak dominant direction, θp) for sea states characterized as intense. The robustness of the climate change signal is evaluated following the guidelines outlined in the Sixth Assessment Report (AR6) of the Intergovernmental Panel on Climate Change. Additionally, using a wave hindcast as a reference, the changes in future extreme events are assessed by fitting a GEV to a unique and coherent set of bias-corrected annual maxima from each model.</p> <p>We provide data used to obtain the results of our study, comprising: 1) wave climate statistics for Hs, Tp, and θp for each model and each period studied; 2) two sets of annual maxima distribution for each model, which were bias-corrected assuming that the set of extreme events follows either a Gumbel distribution or a GEV distribution. For more details on the methods to obtain these files describing wave climate statistical as used in the study, please refer to Toomey et al., 2024: "Future wave climate in the Mediterranean Sea and associated uncertainty from an ensemble of GCM-RCM wave simulations."</p>
Effects of nonlinear observation operators for visible and infrared radiances in ensemble data assimilation
<p><br>Description</p> <p>Dataset to accompany the first revision of the manuscript "Effects of nonlinear observation operators for visible and infrared radiances in ensemble data assimilation" for the QJRMS.</p> <p>Contains <br>- experiments (one folder per experiment)<br> each contains <br> - diagnostics (DART format) for each assimilation time<br> - obs_seq.final: assimilation output <br> - obs_seq.final-linear: contains linear posterior as "prior" ("evaluate" at +1s after analysis)<br> - obs_seq.final-evaluate: at analysis time (includes DART clamping), at 1s after analysis time: "nonlinear posterior"<br> - config file DART-WRF for each experiment (python format)</p> <p>- code<br> assim_tools_mod.f90: DART code modification to compute the linear posterior</p> <p>- other data: see v1 of this repository!<br> nature run initial conditions (WRF format)<br> forecast ensemble initial conditions (WRF format)<br> script to read "obs_seq.final"-files into pandas.DataFrame (python format)</p> <p><br>Experiment naming:<br>VIS ... visible reflectance 0.6 micrometer assimilation<br>WV73 ... infrared 7.3 micrometer assimilation<br>obs10 ... observation density: 1 observation per 10x10 km of domain area<br>obs30 ... 1 observation per 30x30 km<br>loc10 ... localization radius: 10 km<br>inf0 ... no prior nor posterior covariance inflation in the assimilation<br>sec0 ... no sampling error correction<br>reject ... not assimilating observations with a small first-guess departure <= 0.03</p> <p><br>Experiment names used in Figures</p> <p>Section 3.1<br>- Figure 2a: exp_v1.23_P2_rr+1_VIS_obsi211_loc20_inf0<br>- Figure 2b: exp_v1.23_P2_rr+1_VIS_obsi360_loc20_inf0</p> <p>Section 3.2.1 <br>- Figure 3: exp_v1.23_P2_rr+1_VIS_obs30_loc14_inf0</p> <p>Section 3.2.2<br>- Figure 4: exp_v1.23_P2_rr+1_VIS_obs10_loc20, exp_v1.23_P2_rr+1_VIS_obs10_loc20_reject</p> <p>Section 3.3<br>- Figures 5a,c: exp_v1.23_P2_rr+1_T2M_obs10_loc20_inf0_sec0<br>- Figures 5b,5d,7a: exp_v1.23_P2_rr+1_VIS_obs10_loc20_inf0_sec0<br>- Figure 6a,6b,7b: exp_v1.23_P2_rr+1_WV73_obs10_loc20_inf0_sec0</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.