Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,049
datasets available to search
ShareScore release 0.9.0
Dataset results
1,049 results for “robustness”
Data from: Interplay of robustness and plasticity of life history traits drives ecotypic differentiation in thermally distinct habitats
Phenotypic plasticity describes the ability of an individual to alter its phenotype in response to the environment and is potentially adaptive when dealing with environmental variation. However, robustness in the face of a changing environment may often be beneficial for traits that are tightly linked to fitness. We hypothesized that robustness of some traits may depend on specific patterns of plasticity within and among other traits. We used a reaction norm approach to study robustness and phenotypic plasticity of three life history traits of the collembolan Orchesella cincta in environments with different thermal regimes. We measured adult mass, age at maturity and growth rate of males and females from heath and forest habitats at two temperatures (12 and 22 °C). We found evidence for ecotype-specific robustness of female adult mass to temperature, with a higher level of robustness in the heath ecotype. This robustness is facilitated by plastic adjustments of growth rate and age at maturity. Furthermore, female fecundity is strongly influenced by female adult mass, explaining the importance of realizing a high mass across temperatures for females. These findings indicate that different predicted outcomes of life history theory can be combined within one species' ontogeny and that models describing life history strategies should not assume that traits like growth rate are maximized under all conditions. On a methodological note, we report a systematic inflation of variation when standard deviations and correlation coefficients are calculated from family means as opposed to individual data within a family structure.
Data from: Experimental evaluation of the robustness of the growth-stress tolerance trade-off within the perennial grass Dactylis glomerata
1. A core tenet of functional ecology is that the vast phenotypic diversity observed in the plant kingdom could be partly generated by a trade- off between the ability of plants to grow quickly and acquire resources in rich environments vs. the ability to conserve resources and avoid mortality under stress. However, experimental demonstrations remain scarce and potentially blurred by phylogenetic constraints in cross-species analyses. Here we experimentally decoupled growth potential and stress survival by applying an off-season stress on contrasting populations of the perennial grass Dactylis glomerata exhibiting a range of seasonal dormancy. 2. Seventeen populations of D. glomerata, originating from a latitudinal gradient from Norway to Morocco, were subjected to three types of dehydration stress: winter frost in Norway, summer drought, and early spring (off-season) drought stress in the south of France. Growth rate, and two leaf traits (leaf width and leaf dry matter content) suspected to be involved in the adaptation to dehydration stress, were monitored under optimal conditions. We quantified plant dehydration survival as the amount of plant recovery after a severe stress. 3. Nordic populations were found to be winter dormant. Winter and summer dormant populations better survived frost and summer drought, respectively. However, no trade-off between growth potential and dehydration survival was detected in non-dormant plants in early spring when dehydration occurred unseasonably for all populations. Furthermore, Mediterranean populations better survived an early spring drought. 4. Our results highlight the importance of assessing plant growth potential as a response to seasonal environmental cues. They suggest that growth potential and stress survival trade-off when plants exhibit seasonal dormancy but can be functionally independent at other seasons. Consequently, the growth-stress survival relationship could be better described as a dynamic linkage rather than a constant and general trade-off. Moreover, leaf trait values, such as thinner and more lignified leaves reflecting drought adaptation, may have contributed to the improved drought-stress survival without resulting in a cost to growth. 5. Further exploration of the growth-stress survival relationship should permit deciphering the suite of plant traits and trait covariations involved in plants' responses to increasing stress.
Data from: Mutation is a sufficient and robust predictor of genetic variation for mitotic spindle traits in Caenorhabditis elegans
Different types of phenotypic traits consistently exhibit different levels of genetic variation in natural populations. There are two potential explanations: either mutation produces genetic variation at different rates, or natural selection removes or promotes genetic variation at different rates. Whether mutation or selection is of greater general importance is a longstanding unresolved question in evolutionary genetics. We report mutational variances (VM) for 19 traits related to the first mitotic cell division in C. elegans, and compare them to the standing genetic variances (VG) for the same suite of traits in a worldwide collection C. elegans. Two robust conclusions emerge. First, the mutational process is highly repeatable: the correlation between VM in two independent sets of mutation accumulation lines is ~0.9. Second, VM for a trait is a good predictor of VG for that trait: the correlation between VM and VG is ~0.9. This result is predicted for a population at mutation-selection balance; it is not predicted if balancing selection plays a primary role in maintaining genetic variation.
Data from: Evaluating population receptive field estimation frameworks in terms of robustness and reproducibility
Within vision research retinotopic mapping and the more general receptive field estimation approach constitute not only an active field of research in itself but also underlie a plethora of interesting applications. This necessitates not only good estimation of population receptive fields (pRFs) but also that these receptive fields are consistent across time rather than dynamically changing. It is therefore of interest to maximize the accuracy with which population receptive fields can be estimated in a functional magnetic resonance imaging (fMRI) setting. This, in turn, requires an adequate estimation framework providing the data for population receptive field mapping. More specifically, adequate decisions with regard to stimulus choice and mode of presentation need to be made. Additionally, it needs to be evaluated whether the stimulation protocol should entail mean luminance periods and whether it is advantageous to average the blood oxygenation level dependent (BOLD) signal across stimulus cycles or not. By systematically studying the effects of these decisions on pRF estimates in an empirical as well as simulation setting we come to the conclusion that a bar stimulus presented at random positions and interspersed with mean luminance periods is generally most favorable. Finally, using this optimal estimation framework we furthermore tested the assumption of temporal consistency of population receptive fields. We show that the estimation of pRFs from two temporally separated sessions leads to highly similar pRF parameters.
Data from: Tree diversity increases robustness of multi-trophic interactions
Multi-trophic interactions maintain critical ecosystem functions. Biodiversity is declining globally, while responses of trophic interactions to biodiversity change are largely unclear. Thus, studying responses of multi-trophic interaction robustness to biodiversity change is crucial for understanding ecosystem functioning and persistence. We investigate plant-Hemiptera- (antagonism) and Hemiptera-ant- (mutualism) interaction networks in response to experimental manipulation of tree diversity. We show increased diversity at both higher trophic levels (Hemiptera and ants) and increased robustness through redundancy of lower level species of multi-trophic interactions when tree diversity increased. Hemiptera and ant diversity increased with tree diversity through non-additive diversity effects. Network analyses identified that tree diversity also increased the number of tree and Hemiptera species utilized by Hemiptera and ant species, and decreased the specialization on lower trophic level species in both mutualistic and antagonist interactions. Our results demonstrate that bottom-up effects of tree diversity ascend through trophic levels regardless of interaction type. Thus, local tree diversity is a key driver of multi-trophic community diversity and interaction robustness in forests.
Data from: The robustness and restoration of a network of ecological networks
Understanding species' interactions and the robustness of interaction networks to species loss is essential to understand the effects of species' declines and extinctions. In most studies, different types of networks (such as food webs, parasitoid webs, seed dispersal networks, and pollination networks) have been studied separately. We sampled such multiple networks simultaneously in an agroecosystem. We show that the networks varied in their robustness; networks including pollinators appeared to be particularly fragile. We show that, overall, networks did not strongly covary in their robustness, which suggests that ecological restoration (for example, through agri-environment schemes) benefitting one functional group will not inevitably benefit others. Some individual plant species were disproportionately well linked to many other species. This type of information can be used in restoration management, because it identifies the plant taxa that can potentially lead to disproportionate gains in biodiversity.
Data from: Applying the multistate capture-recapture robust design to characterize metapopulation structure
1. Population structure must be considered when developing mark-recapture (MR) study designs as the sampling of individuals from multiple populations (or subpopulations) may increase heterogeneity in individual capture probability. Conversely, the use of an appropriate MR study design which accommodates heterogeneity associated with capture-occasion varying covariates due to animals moving between 'states' (i.e. geographic sites) can provide insight into how animals are distributed in a particular environment and the status and connectivity of subpopulations. 2. The Multistate Closed Robust Design was chosen to investigate: 1) the demographic parameters of Indo-Pacific bottlenose dolphins (Tursiops aduncus) subpopulations in coastal and estuarine waters of Perth, Western Australia; and 2) how they are related to each other in a metapopulation. Using four years of year-round photo-identification surveys across three geographic sites, we accounted for heterogeneity of capture probability based on how individuals distributed themselves across geographic sites and characterized the status of subpopulations based on their abundance, survival and interconnection. 3. MSCRD models highlighted high heterogeneity in capture probabilities and demographic parameters between sites. High capture probabilities, high survival and constant abundances described a subpopulation with high fidelity in an estuary. In contrast, low captures, permanent and temporary emigration and fluctuating abundances suggested transient use and low fidelity in an open coastline site. 4. Estimates of transition probabilities also varied between sites, with estuarine dolphins visiting sheltered coastal embayments more regularly than coastal dolphins visited the estuary, highlighting some dynamics within the metapopulation. 6. Synthesis and applications. To date, bottlenose dolphin studies using mark-recapture approach have focussed on investigating single subpopulations. Here, in a heterogeneous coastal-estuarine environment, we demonstrated that spatially structured bottlenose dolphin subpopulations contained distinct suites of individuals and differed in size, demographics and connectivity. Such insights into the dynamics of a metapopulation can assist in local-scale species conservation. The MSCRD approach is applicable to species/populations consisting of recognizable individuals and is particularly useful for characterizing wildlife subpopulations that vary in their vulnerability to human activities, climate change or invasive species.
Reproducibility and robustness of graph measures of the Associative-Semantic Network
<p>Matfiles and matlab scripts used to study the reproducibility and robustness of graph measures of the Associative-Semantic Network.</p>
Simulated data for "Spot-On: robust model-based analysis of single-particle tracking experiments" (MATLAB format)
<p>See 10.5281/zenodo.834787 for a more complete description.</p>
"Contrast based circular approximation for accurate and robust optic disc segmentation in retinal images" - Code
<p>A new method for automatic optic disc localization and segmentation is presented. The localization procedure combines vascular and brightness information to provide the best estimate of the optic disc center which is the starting point for the segmentation algorithm. A detection rate of 99.58% and 100% was achieved for the Messidor and ONHSD databases, respectively. A simple circular approximation to the optic disc boundary is proposed based on the maximum average contrast between the inner and outer ring of a circle centered on the estimated location. An average overlap coefficient of 0.890 and 0.865 was achieved for the same datasets, outperforming other state of the art methods. The results obtained confirm the advantages of using a simple circular model under non ideal conditions as opposed to more complex deformable models.</p>
Is the isotopic composition of precipitation a robust indicator for reconstructions of past tropical cyclones frequency? A case study on Réunion Island from rain and water vapor isotopic observations
<p>Isotopic composition of precipitation and water vapor and LMDZ-iso and ECHAM6-wiso simulations associated with the manuscript:</p> <p>Françoise Vimeux, Camille Risi, Christelle Barthe, Sören François, Alexandre Cauquoin, Olivier Jossoud, Jean-Marc Metzger, Olivier Cattani, Bénédicte Minster, and Martin Werner (2024). Is the isotopic composition of precipitation a robust indicator for reconstructions of past tropical cyclones frequency? A case study on Réunion Island from rain and water vapor isotopic observations. Journal of Geophysical Research: Atmospheres, 129, e2023JD039794. https://doi.org/10.1029/2023JD039794</p>
ToolBoxSF: Robustly interrogating machine learning-based scoring functions: what are they learning?
<p>This Zenodo repository provides comprehensive resources for the pre-print research paper titled "Robustly interrogating machine learning-based scoring functions: what are they learning?" Our collection includes Singularity containers containing pre-trained models, benchmark datasets, and training/test CSV files, offering valuable insights into the inner workings of machine learning-based scoring functions.</p><p>Key Components:</p><p>Singularity Containers:</p><ul><li>Machine Learning Models: Explore state-of-the-art scoring models used in the study, enabling reproducibility and in-depth analysis.</li><li>Environment Setup: Simplify model deployment and experimentation by utilizing our pre-configured environments.</li></ul><p>Benchmark Datasets:</p><ul><li>Curated benchmark datasets used in the pre-print, facilitating validation and evaluation of scoring functions.</li></ul><p>Training and Test CSV Files:</p><ul><li>Training and test data in CSV format, along with associated metadata.</li><li>Facilitate model testing and comparison using the provided data.</li></ul><p>This Zenodo collection is a valuable resource for researchers, data scientists, and machine learning enthusiasts seeking to replicate the study's findings, explore model behaviors, and conduct further investigations into machine learning-based scoring functions. Detailed documentation and usage instructions are included to support your research efforts at <a href="https://github.com/guydurant/toolboxsf">https://github.com/guydurant/toolboxs</a>f.</p><p>Citation Information: Please cite this Zenodo repository when using our resources in your work, and consider acknowledging the original pre-print when publishing research based on these materials.</p>
Robust CO2-abatement from early end-use electrification under uncertain power transition speed in China's netzero transition - Data and Plotting script
Open the record for dataset details and reuse information.
Robust candidate solutions derived by our explicit ROPAR algorithm
<p># sset25 represents the optimal robust candidate solutions with expected hydrological laternatin degree being 0.25.</p> <p> </p>
Evaluating the Robustness of Deep-learning Algorithm-selection Models by Evolving Adversarial Instances - Code and Data
<p>This repository contains the code and data for reproducibility of the paper 'Evaluating the Robustness of Deep-learning Algorithm-selection Models by Evolving Adversarial Instances'. </p> <p>The following files are included:</p> <ul> <li>Data.zip : contains the original instances in the datasets;</li> <li>Models.zip : trained Deep Neural Networks models used in the paper;</li> <li>New_instances.zip : generated instances using the approach;</li> <li>Parsed_data.zip : results and statistics of the experiments;</li> <li>script_adversarial_v3.py : Python script used to generate the results</li> </ul>
Data for article: Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists
<p><strong>### Update 03/02/2025: the most up to date version of the processing pipeline presented here, also working on Linux, is available on github: https://github.com/fischer-fjd/GCA/tree/main, and a worked example with open data from the Dutch AHN surveys is available on Zenodo: https://zenodo.org/records/14722001 ###<br></strong></p> <p>This is a collection of scripts and research data to assess the robustness of forest structure characterization from airborne laser scanning (ALS). It accompanies the article "Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists" (accepted in Methods in Ecology and Evolution on 25/08/2024). </p> <p>In the article, we assess the derivation of canopy height models (CHMs) from point cloud data, how sensitive CHM algorithms are to point cloud degradation (pulse density thinning, large scan angles, loss of higher-order returns) and how uncertainties and biases propagate to commonly used forest structure metrics. In addition, we provide a standardized processing pipeline in R to convert point clouds into CHMs. </p> <p>The main data source for this study are ALS point clouds from nine Australian research sites belonging to the Terrestrial Ecosystem Research Network (TERN, 5 km x 5 km extent each). The underlying data can be found here: https://portal.tern.org.au/metadata/TERN/4ff0b4c9-cfa0-4d09-9520-b5402adc583f. For one site (Robson Creek), we also used field data to assess the sensitivity of aboveground biomass estimates to ALS point cloud characteristics. Data are available here: https://portal.tern.org.au/metadata/supersite.174. </p> <p>To characterize climatic/environmental differences between sites, we used climatic data from the CHELSA/BIOCLIM+ climatology 1981-2010 (Brun et al. 2022: Global climate-related predictors at kilometer resolution for the past and future. Earth System Science Data, 14(12), 5573–5603. https://doi.org/10.5194/essd-14-5573-2022; Karger et al. 2017: Climatologies at high resolution for the earth's land surface areas. Scientific Data, 4(1), 170122. https://doi.org/10.1038/sdata.2017.122). </p> <p>We note that the enormous size of the full set of manipulated point clouds (original + thinned + individual flightlines: ~400 GB) and the derived raster products (~200 GB) by far exceeds limits on data storage in Zenodo. However, all analyses can be recreated from scratch from the openly available data and the R code in this repository. In addition, we include derived products for the nine study sites that allow to replicate results in the main text without any point cloud processing (CHMs and other rasters across thinned point clouds + summary statistics). </p> <p>The different data layers are:</p> <p><strong>01_rscripts.zip:</strong></p> <ul> <li>contains a sample script to test the processing pipeline (<em>test.processing.R</em>) as well as a collection of helper functions (<em>ALS_processing_helperfunctions_v40.R</em>); the script can be run directly after unzipping the folder, but an installation of LAStools (https://rapidlasso.de) is necessary (path_lastools = "PATH/TO/LASTOOLS/BIN"); we note that the script was developed on Windows PCs, its application with the recent Linux distribution of LAStools has not yet been tested</li> <li>contains the full set of scripts necessary to reproduce the analyses, including point cloud manipulations and derivation of CHMs from the raw data (<em>create.CHMs.R)</em> as well as the overall robustness analysis (<em>analyze.CHMs.R</em>); to replicate the processing of the raw point clouds step by step, .laz files should be downloaded from the TERN repository (cf. citation above) and placed in a "data" folder, with subfolders for each site and with the same naming conventions as in this repository (e.g., "/data/Alice Mulga")</li> </ul> <p><strong>02_reference.zip</strong></p> <ul> <li>contains reference digital surface models (DSMs), canopy height models (CHMs) and digital terrain models (DTMs) for all nine TERN sites, based on the original ALS point clouds</li> <li>note that these reference layers are produced with the "CHMhighest" algorithm, which provides an easily interpretable canopy description as long as pulse densities are high (>= 20 shots per squaremetre)</li> </ul> <p><strong>03_climate.zip</strong></p> <ul> <li>contains site coordinates</li> <li>contains the climate layers from the CHELSA climatology (cf. citation above, only used to evaluate climatic ranges of sites)</li> </ul> <p><strong>04_robson_additional.zip</strong></p> <ul> <li>contains biomass estimates for Robson Creek</li> <li>contains shapefiles for large trees at Robson Creek (only used for visualization purposes)</li> </ul> <p><strong>05_downsampling_pulse_[Site name].zip</strong></p> <ul> <li>[Site name] is a stand-in for the nine TERN sites (e.g., "Alice Mulga.zip", "Credo.zip", etc.)</li> <li>contains the data necessary to reproduce results in the main text of the study, i.e. DSMs, CHMs, and DTMs for all nine TERN sites, and at different pulse density levels (from 16 down to 0.5 laser shots per squaremetre)</li> <li>also contains calculated summary statistics for each site</li> </ul> <p>All zip files should be extracted into the same folder, except for 05_downsampling_pulse_[Site name].zip which should all be moved to a subfolder called "downsampling_pulse".</p>
Mesh used for Robust Discontinuity Indicators workflows
<p>Mesh used in the paper "Robust Discontinuity Indicators for High-Order Reconstruction of Piecewise Smooth Functions."<br>We have three types of spherical meshes: cubed-sphere (CS), quasi-uniform Voronoi (ICOD), and regionally refined (RRM) grids of the CS and ICOD types. They are NetCDF (.nc) meshes. Spherical meshes have different levels, from coarse to fine. All spherical meshes do not contain any function information.<br>We also uploaded the general surfaces of the plane, cylinder, and reamer bit. They are VTK meshes. The grids in the general surfaces are the same, but they contain different function values and generated RDI information.</p>
Dataset for: Online Virtual Machine Provisioning under Uncertainty: A Robust Approach to Ultimate Resource Utilization
<div> <div>In cloud resource scheduling and management, efficiency and quality are two vital yet conflicting objectives. Cloud providers, such as Huawei, prioritize quality by avoiding hotspots and aim to optimize and increase the utilization efficiency of PM resources without compromising quality. The typical online VM scheduling problem in cloud practice can be stated as follows: given a fixed number of PMs and a queue of arriving VMs, the goal is to place as many VMs as possible onto the PMs while ensuring that no hotspots occur.</div> <br> <div>The trace consists of a total of 297 instances, where each instance is represented by a JSON file named in the format X-1.json, X-2.json, and X-3.json. Here, X represents the number of PMs capable of hosting the arriving VMs, ranging from 2 to 100. For example, 50-1.json indicates that there are 50 available PMs to host the arriving VMs. For each value of X, there are three replicas denoted by the suffixes 1, 2, and 3.</div> <br> <div>There are multiple flavors of PM available in our cloud service provider. Readers can refer to our online paper, which we will provide the link for below, to learn about the specific flavor we used for academic and testing purposes. However, they are also free to use other typical flavors as per their requirements.</div> <br> <div>Each JSON file contains information about the VMs waiting to be assigned to PMs, and the fields for these VMs are as follows (each line represents a VM):</div> <br> <div> <ul> <li><strong>Created_at_point</strong>: the timestamp when the VM arrives. The list is already sorted in ascending order based on the arrival times.</li> </ul> </div> <ul> <li><strong>memory</strong>: the memory capacity of the VM defined by its flavor, measured in gigabytes (GB).</li> <li><strong>duration_point</strong>: the timestamps indicating the duration of the VM's usage. Each timestamp represents a 5-minute interval in practice.</li> <li><strong>vm_util</strong>: the real utilized capacity of the VM at each timestamp, with a similar meaning as the Hotspot Resolution trace. The length of the list is equal to the value of "duration_point".</li> </ul> </div> <div> </div>
Representative data accompanying the manuscript: Four-dimensional quantitative analysis of cell plate development in Arabidopsis using lattice light sheet microscopy identifies robust transition points between growth phases
<p>Representative data accompanying the manuscript: Sinclair R, Wang M, Jawaid MZ, Longkumer T, Aaron J, Rossetti B, Wait E, McDonald K, Cox D, Heddleston J, Wilkop T, Drakakaki G. (2024). <em>Four-dimensional quantitative analysis of cell plate development in Arabidopsis using lattice light sheet microscopy identifies robust transition points between growth phases.</em> J Exp Bot. 2024 Mar 4: erae091. doi: 10.1093/jxb/erae091.</p> <p>The data show YFP–RABA2a dynamics in dividing cells of Arabidopsis root tips using lattice light sheet microscopy. Treatments with or without Endosidin 7, a cytokinesis-specific callose deposition inhibitor, are shown.</p> <p>Data: </p> <p>22.3 YFP-RABA2A. </p> <p>23.9 YFP-RABA2A ES7 </p>
Dataset related to "Heterogeneous and higher-order cortical connectivity undergirds efficient, robust and reliable neural codes"
<p>This is an accompanying dataset to the article with the title "Heterogeneous and higher-order cortical connectivity undergirds efficient, robust and reliable neural codes" (DOI: <a href="https://doi.org/10.1101/2024.03.15.585196">10.1101/2024.03.15.585196</a>). It contains structural and activity data related to the morphologically detailed model of the rat somatosensory cortex (<a href="https://doi.org/10.1016/j.cell.2015.09.029" target="_blank" rel="noopener">Markram et al., 2015</a>), refered to as "BBP" data in the article. Specificaly, following data items are included:</p> <p><strong>Simulation data:</strong> <code>simulation.xz</code></p> <p><em>"Reliability" protocol:</em></p> <p>Separate folders <code>BlobStimReliability_O1v5-SONATA_<circuit-name></code> with simulation data using the baseline and all manipulated connectomes respectively (see Technical info below), each of which containing:</p> <ul> <li><code>working_dir/connectome.h5</code>: Connectivity matrix in <code>ConnectivityMatrix</code> format, which can be loaded using <a href="https://github.com/BlueBrain/ConnectomeUtilities" target="_blank" rel="noopener">ConnectomeUtilities</a>.</li> <li><code>working_dir/raw_spikes_exc_<sim>.npy</code>: Raw (excitatory) spikes in numpy .npy format, containing an array of spike times (first column) and corresponding neuron GIDs (second column). One file for each of the 10 simulations with different simulator seeds, i.e., <sim> is 0..9.</li> <li><code>working_dir/stim_stream.npy</code>: Stimulus train in numpy .npy format, containing the sequence of stimulus identities.</li> <li><code>working_dir/time_windows.npy</code>: Time windows in ms in numpy .npy format, corresponding to the stimulus train.</li> <li><code>working_dir/processed_data_store.h5</code>: Data store in HDF5 format with preprocessed spike signals (e.g., required for <em>Gaussian kernel reliability</em> computations), which contains... <ul> <li><code>spike_signals_exc</code>: Group of simulations datasets "sim_0" to "sim_9", each of which is an array of size <#gids x #t_bins> and contains binned spike signals filtered with a Gaussian kernel.</li> <li><code>sigma</code>: Sigma in ms of Gaussian kernel used for smoothing.</li> <li><code>gids</code>: List of excitatory neuron GIDs.</li> <li><code>t_bins</code>: List of time bins in ms.</li> <li><code>firing_rates</code>: Average firing rates per simulation (0..9; rows) and (excitatory) neuron GID (columns); average firing rates were computed as the inverse of the mean inter-spike interval per neuron.</li> </ul> </li> </ul> <p><em>"Classification" protocol:</em></p> <p>Single folder <code>Toposample_O1v5-SONATA</code> with simulation data using the baseline connectome, stored in a format compatible with the <a href="https://github.com/JasonPSmith/TriDy" target="_blank" rel="noopener">TriDy</a> (<a href="https://doi.org/10.1162/netn_a_00228">Conceição et al., 2022</a>) and <a href="https://github.com/BlueBrain/topological_sampling/tree/merge_samples" target="_blank" rel="noopener">TopoSampling</a> (<a href="https://doi.org/10.1371/journal.pone.0261702">Reimann et al., 2022</a>) pipelines, containing:</p> <ul> <li><code>toposample_input/connectivity.npz</code>: Sparse connectivity matrix in Compressed Sparse Column format , which can be loaded using <code>scipy.sparse.load_npz</code>.</li> <li><code>toposample_input/neuron_info.pickle</code>: Pandas dataframe in pickle format, which can be loaded using <code>pandas.read_pickle</code>, containing additional information about each neuron.</li> <li><code>toposample_input/raw_spikes_exc.npy</code>: Raw (excitatory) spikes in numpy .npy format, containing an array of spike times (first column) and corresponding neuron GIDs (second column).</li> <li><code>toposample_input/stim_stream.npy</code>: Stimulus train in numpy .npy format, containing the sequence of stimulus identities.</li> <li><code>toposample_input/time_windows.npy</code>: Time windows in ms in numpy .npy format, corresponding to the stimulus train.</li> </ul> <p> </p> <p><strong>Classification data:</strong> <code>classification.xz</code></p> <p>Selected neighborhoods and classification results for the "PCA" method (<a href="https://github.com/BlueBrain/topological_sampling/tree/merge_samples" target="_blank" rel="noopener">TopoSampling</a> pipeline) as well as the "network_based" method using active subnetworks (<a href="https://github.com/JasonPSmith/TriDy" target="_blank" rel="noopener">TriDy</a> pipeline, using <a href="https://github.com/jlazovskis/TriDy-tools" target="_blank" rel="noopener">TriDy-tools</a> wrapper), stored as:</p> <p><em>"PCA" method:</em></p> <ul> <li><code>PCA/community_database_PCA.pkl</code>: Pandas dataframe in pickle format, containing the binary selection of 50 neighborhood centers (neuron GIDs; rows) for the different selection parameters (columns).</li> <li><code>PCA/classification_results_PCA.pkl</code>: Pandas dataframe in pickle format, containing the classification accuracies for all selection parameters (rows) and 6 cross-validation folds plus mean (columns).</li> </ul> <p><em>"network_based" method:</em></p> <ul> <li><code>network_based/selections_reliability.pkl</code>: Pandas dataframe in pickle format, containing different combinations of first/second selection parameters for the <em>double selection procedure</em> (rows) together with neuron indices (w.r.t. the EXC subcircuit) of the corresponding 50 neighborhood centers (chief0..49; columns).</li> <li><code>network_based/partition_reliability.npy</code>: Partition indices in numpy .npy format required to launch the pipeline using <a href="https://github.com/jlazovskis/TriDy-tools" target="_blank" rel="noopener">TriDy-tools</a>, which is an array of 50 neuron indices (w.r.t. the full circuit!) of the neighborhood centers (columns) for each combination of selections as in <code>selections_reliability.pkl</code> (rows).</li> <li><code>network_based/results/...</code>: Subfolder containing a list of pickle files with the classification results using different featurization parameters, as indicated by the filename. Each file contains a Pandas dataframe with the classification accuracies and numbers of (non-zero) features (columns) for each combination of selections as in <code>selections_reliability.pkl</code> (rows).</li> </ul> <p> </p> <blockquote> <p><strong><em>Funding</em></strong></p> <p><em>Funding provided by the Swiss government’s ETH Board to the Blue Brain Project, a research center of the École polytechnique fédérale de Lausanne (EPFL).</em></p> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.