Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
750
datasets available to search
ShareScore release 0.9.0
Dataset results
750 results for “heterogeneous data”
Data package supporting manuscript "Widespread Heterogeneity in Density-Dependent Mortality of Nearshore Fishes"
This repository contains the complete data synthesis and analysis pipeline for a global meta-analysis on density-dependent mortality in reef fishes. We estimated mortality parameters (α and β) from >30 ecological studies and explored how ecological traits, experimental methods, and phylogenetic history explain variation in density dependence. It comprises eight data tables in csv format, three .tre files for phylogenetic trees (see method document for data sources), and the zipped code folder (including 12 R scripts) to ensure transparent, end-to-end reproducibility of data processing, analysis, and visualization. This package supports the manuscript “Widespread Heterogeneity in Density-Dependent Mortality of Nearshore Fishes” by Stier & Osenberg (Ecology Letters).
Data belonging to: Teurlincx, S., Verhofstad, M. J., Bakker, E. S., & Declerck, S. A. (2018). Managing successional stage heterogeneity to maximize landscape-wide biodiversity of aquatic vegetation in ditch networks. Frontiers in plant science, 9, 1013.
<p>Data belonging to the paper Teurlincx, S., Verhofstad, M. J., Bakker, E. S., & Declerck, S. A. (2018). Managing successional stage heterogeneity to maximize landscape-wide biodiversity of aquatic vegetation in ditch networks. Frontiers in plant science, 9, 1013.</p> <p>Data includes analysis scripts (R Language) and all used data files. Data is composed of location information of the different sites, environmental conditions on site and vegetation composition.</p>
Data from: Heterogeneity in habitat and nutrient availability facilitate the co-occurrence of N2 fixation and denitrification across wetland - stream - lake ecotones of Lakes Superior and Huron
Great Lakes coastlines are mosaics of wetland, stream, and lake habitats, characterized by a high degree of spatial heterogeneity that may facilitate the co-occurrence of seemingly incompatible biogeochemical processes due to variation in environmental factors that favor each process. We measured nutrient limitation and rates of N2 fixation and denitrification along transects in 5 wetland - stream - lake ecotones with different nutrient loading in Lakes Superior and Huron and hypothesized that rates of both processes would be related to nutrient limitation status, habitat type, and environmental characteristics including temperature, nutrient concentrations, and organic matter quality. This data package includes information on sampling sites, dates and locations; rates of N fixation and denitrification measured at each site, date and transect location; and biomass information from nutrient diffusing substrates deployed on the study transects.
Long-term live imaging and multiscale analysis identify heterogeneity and core principles of epithelial organoid morphogenesis - Image data
<p>The dataset contains raw imaging data from the work:</p> <p>"Long-term live imaging and multiscale analysis identify heterogeneity and core principles of epithelial organoid morphogenesis"</p> <p>The dataset is organized as the following: the "FigureX_" or SupplementaryFigure_X" suffix in the filename refers to the figure in the paper in which the raw data is analyzed and/or visualized. The data is "raw", i.e. not processed. However, in many cases, maximum projections of the original 3D image stacks have been uploaded due to size limitations. The total size of the image stacks approaches 0.5TB. To access the full 3D image stacks please contact the corresponding author (Francesco Pampaloni, fpampalo@bio.uni-frankfurt.de).</p> <p><strong>Authors</strong></p> <p>Lotta Hof<sup>1</sup>*, Till Moreth<sup>1</sup>*, Michael Koch<sup>1</sup>, Tim Liebisch<sup>2</sup>, Marina Kurtz<sup>3</sup>, Julia Tarnick<sup>4</sup>, Susanna M. Lissek<sup>5</sup>, Monique M.A. Verstegen<sup>6</sup>, Luc J.W. van der Laan<sup>6</sup>, Meritxell Huch<sup>7</sup>, Franziska Matthäus<sup>2</sup>, Ernst H.K. Stelzer<sup>1</sup>, Francesco Pampaloni<sup>1§</sup></p> <p><sup>1</sup>Physical Biology Group, Buchmann Institute for Molecular Life Sciences (BMLS), Goethe-Universität Frankfurt am Main, Frankfurt am Main, Germany</p> <p><sup>2</sup>Faculty of Biological Sciences, Goethe-Universität Frankfurt am Main, Frankfurt am Main, Germany</p> <p><sup>3</sup>Department of Physics, Goethe-Universität Frankfurt am Main, Frankfurt am Main, Germany</p> <p><sup>4</sup>Deanery of Biomedical Science, University of Edinburgh, Edinburgh, United Kingdom</p> <p><sup>5</sup>Experimental Medicine and Therapy Research, University of Regensburg, Regensburg, Germany</p> <p><sup>6</sup>Department of Surgery, Erasmus MC – University Medical Center, Rotterdam, The Netherlands</p> <p><sup>7</sup>The Wellcome Trust/CRUK Gurdon Institute, University of Cambridge, Cambridge, United Kingdom. Present address: Max Planck Institute of Molecular Cell Biology and Genetics, Dresden, Germany</p> <p>*contributed equally</p> <p><sup>§</sup>corresponding author: fpampalo@bio.uni-frankfurt.de</p> <p><strong>Abstract</strong></p> <p><em>Background</em></p> <p>Organoids are morphologically heterogeneous three-dimensional cell culture systems and serve as an ideal model for understanding the principles of collective cell behaviour in mammalian organs during development, homeostasis, regeneration and pathogenesis. To investigate the underlying cell organisation principles of organoids, we imaged hundreds of pancreas and cholangio carcinoma organoids in parallel using light sheet and bright field microscopy for up to seven days.</p> <p><em>Results</em></p> <p>We quantified organoid behaviour at single-cell (microscale), individual-organoid (mesoscale), and entire-culture (macroscale) levels. At single-cell resolution, we monitored formation, monolayer polarisation and degeneration, and identified diverse behaviours, including lumen expansion and decline (size oscillation), migration, rotation and multi-organoid fusion. Detailed individual organoid quantifications lead to a mechanical 3D agent-based model. A derived scaling law and simulations support the hypotheses that size oscillations depend on organoid properties and cell division dynamics, which is confirmed by bright field microscopy analysis of entire cultures.</p> <p><em>Conclusion</em></p> <p>Our multiscale analysis provides a systematic picture of the diversity of cell organisation in organoids by identifying and quantifying the core regulatory principles of organoid morphogenesis.</p>
Research data supporting "Impact of global heterogeneity of renewable energy supply on heavy industrial production and green value chains"
<p>Research data supporting the peer-reviewed article "Impact of global heterogeneity of renewable energy supply on heavy industrial production and green value chains" by the same authors.</p>
scRNA-seq data for article: Kupffer cell and recruited macrophage heterogeneity orchestrate granuloma maturation and hepatic immunity in visceral leishmaniasis
<p>Single-cell RNA-seq dataset from sorted CD11bInt, F4/80Hi, CD64+ mouse liver cells in naive or Leishmania infantum-infected animals at 42 d.p.i.. Data analyses and results are described in manuscript: "Kupffer cell and recruited macrophage heterogeneity orchestrate granuloma maturation and hepatic immunity in visceral leishmaniasis". Data files are Seurat objects in RDS format. Filtered-out potential doublets, low quality cells and dying cells (excluded cells with <1000 genes detected, cells with >6000 genes detected, cells with mitochondrial gene expression > 10% and cells with <5000 transcript molecules). Data normalization, scaling and integration performed using Seurat.</p> <p>Filtered dataset containing all KCs and macrophages is in the "pessenda_KC_Macro_seurat" file.</p> <p>Our data were then mapped onto a reference dataset published by Remmerie et al. (DOI: 10.1016/j.immuni.2020.08.004) for annotation consistent with the literature. The reference mapped object can be found in the "pessenda_refmap_KC_Macro_seurat" file.</p> <p>Dataset containing the additional analysis of CLEC4F-TIM4+ FACS-sorted KCs can be found in the "pessenda_refmap_KCTimPos_seurat" file.</p>
Synthetic Data for Uplift Modeling and Heterogenous Treatment Effect with Known Counterfactuals and ITE
<p>This dataset is designed and simulated for evaluating uplift modeling. The data generation process is based on a logistic regression model - no real data is included or used for generating this dataset.</p> <p>This dataset has several signatures:</p> <ul> <li>It generates features with various patterns associated with the outcome variable and the causal effect (or treatment effect). Thus it is suitable for evaluating feature importance and model interpretation for uplift modeling.</li> <li>The true counterfactual outcomes under control and treatment are known for each user, as well as the true ITE (Individual treatment effect).</li> </ul> <p>This dataset consists of 50 trials (replicates with different random seeds), each trial with 20,000 samples and 36 features. The outcome variable is binary, which makes this dataset for classification problems. The samples are equally split for the control and treatment groups (10,000 samples in each group in each trial).</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect.</p> <p>To simulate the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <p> Trial ID: 'trial_id'<br> Experiment group label: 'treatment_group_key'<br> Outcome variable (classification label): 'conversion'<br> Feature names: ['x1_informative',<br> 'x2_informative',<br> 'x3_informative',<br> 'x4_informative',<br> 'x5_informative',<br> 'x6_informative',<br> 'x7_informative',<br> 'x8_informative',<br> 'x9_informative',<br> 'x10_informative',<br> 'x11_irrelevant',<br> 'x12_irrelevant',<br> 'x13_irrelevant',<br> 'x14_irrelevant',<br> 'x15_irrelevant',<br> 'x16_irrelevant',<br> 'x17_irrelevant',<br> 'x18_irrelevant',<br> 'x19_irrelevant',<br> 'x20_irrelevant',<br> 'x21_irrelevant',<br> 'x22_irrelevant',<br> 'x23_irrelevant',<br> 'x24_irrelevant',<br> 'x25_irrelevant',<br> 'x26_irrelevant',<br> 'x27_irrelevant',<br> 'x28_irrelevant',<br> 'x29_irrelevant',<br> 'x30_irrelevant',<br> 'x31_uplift_increase',<br> 'x32_uplift_increase',<br> 'x33_uplift_increase',<br> 'x34_uplift_increase',<br> 'x35_uplift_increase',<br> 'x36_uplift_increase']<br> True underlying control conversion probability: 'control_conversion_prob'<br> True underlying treatment conversion probability: 'treatment1_conversion_prob'<br> True treatment effect: 'treatment1_true_effect'</p>
Data Set for the Journal Article "Autonomous Reaction Network Exploration in Homogeneous and Heterogeneous Catalysis"
<p>This dataset includes the XYZ structures of the centroids of all compounds found. Charge and multiplicity are given in the comment line of each XYZ file.</p>
Row sequcenes data for assessing the risks of potential pathogens and antibiotic resistance genes among heterogeneous habitats in a temperate estuary wetland
<p>The study included 118 usable samples within three different habitats (water, soil, and sediment) across the Liaohe River basin to the Red Beach wetland collected from seven papers, and all of the sequence files were uploaded for availability.</p>
Data for "SeaMoon: from protein language models to continuous structural heterogeneity"
<p>Datasets used for development of SeaMoon: <br><a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</p> <p>This upload contains the following data:</p> <ul> <li><strong>precomputed_emb.tar.gz</strong> is a compressed archive containing the precomputed data used for training and testing the models of the SeaMoon method, in Torch <strong>.pt </strong>format. <br>The file prefixes consist of two IDs, "ID1_ID2_", identifying the <a href="https://github.com/PhyloSofS-Team/DANCE">DANCE</a> [1] protein conformational collection used for its generation. "ID1" represents the first member of the collection in alphabetical order, while "ID2" is the reference conformation for the structural alignment. The "ESM_data" or "ProstT5_data" suffixes designate the type of embeddings, generated by either ESM2 [2] or ProstT5 [3].<br>The dictionnary contains the following keys: <ul> <li><strong>emb:</strong> The per-residue embedding.</li> <li><strong>data: </strong>A tuple containing "ID2" (the reference), the amino acid sequence, and the coverage of the positions in the original DANCE collection.</li> <li><strong>eigvect:</strong> The eigenvectors of the covariance matrix of the "ID1_ID2" collection, centered on reference conformaton "D2".</li> <li><strong>eigval: </strong>The associated eigenvalues.</li> <li><strong>ref:</strong> The coordinates of the C-alpha atoms of the reference conformaton "ID2".</li> </ul> </li> <li><strong>train_list.txt, train_list_5ref.txt, val_list.txt </strong>and<strong> test_list.txt</strong> contain the identifiers of the samples used for training and evaluating the SeaMoon models. In the "5ref" setting, we used up to 5 reference conformations per collection. </li> </ul> <p>For details on SeaMoon see:</p> <div> <div>SeaMoon: Prediction of molecular motions based on language models</div> </div> <div>Valentin Lombard, Dan Timsit, Sergei Grudinin, Elodie Laine</div> <div>bioRxiv 2024.09.23.614585; doi: https://doi.org/10.1101/2024.09.23.614585</div> <div> </div> <div>For more information on data usage and generation please see <a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</div> <div> </div> <div>Abstract:</div> <p>How protein move and deform determines their interactions with the environment and is thus of utmost importance for cellular functioning. Following the revolution in single protein 3D structure prediction, researchers have focused on repurposing or developing deep learning models for sampling alternative protein conformations. In this work, we explored whether continuous compact representations of protein motions could be predicted directly from protein sequences, without exploiting nor sampling protein structures. Our approach, called SeaMoon, leverages protein Language Model (pLM) embeddings as input to a lightweight (~1M trainable parameters) convolutional neural network. SeaMoon achieves a success rate of up to 40% when assessed against ~1,000 collections of experimental conformations exhibiting a wide range of motions. SeaMoon capture motions not accessible to the normal mode analysis, an unsupervised physics-based method relying solely on a protein structure's 3D geometry, and generalises to proteins that do not have any detectable sequence similarity to the training set. SeaMoon is easily retrainable with novel or updated pLMs. </p> <p> </p> <p>[1] Lombard, V.; Grudinin, S.; Laine, E. Explaining Conformational Diversity in Protein Families through Molecular Motions. Scientific Data 2024, 11, 752.</p> <p>[2] Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.; Lu, W.; Smetanin, N.; Verkuil, R.; Kabeli, O.; Shmueli, Y.; Dos Santos Costa, A.; Fazel-Zarandi, M.; Sercu, T.; Candido, S.; Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123–1130.</p> <p>[3] Heinzinger, M.; Weissenow, K.; Sanchez, J. G.; Henkel, A.; Steinegger, M.; Rost, B. ProstT5: Bilingual language model for protein sequence and structure. bioRxiv 2023, 2023–07.</p>
Supplementary Movies and Source Data for: Quantitative real-time in-cell imaging reveals heterogeneous clusters of proteins prior to condensation
<p>Supplementary Movies and raw data for the manuscript: "Quantitative real-time in-cell imaging reveals heterogeneous clusters of proteins prior to condensation":</p> <p>Source_Data.zip: Supplementary Code, Supplementary Data and Weka Analysis</p> <p>Lan_supplementary_movies_AVI.zip: Supplementary movies as AVI</p> <p>Lan_supplementary_movies_MP4.zip: Supplementary movies as MP4</p> <p>Lan_raw_movies.zip: Raw TIFF stacks of the movies.</p> <p>Lan_supplementary_movies.zip: Old version of the movies.</p>
Code and data for manuscript: Incorporating environmental heterogeneity and observation effort to predict host distribution and viral spillover from a bat reservoir.
<p>This is the source code and data required to reproduce data analysis and figures from the manuscript, "Incorporating environmental heterogeneity and observation effort to predict host distribution and viral spillover from a bat reservoir". </p>
Data from Citizen science data reveal regional heterogeneity in phenological response to climate in the large milkweed bug, Oncopeltus fasciatus
These data include annotations for life stage, mating behavior, and plant part occupancy of large milkweed bug observations in North America as well as information about climate and environment.
Data for "Plasmon excitations in chemically heterogeneous nanoarrays"
<p>The data includes atomic structures, photoabsorption spectra, and noninteracting spectra of the systems modeled in the article "Plasmon excitations in chemically heterogeneous nanoarrays" by Kevin Conley <em>et al</em>.</p> <p>See <em>README.md</em> in the archive for a detailed description.</p>
Data from: Using genetic relatedness to understand heterogeneous distributions of urban rat-associated pathogens
<p>Urban Norway rats (<i>Rattus norvegicus</i>) carry several pathogens transmissible to people. However, pathogen prevalence can vary across fine spatial scales (i.e., by city block). Using a population genomics approach, we sought to describe rat movement patterns across an urban landscape, and to evaluate whether these patterns align with pathogen distributions. We genotyped 605 rats from a single neighborhood in Vancouver, Canada and used 1,495 genome-wide single nucleotide polymorphisms to identify parent-offspring and sibling relationships using pedigree analysis. We resolved 1,246 pairs of relatives, of which only 1% of pairs were captured in different city blocks. Relatives were primarily caught within 33 meters of each other leading to a highly leptokurtic distribution of dispersal distances. Using binomial generalized linear mixed models we evaluated whether family relationships influenced rat pathogen status with the bacterial pathogens <i>Leptospira interrogans</i>, <i>Bartonella tribocorum</i>, and <i>Clostridium difficile</i>, and found that an individual's pathogen status was not predicted any better by including disease status of related rats. The spatial clustering of related rats and their pathogens lends support to the hypothesis that spatially restricted movement promotes the heterogeneous patterns of pathogen prevalence evidenced in this population. <span>Our findings also highlight the utility of evolutionary tools to understand movement and rat-associated health risks in urban landscapes.</span></p>
Single-cell mouse and PC9 data for "TP53 loss with whole genome doubling mediates heterogeneous intra-patient therapy response through Chromosomal Instability"
<p>This repository includes the processed data (including copy number profiles and related analysis) for the E/EP mouse tumors and for the PC9 resistance cell lines for all the analyses of the manuscript "TP53 loss with whole genome doubling mediates heterogeneous intra-patient therapy response through Chromosomal Instability".</p><p>The code for the related analyses is available in GitHub at https://github.com/zaccaria-lab/TP53loss_WGD</p>
Supplementary data to "The effects of small-scale heterogeneity on biomonitoring of desmid phytobenthos in Central European temperate mountain peatlands"
<p>The supplementary data consist of the files including the species-in-samples data and their associated NCV scores used for the analyses described in the manuscript submitted to hydrobiologia. In addition, two R scripts used for the analyses are included, too.</p> <p> </p>
Resources of IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources
<h2>IncRML resources</h2> <p>This Zenodo dataset contains all the resources of the paper 'IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources' submitted to the Semantic Web Journal's Special Issue on Knowledge Graph Construction. This resource aims to make the paper experiments fully reproducible through our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a> written in Python which was already used before in the <a href="https://doi.org/10.5281/zenodo.7837289" target="_blank" rel="noopener">Knowledge Graph Construction Challenge by the ESWC 2023 Workshop on Knowledge Graph Construction</a>. The exact Java JAR file of the RMLMapper (rmlmapper.jar) is also provided in this dataset which was used to execute the experiments. This JAR file was executed with Java OpenJDK 11.0.20.1 on Ubuntu 22.04.1 LTS (Linux 5.15.0-53-generic). Each experiment was executed 5 times and the median values are reported together with the standard deviation of the measurements.</p> <h2>Datasets</h2> <p>We provide both dataset dumps of the GTFS-Madrid-Benchmark and of real-life use cases from Open Data in Belgium.<br>GTFS-Madrid-Benchmark dumps are used to analyze the impact on execution time and resources, while the real-life use cases aim to verify the approach on different types of datasets since the GTFS-Madrid-Benchmark is a single type of dataset which does not advertise changes at all.</p> <h3>Benchmarks</h3> <ul> <li>GTFS-Madrid-Benchmark: change types with fixed data size and amount of changes: additions-only, modifications-only, deletions-only (11 versions)</li> <li>GTFS-Madrid-Benchmark: amount of changes with fixed data size: 0%, 25%, 50%, 75%, and 100% changes (11 versions)</li> <li>GTFS-Madrid-Benchmark: data size with fixed amount of changes: scales 1, 10, 100 (11 versions)</li> </ul> <h3>Real-world datasets</h3> <ul> <li>Traffic control center Vlaams Verkeerscentrum (Belgium): traffic board messages data (1 day, 28760 versions)</li> <li>Meteorological institute KMI (Belgium): weather sensor data (1 day, 144 versions)</li> <li>Public transport agency NMBS (Belgium): train schedule data (1 week, 7 versions)</li> <li>Public transport agency De Lijn (Belgium): busses schedule data (1 week, 7 versions)</li> <li>Bike-sharing company BlueBike (Belgium): bike-sharing availability data (1 day, 1440 versions)</li> <li>Bike-sharing company JCDecaux (EU): bike-sharing availability data (1 day, 1440 versions)</li> <li>OpenStreetMap (World): geographical map data (1 day, 1440 versions)</li> </ul> <h3>Ingestion</h3> <p>Real-world datasets LDES output was converted into SPARQL UPDATE queries and executed against Virtuoso to have an estimate for non-LDES clients how incremental generation impacted ingestion into triplestores.</p> <h2>Remarks</h2> <ol> <li>The first version of each dataset is always used as a baseline. All next versions are applied as an update on the existing version. The reported results are only focusing on the updates since these are the actual incremental generation.</li> <li>GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz datasets are not uploaded as GTFS-Madrid-Benchmark scale 100 because both share the same parameters (50% changes, scale 100). Please use GTFS-Scale-100-{ALL, CHANGE}.tar.xz for GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz</li> <li>All datasets are compressed with XZ and provided as a TAR archive, be aware that you need sufficient space to decompress these archives! 2 TB of free space is advised to decompress all benchmarks and use cases. The expected output is provided as a ZIP file in each TAR archive, decompressing these requires even more space (4 TB).</li> </ol> <h2>Reproducing</h2> <p>By using our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a>, you can easily reproduce the experiments as followed:</p> <ol> <li>Download one of the TAR.XZ archives and unpack them.</li> <li>Clone the GitHub repository of our experiment tool and install the Python dependencies with '<em>pip install -r requirements.txt'.</em></li> <li>Download the rmlmapper.jar JAR file from this Zenodo dataset and place it inside the experiment tool root folder.</li> <li>Execute the tool by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive --runs=5 run</em>'. The argument '<em>--runs=5</em>' is used to perform the experiment 5 times.</li> <li>Once executed, you can generate the statistics by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive stats</em>'.</li> </ol> <h2>Testcases</h2> <p>Testcases to verify the integration of RML and LDES with IncRML, see <a href="https://doi.org/10.5281/zenodo.10171394">https://doi.org/10.5281/zenodo.10171394</a></p>
Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity"
<p>Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity".</p> <p>Data and further information at GitHub repository https://github.com/kreutz-lab/dia-benchmarking (DOI: 10.5281/zenodo.6371925)</p>
Using single-worm data to quantify heterogeneity in Caenorhabditis elegans-bacterial interactions
<p>The nematode <em>Caenorhabditis elegans</em> is a model system for host-microbe and host-microbiome interactions. Many studies to date use batch digests rather than individual worm samples to quantify bacterial load in this organism. Here it is argued that the large inter-individual variability seen in bacterial colonization of the <em>C. elegans</em> intestine is informative, and that batch digest methods discard information that is important for accurate comparison across conditions. As describing the variation inherent to these samples requires large numbers of individuals, a convenient 96-well plate protocol for disruption and colony plating of individual worms is established.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.