Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,329
datasets available to search
ShareScore release 0.7.1
Dataset results
3,329 results for “heterogeneity”
Heterogeneous selectivity and morphological evolution of marine clades during the Permian-Triassic mass extinction
<p>This is a supplementary repository, including the dataset and codes we used in this manuscript. we developed a new method, called DeepMorph to analyze the morphological evolution of six marine clades (i.e., ammonoids, bivalves, brachiopods, gastropods, ostracods, and conodonts ) during the Permian-Triassic mass extinction events. The taxonomy dataset was uploaded and contains 599 genera and 656 images, spanning from the latest Permian (Changhsingian) to the earliest Triassic (Induan). </p>
RafanoSet: Dataset of raw, manual and automatically annotated Raphanus Raphanistrum weed images for object detection and segmentation in Heterogenous Agriculture Environment
<p>This dataset is a collection of raw and annotated Multispectral (MS) images acquired in a heterogenous agricultural environment with MicaSense RedEdge-M camera. The spectra particularly Green, Blue, Red, Red Edge and Near Infrared (NIR) were acquired at sub-metre level.. <br><br>The MS images were labelled manually using VIA and automatically using Grounding DINO in combination with Segment Anything Model. The segmentation masks obtained using these two annotation techniqes over as well as the source code to perform necessary image processing operations are provided in the repository. The images are focussed over Horseradish (Raphanus Raphanistrum) infestations in Triticum Aestivum (wheat) crops.</p> <p>The nomenclature of sequecncing and naming images and annotations has been in this format: IMG_<scene number>_<spectral channel number><br><strong>_1</strong>: Blue<br><strong>_2</strong>: Green<br><strong>_3</strong>: Red<br><strong>_4</strong>: Near Infrared<br><strong>_5</strong>: RedEdge<br><br>Example: An image name <strong>IMG_0200_3 </strong>represents the scene number<strong> 200</strong> in <strong>Red channel</strong></p> <p>This dataset 'RafanoSet'is categorized in 6 directories namely 'Raw Images', 'Manual Annotations', 'Automated Annotations', 'Binary Masks - Manual', 'Binary Masks - Automated' and 'Codes'. The sub-directory 'Raw Images' consists of manually acquired 85 images in .PNG format. over 17 different scenes. The sub-directory 'Manual Annotations' consists of annotation file 'region_data' in COCO segmentation format. The sub-directory 'Automated Annotations' consists of 80 automatically annotated images in .JPG format and 80 .XML files in Pascal VOC annotation format.</p> <p>The scientific framework of image acquisition and annotations are explained in the Data in Brief paper which is the course of peer review. This is just a prerequisite to the data article. <br><br>Field experimentation roles:</p> <p>The image acquisition was performed by Mariano Crimaldi, a researcher, on behalf of Department of Agriculture and the hosting institution University of Naples Federico II, Italy.</p> <p>Shubham Rana has been the curator and analyst for the data under the supervision of his PhD supervisor Prof. Salvatore Gerbino. They are affiliated with Department of Engineering, University of Campania 'Luigi Vanvitelli'. </p> <p>Domenico Barretta, Department of Engineering has been associated in consulting and brainstorming role particularly with data validation, annotation management and litmus testing of the datasets.</p>
Energy input, habitat heterogeneity, and host specificity on avian haemosporidian diversity at continental scales
<p>The correct identification of biotic and abiotic drivers affecting parasite diversity and assemblage composition at different spatial scales is crucial for understanding how pathogen distribution responds to anthropogenic disturbance and climate change. Here, we used a database of avian haemosporidian parasites to identify such drivers and their effect on the taxonomic and phylogenetic diversity of genera Plasmodium, Haemoproteus, and Leucocytozoon from three zoogeographic regions. We explored how parasite diversity is related to energy input (i.e., temperature, precipitation, and potential evapotranspiration [PET]), to habitat heterogeneity (i.e., climatic seasonality, vegetation density, ecosystem heterogeneity, human disturbance, and host richness), and to a novel assemblage-level metric related to parasite niche overlap (degree of generalism). We found that the relative importance of the predictors differed between the three studied parasite genera and across diversity metrics. Among the most consistent predictors, host richness was positively related to the taxonomic diversity of the three genera. Energy input and human footprint explained the phylogenetic diversity of Haemoproteus. Finally, the degree of generalism explained the diversity of Plasmodium and Leucocytozoon. Our results suggest that different dimensions of haemosporidian diversity are shaped by energy input, host heterogeneity, and assembly processes related to parasite resource use within local parasite assemblages.</p>
Dataset for "Experimental and Modeling Insights into Mixing-Limited Reactive Transport in Heterogeneous Porous Media: Role of Stagnant Zones"
<p>This dataset contains the observed and simulated BTC of bimolecular transport experiment that was involved in "Yin et al., Experimental and Modeling Insights into Mixing-Limited Reactive Transport in Heterogeneous Porous Media: Role of Stagnant Zones".</p>
Plate interface geometry complexity and persistent heterogenous coupling revealed by a high-resolution earthquake focal mechanism catalog in Mentawai, Sumatra
<p>This website contains all the outputs from the study entitled “Plate interface geometry complexity and persistent heterogenous coupling revealed by a high-resolution earthquake focal mechanism catalog in Mentawai, Sumatra”. The contents include the seismic stations used in this study, obtained focal mechanism solutions, corresponding waveform fits, relocation results, and depth-phase modeling results. Each figure (started with ${ID}) is corresponding to the Earthquake ID as shown in Table S1.txt.</p>
Transcription start site analysis for heterogenous CD4+ T cells using 5′ scRNA-seq
<p>These datasets are generated by ReapTEC (read-level pre-filtering and transcribed enhancer call) using 5' single-cell RNA-seq data on human heterogenous CD4+ T cells. By taking advantage of a unique "cap signature" derived from the 5′-end of a transcript, ReapTEC simultaneously profiles gene expression and enhancer activity at nucleotide resolution using 5′-end single-cell RNA-sequencing (5′ scRNA-seq). The detail of ReapTEC pipeline is described in https://github.com/MurakawaLab/ReapTEC.</p>
Supplementary data to "The effects of small-scale heterogeneity on biomonitoring of desmid phytobenthos in Central European temperate mountain peatlands"
<p>The supplementary data consist of the files including the species-in-samples data and their associated NCV scores used for the analyses described in the manuscript submitted to hydrobiologia. In addition, two R scripts used for the analyses are included, too.</p> <p> </p>
Resources of IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources
<h2>IncRML resources</h2> <p>This Zenodo dataset contains all the resources of the paper 'IncRML: Incremental Knowledge Graph Construction from Heterogeneous Data Sources' submitted to the Semantic Web Journal's Special Issue on Knowledge Graph Construction. This resource aims to make the paper experiments fully reproducible through our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a> written in Python which was already used before in the <a href="https://doi.org/10.5281/zenodo.7837289" target="_blank" rel="noopener">Knowledge Graph Construction Challenge by the ESWC 2023 Workshop on Knowledge Graph Construction</a>. The exact Java JAR file of the RMLMapper (rmlmapper.jar) is also provided in this dataset which was used to execute the experiments. This JAR file was executed with Java OpenJDK 11.0.20.1 on Ubuntu 22.04.1 LTS (Linux 5.15.0-53-generic). Each experiment was executed 5 times and the median values are reported together with the standard deviation of the measurements.</p> <h2>Datasets</h2> <p>We provide both dataset dumps of the GTFS-Madrid-Benchmark and of real-life use cases from Open Data in Belgium.<br>GTFS-Madrid-Benchmark dumps are used to analyze the impact on execution time and resources, while the real-life use cases aim to verify the approach on different types of datasets since the GTFS-Madrid-Benchmark is a single type of dataset which does not advertise changes at all.</p> <h3>Benchmarks</h3> <ul> <li>GTFS-Madrid-Benchmark: change types with fixed data size and amount of changes: additions-only, modifications-only, deletions-only (11 versions)</li> <li>GTFS-Madrid-Benchmark: amount of changes with fixed data size: 0%, 25%, 50%, 75%, and 100% changes (11 versions)</li> <li>GTFS-Madrid-Benchmark: data size with fixed amount of changes: scales 1, 10, 100 (11 versions)</li> </ul> <h3>Real-world datasets</h3> <ul> <li>Traffic control center Vlaams Verkeerscentrum (Belgium): traffic board messages data (1 day, 28760 versions)</li> <li>Meteorological institute KMI (Belgium): weather sensor data (1 day, 144 versions)</li> <li>Public transport agency NMBS (Belgium): train schedule data (1 week, 7 versions)</li> <li>Public transport agency De Lijn (Belgium): busses schedule data (1 week, 7 versions)</li> <li>Bike-sharing company BlueBike (Belgium): bike-sharing availability data (1 day, 1440 versions)</li> <li>Bike-sharing company JCDecaux (EU): bike-sharing availability data (1 day, 1440 versions)</li> <li>OpenStreetMap (World): geographical map data (1 day, 1440 versions)</li> </ul> <h3>Ingestion</h3> <p>Real-world datasets LDES output was converted into SPARQL UPDATE queries and executed against Virtuoso to have an estimate for non-LDES clients how incremental generation impacted ingestion into triplestores.</p> <h2>Remarks</h2> <ol> <li>The first version of each dataset is always used as a baseline. All next versions are applied as an update on the existing version. The reported results are only focusing on the updates since these are the actual incremental generation.</li> <li>GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz datasets are not uploaded as GTFS-Madrid-Benchmark scale 100 because both share the same parameters (50% changes, scale 100). Please use GTFS-Scale-100-{ALL, CHANGE}.tar.xz for GTFS-Change-50_percent-{ALL, CHANGE}.tar.xz</li> <li>All datasets are compressed with XZ and provided as a TAR archive, be aware that you need sufficient space to decompress these archives! 2 TB of free space is advised to decompress all benchmarks and use cases. The expected output is provided as a ZIP file in each TAR archive, decompressing these requires even more space (4 TB).</li> </ol> <h2>Reproducing</h2> <p>By using our <a href="https://github.com/kg-construct/exectool" target="_blank" rel="noopener">experiment tool</a>, you can easily reproduce the experiments as followed:</p> <ol> <li>Download one of the TAR.XZ archives and unpack them.</li> <li>Clone the GitHub repository of our experiment tool and install the Python dependencies with '<em>pip install -r requirements.txt'.</em></li> <li>Download the rmlmapper.jar JAR file from this Zenodo dataset and place it inside the experiment tool root folder.</li> <li>Execute the tool by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive --runs=5 run</em>'. The argument '<em>--runs=5</em>' is used to perform the experiment 5 times.</li> <li>Once executed, you can generate the statistics by running: '<em>./exectool --root=/path/to/the/root/of/the/tarxz/archive stats</em>'.</li> </ol> <h2>Testcases</h2> <p>Testcases to verify the integration of RML and LDES with IncRML, see <a href="https://doi.org/10.5281/zenodo.10171394">https://doi.org/10.5281/zenodo.10171394</a></p>
TMax index: spatial heterogeneity of climate change as an experiential basis for skepticism
<p>To evaluate how the spatial heterogeneity of climate change affects the public’s willingness to accept scientific results that the climate is changing, we propose an index that accurately measures local changes in climate based on the number of days per year for which the year of the record high temperature is more recent than the year of the record low temperature. TMax index is calculated using the Global Historical Climatology Network (GHCN) dataset. Please refer to the PNAS paper for details on how the index is calculated.</p>
Correlation of urban avian species diversity present in heterogenous habitat types of the Silk city, Odisha, Eastern India
<p>This is the complete metadata and the R code required to do the analysis of the paper regarding birds of Berhampur city.</p>
Original datasets for : A computational homogenization framework with enhanced localization criterion for macroscopic cohesive failure in heterogeneous materials
<p>The original datasets from tests in the article: <strong> A computational homogenization framework with enhanced localization criterion for macroscopic cohesive failure in heterogeneous materials</strong>. The results are produced by the in-house fem codes of the Computational Mechanics group, CiTG, TU delft.</p>
Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity"
<p>Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity".</p> <p>Data and further information at GitHub repository https://github.com/kreutz-lab/dia-benchmarking (DOI: 10.5281/zenodo.6371925)</p>
Idealized Planar Array study for Quantifying Spatial heterogeneity (IPAQS) - Numerical Simulations
<p>The Idealized Planar Array study for Quantifying Spatial heterogeneity (IPAQS) is the result of a National Science Foundation (US) funded project, that aims at studying the effect of surface thermal heterogeneities of different length-scale on the atmospheric boundary layer. This project consisted of a computational effort (dataset here included), and an experimental effort (dataset being prepared for publication). </p> <p><strong>Overview of the numerical (Large Eddy) simulations:</strong></p> <p>The simulations are separated into two sets to study the differences between heterogeneous and homogeneous surfaces. In the first set, a total of seven configurations are considered, all with a homogeneous surface temperature fixed at a value of <span class="math-tex">\(T_s\)</span> = 290 K, and for which the geostrophic wind speed has been increased from 1 to 15 m s<sup>-1</sup> (i.e., U<sub>g</sub> = 1, 2, 3, 4, 6, 9, 15 m s<sup>−1</sup> ). These homogeneous cases are referred to as Homog-X, where X indicates the geostrophic wind speed corresponding case (see Margairaz et al. 2020a). In the second set, the surface temperature is distributed amongst square patches, where the temperature of each patch is determined by sampling a Gaussian distribution with a mean temperature of 290 K and a standard deviation of 5 K. In this case, three different patch sizes were considered (i.e., l<sub>h</sub> = 800, 400, and 200 m). The sizes of the heterogeneities were chosen to be of similar size (l<sub>h</sub> /l<sub>d</sub> ≈ 1), half the size (l<sub>h</sub> /l<sub>d</sub> ≈ 1/2), and about a quarter of the size (l<sub>h</sub> /l<sub>d</sub> ≈ 1/4) of the largest flow motions within the represented thermal boundary layer, assuming that this is of the order of the boundary-layer height (l<sub>d</sub> ∼ z <sub>i</sub> ). These heterogeneities are typically not resolved in NWP models. These cases have been studied for the same geostrophic wind speeds indicated above, and hereafter are referred to as PYYY-X-, where X indicates the corresponding geostrophic wind speed, and YYY refers to the size of the patches (e.g., P800_Ug1_ would be the heterogeneous case with patches of 800 m, and forced with Ug = 1 m s<sup>−1</sup> ). Additionally, for the case with larger patches, three different random distributions of the patches were considered to evaluate the potential effect of a given surface distribution for all geostrophic wind speeds. In this dataset we only include case v3. The LES imposed surface temperature distributions emulate the surface thermal conditions observed in Morrison et al. (2017 QJRMS, 2021 BLM, 2022 BLM), where measurements of the surface temperature were taken with a thermal camera at the SLTEST site of the US Army Dugway Proving Ground in Utah, USA. This is an ideal site with uniform roughness and a large unperturbed fetch, where surface thermal heterogeneities are naturally created by differences in surface salinity. In all studied cases, the surface roughness is assumed homogeneous, with z<sub>0</sub> = 0.1 m, and representative of a surface with sparse forest or farmland with many hedges (Brutsaert 1982; Stull 1988). The initial boundary-layer height is set to z<sub>i</sub> = 1000 m. The temperature profile is initialized with a mean air temperature of 285 K. At the top of the initial boundary layer, a capping inversion of 1000 m is used to limit its growth. The strength of this inversion is fixed at Γ = 0.012 K m<sup>−1</sup>. The atmospheric boundary layer (ABL) is considered dry and the latent heat flux is neglected in all cases. Further, in all simulations, the surface heat flux is computed using MOST, as explained in Margairaz et al. 2020a, where the surface temperature is kept constant in time throughout the simulations. Thus, there is no feedback from the atmosphere to the surface as the surface temperature does not cool down or warm up with local changes in velocities. As a consequence, the ABL gradually warms up as the simulations progress, and hence becomes less convective over time. However, the runs are not long enough for this to be significant. In addition, to ensure a degree of homogeneity within each patch and a certain degree of validity of MOST, note that even for the heterogeneous cases with the fewest amount of grid points per patch, a minimum of eight grid points is granted in each horizontal direction. The domain size is set to (L<sub>x</sub>, L<sub>y</sub>, L<sub>z</sub>) = (2π, 2π, 2) km at a grid size of (Nx , Ny , Nz ) = (256, 256, 256) resulting in a horizontal resolution of <span class="math-tex">\(\Delta\)</span>x = <span class="math-tex">\(\Delta\)</span>y = 24.5 m and a vertical grid spacing of <span class="math-tex">\(\Delta\)</span>z = 7.8 m. A timestep of <span class="math-tex">\(\Delta\)</span>t = 0.1 s is used to ensure the stability of the time integration. The two sets of simulations span a large range of geostrophic forcing conditions, allowing the study of the effect on the structure of the convective boundary layer (CBL) above a patchy surface compared to a homogeneous surface. The procedure used to spin up the simulations is the following: a spinup phase of four hours of real time is used to achieve converged turbulent statistics, which is then followed by an evaluation phase. During the latter, running averages are computed for the next hour of real time (dataset here published). Statistics have been computed for averaging times of 5 min to 1 h, showing statistical convergence at 30-min averages with negligible changes between the 30-min and the 60-min averages. The simulations cover a wide range of atmospheric stability regimes ranging from −z<sub>i</sub>/L < 5 to −z<sub>i</sub>/L > 700, and hence spanning from near neutral to highly convective scenarios.</p> <p><strong>Description of the Dataset as included in the NetCDF files:</strong></p> <p>Data for each study case is included in two files, one for momentum related variables, and one for temperature related variables. For example, the following files "P200_Ug1_Momentum.nc" and "P200_Ug1_Scalar.nc", include the 1h averaged variables for momentum and temperature for the case of 200 m surface patches with 1 m/s geostrophic winds. </p> <p>Each corresponding momentum file "PXXX_UgX_Momentum.nc" includes the following variables in a Python Xarray structure:</p> <ul> <li>'avgU' = mean streamwise wind speed; 'avgV' = mean spanwise wind speed; 'avgW' = mean vertical wind speed, 'avgP' = mean dynamic modified pressure field (<span class="math-tex">\(p^*\)</span>, see Margairaz et al 2020a),</li> <li>'avgU2', 'avgV2', 'avgW2' = correspond to <span class="math-tex">\(\overline{UU}\)</span>, <span class="math-tex">\(\overline{VV}\)</span>, and <span class="math-tex">\(\overline{WW}\)</span>, where the capital indicates the LES filtered variable.</li> <li>'avgUV', 'avgUW', 'avgVW' = correspond to <span class="math-tex">\(\overline{UV}\)</span>, <span class="math-tex">\(\overline{UW}\)</span>, and <span class="math-tex">\(\overline{VW}\)</span>. These variables together with the ones above are used to compute the Reynolds stress components (e.g. <span class="math-tex">\(R_{xz} = \overline{U}\overline{W} - \overline{UW}\)</span>).</li> <li>avgU3', 'avgV3', 'avgW3', 'avgU4', 'avgV4', 'avgW4' = correspond to the equivalent but instead of squared they are cubed and to the 4th power.</li> <li>'avgtxx','avgtyy','avgtzz','avgtxy','avgtxz','avgtyz' = These represent the corresponding averaged subgrid scale (SGS) stress.</li> <li>'avgdudz','avgdvdz','avgNut','avgCs' = Represent the averaged vertical derivatives, an averaged subgrid Nusselt number, and the Cs coefficient computed in the SGS model.</li> </ul> <p>Overall, there are a total of 26 variables related to the momentum field. Alternatively, the temperature fields are included in the "PXXX_UgX_Scalar.nc" files. These files include 10 variables, </p> <ul> <li>'avgT' = mean Temperature field, 'avgT2' = corresponds to <span class="math-tex">\(\overline{TT}\)</span>, 'avgUT' = correspond to <span class="math-tex">\(\overline{UT}\)</span>, 'avgVT' = correspond to <span class="math-tex">\(\overline{VT}\)</span>, 'avgWT' = correspond to <span class="math-tex">\(\overline{WT}\)</span>; one can use these terms to compute the corresponding Reynolds averaged turbulent fluxes as is the case for momentum. </li> <li>'avgUT_sgs','avgVT_sgs','avgWT_sgs' = These represent the corresponding subgrid scale fluxes.</li> <li>'avg_nus', avg_ds' = averaged subgrid Nusselt number, and the Ds coefficient computed in the scalar SGS model.</li> </ul> <p>All variables output from the LES are normalized by Tscale = 290 [K] when it includes dimensions of temperature, u_scale = 0.45 [m/s], when it relates to velocity fields, and z<sub>i</sub> = 1000 [m] for length scales.</p> <p>The only output variables that are expressed in dimensional form are those for the surface temperature included in the files "SurfTemp_DXXX.nc"</p> <p>Together with the data files we include a Python script that loads the data and includes it in two Xarray structures that one can then use to work with the datasets. </p>
Fig. 3 in Heterogeneity Studies Of Wild Clarias Gariepinus (Osteichthyes, Clariidae) Using Sds-Polyacrylamide Gel Electrophoresis
Fig. 3. Dendrogram obtained from Classical Cluster analysis using Paired group Bray-Curtis similarity index on C. gariepinus from two natural populations in Ado-Ekiti and Ilesa.
Fig. 2 in Heterogeneity Studies Of Wild Clarias Gariepinus (Osteichthyes, Clariidae) Using Sds-Polyacrylamide Gel Electrophoresis
Fig. 2. Dendrogram from Classical Cluster analysis using Paired group Bray-Curtis similarity index on Clarias gariepinus obtained in Ilesa, Osun State.
Fig. 1 in Heterogeneity Studies Of Wild Clarias Gariepinus (Osteichthyes, Clariidae) Using Sds-Polyacrylamide Gel Electrophoresis
Fig. 1. Dendrogram obtained from Classical Cluster analysis using Paired group Bray-Curtis similarity index on Clarias gariepinus in Ado-Ekiti.
Crop heterogeneity is positively associated with beneficial insect diversity in subtropical farmlands
<p>Increasing crop configurational heterogeneity – smaller crop fields with more field margins – has been repeatedly found to support farmland biodiversity. But research on compositional crop heterogeneity – the number and evenness of crop types – has usually shown only weak effects. However, much of this research has been conducted in large-scale temperate agroecosystems.</p> <p>We examined smallholder subtropical agroecosystems in southern China to assess the effects of crop heterogeneity on beneficial insect biodiversity. In addition to pollinators (bees, apoid and vespid wasps, butterflies), we studied dung beetles and dragonflies/damselflies, which are not usually considered in cropland heterogeneity studies, but are abundant in these multi-functional agroecosystems. We sampled these taxa in 468 transects placed inside 52 farms across three seasons (summer, spring, winter), collecting data on 27,245 insects belonging to 160 species.</p> <p>We found a strong positive effect of crop compositional heterogeneity (measured by Shannon-Wiener index) on dung beetle and dragonfly/damselfly diversity. Bees/wasps and butterflies, conversely, were positively affected by crop configurational heterogeneity (measured by cumulative field margin length).</p> <p>Field margin type, categorized by the structure of the dominant crop types, was consistently an important explanatory variable, with weedy margins having high insect diversity. The presence of a vegetable crop on one side of the field margin, compared to non-vegetable monocultures on both sides, increased diversity in 3/4 taxon-season comparisons made for rice, and 6/9 comparisons made for sugarcane or corn.</p> <p>Synthesis and applications. We demonstrate that crop compositional heterogeneity can support insects that respond to differences among crop types, including taxa that play a key role in nutrient cycling (dung beetles) and natural pest control (dragonflies/damselflies). Incorporating structurally diverse crops into monoculture Asian agroecosystems can reduce the adverse effects these intensive systems have on beneficial insects, and increase crucial ecosystem services.</p>
Riverscape heterogeneity in estimated Chinook Salmon emergence phenology and implications for size and growth
<p>Many salmonid-bearing rivers exhibit thermal and hydrologic heterogeneity at multiple spatial and temporal scales, but how this translates into spatiotemporal patterns of fry emergence is poorly understood. Understanding this variability is important because emergence timing determines the biophysical conditions fish first experience (e.g., temperature, flow, food supply), thereby influencing growth opportunities and survival during this critical life stage. We predicted spring Chinook Salmon (<em>Oncorhynchus tshawytscha</em>) emergence phenology across four NE Oregon subbasins over 5-9 years using empirical spawning and temperature data. We then related inter-annual emergence timing estimates to juvenile salmon size and growth rates at consistent sampling locations. There were clear longitudinal patterns of predicted emergence timing in each subbasin: the shape of these patterns was consistent among years, but not among subbasins. In two subbasins emergence occurred progressively later with distance upstream, whereas in the other two subbasins emergence was earliest at upstream sites. Within each year, median emergence dates among sites within each subbasin ranged between 44 and 58 days. This spatial variation was comparable to inter-annual variation, with median emergence dates for a given location in each subbasin ranging between 47 to 74 days among years. Contrary to our expectations, juvenile salmon were not larger in years with earlier emergence, owing to slower spring and summer growth rates compared to years with later emergence. Despite large inter-annual variation in emergence dates, these results suggest that other factors (e.g., stream flow, temperature, density-dependence) were more important than growth duration in determining juvenile salmon growth rates and size among years. We demonstrated considerable spatial and inter-annual variation in emergence phenology within these subbasins. Understanding how this variation translates to spatiotemporal patterns of juvenile salmon habitat use, growth, and survival has important implications for guiding restoration efforts and understanding how climate change may impact these populations.</p>
Using single-worm data to quantify heterogeneity in Caenorhabditis elegans-bacterial interactions
<p>The nematode <em>Caenorhabditis elegans</em> is a model system for host-microbe and host-microbiome interactions. Many studies to date use batch digests rather than individual worm samples to quantify bacterial load in this organism. Here it is argued that the large inter-individual variability seen in bacterial colonization of the <em>C. elegans</em> intestine is informative, and that batch digest methods discard information that is important for accurate comparison across conditions. As describing the variation inherent to these samples requires large numbers of individuals, a convenient 96-well plate protocol for disruption and colony plating of individual worms is established.</p>
PGB: A PubMed Graph Benchmark for Heterogeneous Network Representation Learning
<p>PubMed Graph Benchmark (PGB) aggregates the metadata associated with the biomedical articles from PubMed into a unified source. The benchmark contains metadata including title, abstract, authors, in/out citations, MeSH terms, MeSH hierarchy, venue, publication type, and chemicals.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.