Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,551
datasets available to search
ShareScore release 0.7.1
Dataset results
1,551 results for “Availability”
Pseudomonas aeruginosa predicted prophages from publicly available genomes
<p>Through the Genome Information by Organism section of the NCBI Genome database, <em>P. aeruginosa</em> bacterial genomic assemblies were downloaded (September 2020). Genome quality was assessed totaling 5,383 genomes total. All 5,383 genomes were then entered into VirSorter v.1 (https://github.com/simroux/VirSorter). The data set provided here includes all category 1 and category 4 predicted prophage sequences.</p>
Fig. 1 in Currently available data on Borneo geometrid moths do not provide evidence for a Pleistocene rainforest refugium
Fig. 1. Sundaland during the Pleistocene. Lowered sea levels exposed much of the Sunda shelf, unifying Borneo, Sumatra, Java, Bali and the Malay Peninsula. Today's coastlines and country borders are shown for orientation. Dark grey areas denote hypothesized rainforest refugia (redrawn from Gathorney-Hardy et al., 2002). Stars indicate sampling locations in Borneo as used for this study; the southern-most location is Camp Foyle (see Appendix for details).
Local fruit availability and en route wind conditions are poor predictors of bird abundance and composition during fall migration in coastal Yucatán Peninsula
<p>In migratory stopover habitats, bird abundance and composition change on a near daily basis. On any given day, the local bird community should reflect local environmental conditions but also the environments that birds encountered previously along their migratory route. For example, during fall migration, the coast of the Yucatán Peninsula in Mexico receives birds that have just crossed the Gulf of Mexico and their abundance and composition may be associated with regional factors such as wind conditions experienced on previous dates but also local factors such as fruit availability. Thus, we used three data sets to quantify the influence of wind and fruit on near daily variation in bird abundance and composition. Our bird data consists of the number of individuals per species captured using mist nets for two coastal national parks in the Yucatán Peninsula during fall migration in 2016 and 2017. The data is provided daily, as are the "net-hours," i.e., the sum of the hours each net was open summed across nets. Thus, we analyzed bird captures standardized by net-hours. Our fruit data consists of the number of unripe and ripe fruits per species counted on each side of each mist net lane at various points during fall migration. Our wind data consists of the wind costs birds arriving at our sites would have experienced when departing the north coast of the Gulf of Mexico two days prior to their arrival. Wind cost reflects wind speed and direction and it was calculated using the wind.dl_2 function in the rWind package. The wind costs are averaged across the entire US Gulf coastline. We used Moran eigenvector maps to quantify the temporal structure of the bird, wind, and fruit data and we partitioned the variance in the bird data into the components explainable by wind or fruit, the temporal structure of wind or fruit, and temporal structure independent of wind or fruit. After running the analysis, we did not find a strong association between daily changes in bird abundance or community composition with wind conditions and ripe fruit availability. Thus, despite wind and fruit being known to be important to individual birds (influencing stopover duration and departure decisions), their effects might not scale up to be drivers of population and community-level variation.</p>
Using unoccupied aerial vehicles to estimate availability and group size error for aerial surveys of coastal dolphins
<p><span>Aerial surveys are frequently used to estimate the abundance of marine mammals, but their accuracy is dependent upon obtaining a measure of the availability of animals for visual detection. Existing methods for characterizing availability have limitations and do not necessarily reflect true availability. Here, we present a method of using small, vessel‐launched, multi‐rotor Unoccupied Aerial Vehicles (UAVs or drones) to collect video of dolphins to characterize availability and investigate errors surrounding group size estimates. We collected over 20 h of aerial video of dive‐surfacing behaviour across 32 encounters with the Australian humpback dolphin </span><em><span>Sousa sahulensis</span></em><span> off north‐western Australia. Mean surfacing and dive periods were 7.85 sec (</span><span>se</span><span> = 0.26) and 39.27 sec (</span><span>se</span><span> = 1.31) respectively. Dolphin encounters were split into 56 focal follows of consistent group composition to which example approaches to estimating availability were applied. Non‐instantaneous availability estimates, assuming a 7-sec observation window, ranged between 0.22 and 0.88, with a mean availability of 0.46 (CV = 0.34). Availability tended to increase with increasing group size. We found a downward bias in group size estimation, with true group size typically one individual more than would have been estimated by a human observer during a standard aerial survey. The variability of availability estimates between focal follows highlights the importance of sampling across a variety of group sizes, compositions, and environmental conditions. Through data re‐sampling exercises, we explored the influence of sample size on availability estimates and their precision, with results providing an indication of target sample sizes to minimize bias in future research. We show that UAVs can provide an effective and relatively inexpensive method of characterizing dolphin availability with several advantages over existing approaches. The example estimates obtained for humpback dolphins are within the range of values obtained for other shallow‐water, small cetaceans, and will directly inform a government‐run program of aerial surveys in the region.</span></p>
Data Availability Statements in the 2020 and 2021 scientific publications of Tampere University
<p>For this dataset, scientific peer-reviewed articles by Tampere University researchers from the years 2020 and 2021 were extracted from the TUNICRIS. A random sample of 40 percent was taken from the listed 4,922 publications according to faculties and years. There were 2,085 analyzed articles, i.e. more than 42 percent of the total number. </p> <p>To find Data Availability Statements, articles were opened one by one and searched for mentions of research data and its availability. For each article, it was written down whether DAS existed and where in the article it was located. From the contents of DAS, information about data availability, location, openness and possible restrictions on use was written down. </p> <p>Dataset also includes information about the journals and publications taken from TUNICRIS. </p> <p>The prevalence of DAS and data openness were examined in relation to different variables. Tampere University faculty information has been removed from the dataset. </p> <p>Related slides: https://doi.org/10.5281/zenodo.7655892</p> <p>Related article (in Finnish): Toikko, T., & Kylmälä, K. (2023). Tutkimusdatan saatavuustiedot tieteellisissä artikkeleissa: Raportti Data Availability Statementien käytöstä Tampereen yliopistossa. <em>Informaatiotutkimus</em>, 42(1-2), 31–50. https://doi.org/10.23978/inf.126098</p>
Time availability of GrInHy2.0 HTE electrolyzer
<p>Time availability in % of the GrInHy2.0 electrolyzer system, averaged over operation months. Corrected values account for compensation of externally forced downtime.</p>
Commercial Available Compounds
<p>This dataset was collected and integrated from ZINC and eMolecules databases, bringing together the unique strengths and features of each. The resulting dataset contains a vast array of information about small molecules, including their chemical structures, properties, and activities. This integrated dataset has the potential to greatly benefit research in drug discovery, chemical biology, and other fields, as it provides a comprehensive and diverse collection of small molecules that can be used for a variety of applications.</p>
Details on the Available Plugins for Tsunami Security Scanner
<p>CSV file containing data from the Tsunami Security Scanner Plugins.</p>
Resource availability and capacity to implement multi-stranded cholera interventions in the north-east region of Nigeria
<p>Limited healthcare facility (HCF) resources and capacity to implement multi-stranded cholera interventions ('water, sanitation, and hygiene (WASH)’, ‘surveillance’, ‘case management’, and ‘community engagement’) can hinder the actualisation of the global strategic roadmap goals for cholera control, especially in settings made fragile by armed conflicts, such as the north-east region of Nigeria. Therefore, we aimed to assess HCF resource availability and capacity to implement these cholera interventions in Adamawa and Bauchi States in Nigeria, as well as assess their coordination in both states and Abuja, where national coordination of cholera is based.<br> We conducted a cross-sectional survey using a face-to-face structured questionnaire to collect data on multi-stranded cholera interventions and their respective indicators in HCFs. We generated scores to describe the resource availability of each cholera intervention and categorised them as: 0-50 (low), 51-70 (moderate), 71-90 (high), and over 90 (excellent). Further, we defined an HCF with a high capacity to implement a cholera intervention as one with a score equal to or above the average intervention score. </p>
Hummingbird blood traits track oxygen availability across space and time
<p>Predictable trait variation across environments suggests shared adaptive responses via repeated genetic evolution, phenotypic plasticity, or both. Matching of trait-environment associations at phylogenetic and individual scales implies consistency between these processes. Alternatively, mismatch implies that evolutionary divergence has changed the rules of trait-environment covariation. Here we tested whether species adaptation alters elevational variation in blood traits. We measured blood for 1,217 Andean hummingbirds of 77 species across a 4,600 m elevational gradient. Unexpectedly, elevational variation in hemoglobin concentration ([Hb]) was scale independent, suggesting that physics of gas exchange, rather than species differences, determine responses to changing oxygen pressure. However, mechanisms of [Hb] adjustment did show signals of species adaptation: Species at either low or high elevations adjusted cell size, whereas species at mid-elevations adjusted cell number. This elevational variation in red blood cell number-versus-size suggests that genetic adaptation to high altitude has changed how these traits respond to shifts in oxygen availability.</p>
Data from: Comparing winter versus summer deepwater dissolved oxygen depletion with the potential for cross-seasonal forecasting of deepwater oxygen availability
<p>Depletion of deepwater dissolved oxygen (DO) in lakes has become increasingly prevalent and severe due to many external stressors, potentially threatening human-derived ecosystem services ranging from drinking water quality to fisheries. Using year-round, high-frequency DO data from 12 dimictic lakes, we compared three measures of deepwater DO depletion during winter and summer: DO depletion rate, DO minimum, and hypoxia duration. Hypoxia (DO < 3 mg L<sup>-1</sup>) occurred in over half of the lakes and persisted an average of 83% longer in summer than in winter. While we found no difference in DO depletion rates between winter versus summer, these rates were strongly related to lake morphology in winter but water transparency and temperature in summer. Winter hypoxia duration was negatively related to summer hypoxia duration, suggesting potential utility for forecasting DO depletion in the subsequent summer. Spring mixing efficacy was strongly related to winter minimum DO saturation and hypoxia duration, and was also a strong predictor of summer minimum DO saturation and hypoxia duration. Hence, these cross-seasonal patterns suggest deepwater DO metrics can be used to forecast DO availability in subsequent seasons, modified by the relative importance of morphology, water transparency, and temperature. These findings can allow for improved, early management when DO is predicted to be critically low based on previous seasons’ DO measurements, which can work to minimize the negative consequences for water quality and fisheries health associated with severe DO depletion.</p>
Data and Code Availability Stament for Halifa-Marín et al., 2023 (Environmental Research: Climate)
<p>This file provides the R codes and datasets to reproduce the analyses made in the mentioned study, Halifa-Marín et al., 2023 (Environmental Research: Climate), and the supplementary material (figures).</p>
Can regime shifts in reproduction be explained by changing climate and food availability?
<p><span>Marine populations often show considerable variation in their productivity, including regime shifts. Of special interest are prolonged shifts to low recruitment and low abundance which occur in many fish populations despite reductions in fishing pressure. One of the possible causes for the lack of recovery has been suggested to be the Allee effect (depensation). Nonetheless, both regime shifts and the Allee effect are empirically emerging patterns but provide no explanation about the underlying mechanisms. Environmental forcing, on the other hand, is known to induce population fluctuations and </span><span>has also been suggested as one of the primary challenges for recovery.</span> <span>Yet, traditional stock-recruitment models used in fisheries management have been time-invariant and considered only density-dependence. </span><span>In the present study, we build upon recently developed Bayesian change-point models to explore the contribution of food and climate as external drivers in recruitment regime shifts, while accounting for density-dependent mechanisms (compensation and depensation). Food availability is approximated by the copepod community. Temperature is included as a climatic driver. Three demersal fish populations in the Irish Sea are studied: Atlantic cod (<em>Gadus morhua</em>), whiting (<em>Merlangius merlangus</em>), and common sole (<em>Solea solea</em>). </span><span>We demonstrate that while spawning stock biomass undoubtedly impacts recruitment, abiotic and biotic drivers can have substantial additional impacts, which can explain regime shifts in recruitment dynamics or low recruitment at low population abundances. Our results stress the fact that traditional abundance-based stock-recruitment models are not sufficient to capture variability in fish recruitment. </span></p>
Characterizing the Folding Transition State Ensembles in the Energy Landscape of an RNA Tetraloop - available data.
<p>Trajectory file and scripts necessary to generate an ELViM [Oliveira, A. B.; Yang, H.; Whitford, P. C.; Leite, V. B. P. JCTC, 2019, 15, p.6482] projection of the conformational space for the GCAA tetraloop.</p>
Data and code for "Turning tables: food availability shapes dynamic aggressive behaviour among asynchronously hatching siblings in red kites Milvus milvus"
<p><strong>Abstract</strong></p> <p>Aggression represents the backbone of dominance acquisition in several animal societies, where the decision to interact is dictated by its relative cost. Among siblings, such costs are weighted in the light of inclusive fitness, but how this translates to aggression patterns in response to changing external and internal conditions remains unclear. Using a null-model-based approach, we investigate how day-to-day changes in food provisioning affect aggression networks and food allocation in growing red kite (<em>Milvus milvus</em>) nestlings, whose dominance rank is largely dictated by age. We show that older siblings, irrespective of age, change from targeting only close-aged peers (close-competitor pattern) when food provisioning is low, to uniformly attacking all other peers (downward heuristic pattern) as food conditions improve. While food allocation was generally skewed towards the older siblings, the youngest sibling in the nest increased its probability of accessing food as more was provisioned and as downward heuristic patterns became more prominent, suggesting that different aggression patterns allow for catch-up growth after periods of low food. Our results indicate that dynamic aggression patterns within the nest modulate environmental effects on juvenile development by influencing the process of dominance acquisition and potentially impacting the fledging body condition, with far-reaching fitness consequences.</p>
Sources of prey availability data alter interpretation of outputs from prey choice null networks
<p><em>Spider surveys</em></p> <p>Data collection was described previously by Cuff, Tercel, et al., (2022). This study pertains to a subset of those data, collected between 1<sup>st</sup> May and 9<sup>th</sup> July 2018 at 19 separate locations, for which paired sticky trap and vacuum sample data were collected (described below). Briefly, money spiders (Araneae: Linyphiidae) and wolf spiders (Araneae: Lycosidae) were visually located along transects in two adjacent barley fields at Burdons Farm, Wenvoe in South Wales (51°26'24.8"N, 3°16'17.9"W) and collected from webs and the ground. Transects were randomly distributed across the entire field. Along these transects, separate 4 m<sup>2</sup> quadrats, at least 10 m apart, were searched and all observed linyphiids and lycosids were collected. Spiders were placed in 100 % ethanol using an aspirator, regularly changing meshing to limit potential cross-contamination. Linyphiids occupying webs were prioritised for collection, but ground-active spiders were also collected. Spiders were taken to Cardiff University, transferred to fresh ethanol and stored at -80 °C in 100 % ethanol until DNA extraction. Extraction, amplification and sequencing of DNA, and bioinformatic analysis is described by Cuff, Tercel, et al., (2022) and Drake et al., (2022), and is also detailed below.</p> <p><em>Extraction and high-throughput sequencing of spider gut DNA</em></p> <p>Given their prevalence in field collections, dietary analysis was carried out for the linyphiid genera <em>Erigone</em>, <em>Tenuiphantes</em>, <em>Bathyphantes</em> and <em>Microlinyphia </em>(Araneae: Linyphiidae), and the Lycosidae genus <em>Pardosa</em>. Spiders were transferred to and washed in fresh 100 % ethanol to reduce external contaminants prior to identification via morphological key (Roberts, 1993). Abdomens were removed from spiders and again transferred to and washed in fresh 100 % ethanol. DNA was extracted from the abdomens via Qiagen TissueLyser II and DNeasy Blood & Tissue Kit (Qiagen) as per the manufacturer protocol, but with an extended lysis time of 12 hours to account for the complex and branched gut system in spider abdomens (Krehenwinkel et al., 2017).</p> <p>For amplification of DNA, two primer pairs were used. BerenF-LuthienR (Cuff et al., 2021) amplified a broad range of invertebrates including spiders, and TelperionF-LaureR (Cuff et al., 2022), amplified a range of invertebrates but fewer spiders. Primers were labelled with unique 10 bp molecular identifier tags (MID-tags) so that each individual had a unique pairing of forward and reverse tags for identification of each spider post-sequencing. PCR reactions of 25 µl contained 12.5 µl Qiagen PCR Multiplex kit, 0.2 µmol (2.5 µl of 2 µM) of each primer and 5 µl template DNA. Reactions were carried out in the same thermocycler, optimised via temperature gradient, with an initial 15 minutes at 95 °C, 35 cycles of 95 °C for 30 seconds, the primer-specific annealing temperature for 90 seconds and 72 °C for 90 seconds, respectively, followed by a final extension at 72 °C for 10 minutes. BerenF-LuthienR and TelperionF-LaureR used annealing temperatures of 52 °C and 42 °C, respectively.</p> <p>Within each PCR 96-well plate, 12 negative controls (extraction and PCR), 2 blank controls and 2 positive controls were included (i.e. 80 samples per plate), based on Taberlet <em>et al. </em>(2018). Positive controls were mixtures of invertebrate DNA comprised of non-native Asiatic species in four different proportions and blanks were empty wells within each plate to identify tag-jumping into unused MID-tag combinations. PCR negative controls were DNase-free water treated identically to DNA samples. A negative control was present for each MID-tag to identify any contamination of primers. All PCR products were visualised in a 2 % agarose gel with SYBRSafe (Thermo Fisher Scientific, Paisley, UK) and placed in categories based on their relative brightness. The concentration of these brightness categories was quantified via Qubit dsDNA High-sensitivity Assay Kits (Thermo Fisher Scientific, Waltham, MA, USA) with at least three representatives of each category per plate. The PCR products were then proportionally pooled according to these concentrations. Each pool was cleaned via SPRIselect beads (Beckman Coulter, Brea, USA), with a left-side size selection using a 1:1 ratio (retaining ~300-1000 bp fragments). The concentration of the pooled DNA was then determined via Qubit dsDNA High-sensitivity Assay Kits and pooled together into one library per primer pair. Library preparation for Illumina sequencing was carried out on the cleaned libraries via NEXTflex Rapid DNA-Seq Kit (Bioo Scientific, Austin, USA) and samples were sequenced on an Illumina MiSeq via a V3 chip with 300-bp paired-end reads (expected capacity ≤25,000,000 reads). Bioinformatic analysis followed Drake et al. (2022).</p> <p><em>Bioinformatic analysis</em></p> <p>The Illumina run generated 11,165,405 and 10,959,010 reads for BerenF-LuthienR and TelperionF-LaureR, respectively, which were quality-checked and paired via FastP (Chen et al., 2018) to retain only sequences of at least 200 bp with a quality threshold of 33, resulting in 10,561,874 and 9,355,112 paired reads. The paired reads were demultiplexed and assigned to their respective spider sample according to their MID-tags via the “trim.seqs” command in Mothur v1.39.5 (Schloss et al., 2009), leaving 7,854,610 and 7,437,929 reads with exact matches to the primer and MID-tags.</p> <p>Replicates were removed, and denoising and clustering to zero-radius operational taxonomic units (ZOTUs; clustered without % identity to avoid multiple species represented within a single operational taxonomic unit (OTU)) completed via Unoise3 in Usearch11 (Edgar, 2010). The resultant sequences were assigned a taxonomic identity from GenBank via BLASTn v2.7.1 (Camacho et al., 2009) using a 97 % identity threshold (Alberdi et al., 2017). The BLAST output was analysed in MEGAN v6.15.2 (Huson et al., 2016). Where the top BLAST hit, determined by lowest e-value, was resolved at a higher taxonomic level than species-level, the results were checked; where possibly erroneous entries were preventing species-level assignment (e.g., poorly resolved identifications on GenBank), finer resolution was assigned based on the next-closest match. Where ZOTUs were assigned the same taxon, these were aggregated.</p> <p>Data clean-up used the optimal minimum sequence copy thresholds identified by Drake et al. (2022). The maximum value for a ZOTU present in blank or negative controls was identified and subtracted from all read counts for that ZOTU to remove background contaminants. Simultaneously, known lab contaminants (e.g., German cockroach <em>Blattella germanica</em>), artefacts and errors of the sequencing process, unexpected reads in positive controls and positive control taxon reads in dietary samples were identified. These were calculated as a percentage of their respective sample’s read count and any read counts lower than the highest of these percentages for their respective sample were removed to eliminate additional instances of contamination. These thresholds were defined as 0.38 % and 0.39 % for BerenF-LuthienR and TelperionF-LaureR, respectively. The data from the two libraries (i.e., from each primer pair) were then aggregated together by sample and aggregated again by taxon. Non-target taxa (e.g., fungi) and instances in which predator DNA was amplified (i.e., ZOTUs with high read counts matching the individual’s morphological identity) were removed. </p> <p>The resultant sequencing read counts were converted into relative proportions (all values made to sum to one within each sample) and a mean value across the two primer pairs retained for each taxon within each sample. Relative read abundances were converted to presence-absence data of each detected prey taxon in each individual spider, but relative read abundance data were also retained for separate analyses to compare experimental outcomes between treatments.</p> <p><em>Invertebrate surveys</em></p> <p>To estimate prey availability using sticky traps, we placed one white dry 100 mm x 125 mm trap (Oecos) in the 4 m<sup>2</sup> quadrat centred at the position where the spider was captured. The trap was suspended with wire approximately 25 mm above the ground to catch falling, crawling and flying invertebrates, and left in place for 72 hours. Invertebrates were identified on the traps under a stereomicroscope. To estimate prey availability using suction sampling, ground and crop stems were sampled using a ‘G-vac’ for approximately 30 seconds at each location. The collected material was emptied into a bag, any organisms immediately killed with ethyl-acetate and material frozen for storage before sorting into 70 % ethanol in the lab. All invertebrates were identified to family level to match the resolution of the least resolved of the metabarcoding-derived trophic interaction data, and due to difficulties associated with identification to finer taxonomic resolution for many taxa. Exceptions included springtails of the superfamily Sminthuroidea (Sminthuridae and Bourletiellidae were often indistinguishable following suction sampling and preservation due to the fine features necessary to distinguish them) which were left at super-family, mites (many of which were immature or in poor condition) which were identified to order level, and wasps of the superfamily Ichneumonoidea which were identified no further due to obscurity of wing venation due to damage following suction sampling.</p> <p><em>Statistical Analysis</em></p> <p>All analyses were conducted in R v4.0.3 (R Core Team, 2021) and carried out on invertebrate data at the family or superfamily level. Alongside the dietary data derived from metabarcoding, and prey availability as determined directly by suction sampling (abundance) and sticky trapping (activity density), three additional datasets were generated where two were designed to combine data from the two trapping methods. The first approach simply set all invertebrate taxa detected in the field to have equal abundance, to provide a baseline against which to assess the effects of different prey abundance estimates. When generating the two combined data sets, it was apparent that simply adding them together would underrepresent one of the datasets as abundance and activity density are measured in different units. Therefore, a ‘proportional combined’ dataset was generated by converting counts to relative proportions of each sample (to equally weight the two methods), which were then combined by summing proportions between the two methods for each sample, multiplied by the total count of individuals across both methods for each sample (to create realistic abundance values), and then rounded to the nearest integer (to return count data). In addition, a ‘frequency of occurrence (FOO) combined’ dataset was generated by converting counts to binary presence-absence values of each sample, which were then summed between the two methods for each sample. To assess the diversity represented by the two sampling methods and their combinations, and the completeness of those datasets, coverage-based rarefaction and extrapolation were carried out, and Hill diversity calculated (Chao et al., 2014; Roswell et al., 2021) using the ‘iNEXT’ package with families represented by frequency-of-occurrence across samples (Chao et al., 2014; Hsieh et al., 2016).</p> <p>The remaining analyses were performed using both presence-absence and relative read abundance dietary data separately to show how differences in the treatment of the observed data are reflected in the outcomes of the analyses. Figures and outputs given in the main text relate to the presence-absence data, while relative read abundance figures and outputs are presented in the Supplementary Information. Prey preferences of spiders were analysed using network-based null models in the ‘econullnetr’ package (Vaughan et al., 2018) with the ‘generate_null_net’ function. Econullnetr generates null models based on prey availability to predict how consumers would forage if based on the availability of resources alone. These null models are then compared against the observed interactions of consumers (e.g., interactions of spiders with their prey based on dietary metabarcoding) to ascertain the extent to which resource consumption deviated from random. In five separate null models, prey availability was represented separately by the datasets described above: abundance (suction sampling), activity density (sticky trapping), proportional combined, FOO combined and equal prey abundance.</p> <p>To compare effect sizes between null models for each resource taxon, mean prey preference standardised effect size (SES) values were calculated from the individual spiders per model. The SES values were plotted and joined between taxa to visualise paired differences using ‘ggplot’ (Wickham, 2016). Null model-predicted trophic interactions were generated via an econullnetr null model with 999 simulations with outputs extended to allow the comparison of the null interactions for individual consumers (generate_null_net_indiv; Cuff, Kitson, et al., 2023). A visualisation of the per-individual differences in null model and observed data was generated via non-metric multi-dimensional scaling (NMDS) using the ‘metaMDS’ function in the ‘vegan’ package (Oksanen et al., 2016) in two dimensions and 9999 simulations, with Euclidean distance. Centroid coordinates for each null model and the observed data were extracted and pairwise distances calculated between model centroids:</p> <p>The ‘observed’ network (i.e., the network determined solely by dietary data, not necessarily the objectively ‘true’ network) and each null network were visualised with the associated prey choice effect sizes as a bipartite network using ‘ggnetwork’ (Briatte, 2021; Wickham, 2016) via an ‘igraph’ object (Csardi & Nepusz, 2006). The degree of each prey node, weighted nestedness and linkage density were generated using the ‘bipartite’ package (Dormann et al., 2008) for each network and compared visually via ggplot2.</p>
LiverHccSeg: A Publicly Available Multiphasic MRI Dataset with Liver and HCC Tumor Segmentations and Inter-Rater Agreement Analysis
<p>Please <strong>cite our data paper </strong>published in "Data in Brief": https://www.sciencedirect.com/science/article/pii/S2352340923007473</p><p> </p><p><strong>Background</strong><br>Liver cancer ranks as the third leading cause of cancer-related mortality worldwide [1] and alarmingly, both the incidence and mortality rates of liver cancer are increasing [2; 3]. Among the various types of primary liver cancer, hepatocellular carcinoma (HCC) stands out as the most prevalent, accounting for approximately 70-85% of liver cancer cases [4]. Leveraging the advantages of magnetic resonance (MR) imaging, HCC can be reliably detected and diagnosed without the requirement of an invasive biopsy [5]. MR imaging offers high tissue contrast, which can be further enhanced through contrast-enhanced multiphasic magnetic resonance imaging (mpMRI) techniques. This enables accurate identification and non-invasive diagnosis of HCC [6].</p><p> </p><p><strong>Objective</strong><br>Precise segmentation of the liver plays a crucial role in volumetry assessment and serves as a vital pre-processing step for subsequent tumor detection algorithms [7]. However, accurate liver segmentation can be particularly challenging in patients with cancer-related tissue alterations and deformations in shape [8]. Accurate HCC tumor segmentation is essential for the extraction of quantitative imaging biomarkers such as radiomics and can be used for studies on treatment response assessment and prognosis evaluation and provides critical information about the tumor biology. In order to enhance the reproducibility of liver and tumor segmentation, automated methods utilizing image analysis techniques and machine learning have been developed. These methods have demonstrated promising results [7; 8]; however, most algorithms were tested only on small internal test sets and therefore do not guarantee generalizable and consistent performance on external data.</p><p>Publicly available datasets allow for fair and objective comparisons between different algorithms, techniques, or approaches. Researchers can evaluate the strengths and weaknesses of their methods in relation to existing solutions and establish benchmarks for performance evaluation. In addition to providing a benchmark with this dataset, we also assess the inter-rater variability between two different sets of tumor segmentations. This analysis serves as a measure of reproducibility for human segmentations, highlighting the consistency or variability that may exist among different human raters. Understanding the reproducibility of human segmentations is essential in assessing the reliability of manual annotations and establishing a baseline for algorithm performance comparison. By introducing LiverHccSeg, we aim to fill the gap of lacking publicly available mpMRI HCC datasets and offer researchers and developers a valuable resource for algorithmic evaluation on external data and imaging biomarker analyzes.</p><p> </p><p><strong>Materials and Methods</strong></p><p><i><strong>Inclusion of Patients</strong></i><br>All available scans from The Cancer Genome Atlas Liver Hepatocellular Carcinoma Collection (TCGA-LIHC) (<a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=6885436">https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=6885436</a>) were downloaded [9]. One multiphasic MRI scan (pre and triphasic post contrast) per patient was included. Patients who did not exhibit a tumor or residual tumor were excluded from the tumor segmentation dataset; however, they were included in the liver segmentation dataset.</p><p> </p><p><i><strong>MR Imaging Data</strong></i><br>Subsequently, all imaging data was converted to the Neuroimaging Informatics Technology Initiative (NIfTI) format with the dcm2nii (v2.1.53) package [10] and available header information was extracted using the pydicom (v.2.1.2) package [11]. Multiparametric MR sequences were labeled with a consistent syntax ('pre', 'art', 'pv', 'del', for the pre-contrast, arterial, portal-venous and delayed contrast phases, respectively). All images were already de-identified by the TCIA website. Images were acquired between the years 1993 and 2007 on Philips and Siemens scanners with field strengths of 1.5 and 3 Tesla. Full details of the imaging parameters can be found in Table 5. Briefly, the median repetition time (TR) and median echo time (TE) were 365.8 ms and 26.4 ms, respectively. The median slice thickness was 9.5 mm, the median bandwidth 536.9 Hz.</p><p> </p><p><i><strong>Scientific Reading</strong></i><br>After conversion, all images were read in a scientific reading by two board-certified abdominal radiologists (S.A. and S.H with 9 and 10 years of experience, respectively). Any disagreement between the two raters was discussed in a consensus meeting. All HCC lesions were classified according to LI-RADS criteria [6].</p><p> </p><p><i><strong>Image Registration</strong></i><br>The co-registration of pre-contrast, portal-venous, and delayed-phase images with arterial phase images was performed using the software BioImage Suite (v3.5) [12]. A non-rigid intensity-based registration approach was applied, employing a parameterized free-form deformation (FFD) with 3D B-splines [13]. The optimal FFD transformation was estimated by maximizing the normalized mutual information similarity metric [14] through gradient descent optimization. To enhance the optimization process, a multi-resolution image pyramid with three levels was utilized. The final B-spline control point spacing was set to 80 mm. The estimated transformation was then employed to warp the moving images (pre-contrast, portal-venous, and delayed-phase) into the reference image space, specifically the arterial phase image.</p><p> </p><p><i><strong>Liver and Tumor Segmentation and Statistical Analysis</strong></i><br>All livers and tumors were manually segmented under the supervision of two board-certified abdominal radiologists using the software 3D Slicer (v4.10.2) [15]. To compare the segmentation agreement between the two sets of liver and tumor segmentations, we calculated segmentation metrics using the Python package seg-metrics (v1.0.0) [16]. All segmentation metrics and statistics were calculated in Python (v3.7).</p><p> </p><p><strong>Data description</strong><br>The data that appears in this article include:</p><ol><li>dicoms.zip: This zip file contains all the raw MR images from The Cancer Genome Atlas Liver Hepatocellular Carcinoma Collection (TCGA-LIHC) [1] in the Digital Imaging and Communications in Medicine (DICOM) format used for the curation of this dataset. The data is structured as Patient-ID/DATE/SEQUENCE where Patient-ID is the unique unidentified patient ID, DATE is the date of the image acquisition, and SEQUENCE is the name of the MR sequence.<br> </li><li>LiverHccSeg_MetaData.xlsx: This spreadsheet contains all the metadata from the DICOM headers along with the data from the scientific image readings.<br> </li><li>nifti_and_segms.zip: This zip file contains all MR images along with the liver and tumor segmentations in the Neuroimaging Informatics Technology Initiative (NIfTI) format.<br>The data is structured as Patient-ID/DATE/SEQUENCE where Patient-ID is the unique anonymized patient identifier, DATE is the date of the image acquisition, and SEQUENCE is the name of the MRI sequence or segmentation image.<br><br>The NIfTI files are named as follows:<br><strong>pre.nii.gz</strong> : Pre-contrast T1-weighted MRI<br><strong>art.nii.gz</strong>: Arterial-phase T1-weighted MRI<br><strong>pv.nii.gz</strong>: Portal-venous-phase T1-weighted MRI<br><strong>del.nii.gz</strong>: Delayed-phase T1-weighted MRI<br><strong>art_pre.nii.gz</strong>: Pre-contrast T1-weighted MRI registered to the corresponding arterial-phase T1-weighted image<br><strong>art_pv.nii.gz</strong>: Portal-venous-phase T1-weighted MRI registered to the corresponding arterial-phase T1-weighted MRI<br><strong>art_del.nii.gz</strong>: Delayed-phase T1-weighted MRI registered to the corresponding arterial-phase T1-weighted MRI<br><br>The corresponding manual segmentations are named after the rater and the type of segmentation and follow the format 'RATER_ROI.nii.gz' where RATER denotes the human rater and ROI denotes the region of interest that was segmented, for example, '<strong>rater1_liver.nii.gz</strong>', '<strong>rater2_liver.nii.gz</strong>', '<strong>rater1_tumor1.nii.gz</strong>', and '<strong>rater2_tumor1.nii.gz</strong>'. For tumor segmentations, an integer indicates the tumor identification number for different tumor ROIs, for example, 'rater1_tumor1.nii.gz' and 'rater2_tumor1.nii.gz'. The segmentations can be used for the arterial phase NIfTI file as well as the corresponding co-registered pre-contrast (art_pre.nii.gz), portal-venous (art_pv.nii.gz), and delayed-phase (art_del.nii.gz) images.<br> </li><li>segm_metrics.xlsx: This spreadsheet summarizes the segmentation agreement between the two sets of liver and tumor segmentations by the two board-certified abdominal radiologists.</li></ol><p> </p><p><strong>References</strong></p><p>1 Sung H, Ferlay J, Siegel RL et al (2021) Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin 71:209-249</p><p>2 Siegel RL, Miller KD, Jemal A (2019) Cancer statistics, 2019. CA Cancer J Clin 69:7-34</p><p>3 White DL, Thrift AP, Kanwal F, Davila J, El-Serag HB (2017) Incidence of Hepatocellular Carcinoma in All 50 United States, From 2000 Through 2012. Gastroenterology 152:812-820.e815</p><p>4 Perz JF, Armstrong GL, Farrington LA, Hutin YJ, Bell BP (2006) The contributions of hepatitis B virus and hepatitis C virus infections to cirrhosis and primary liver cancer worldwide. J Hepatol 45:529-538</p><p>5 Hamer OW, Schlottmann K, Sirlin CB, Feuerbach S (2007) Technology insight: advances in liver imaging. Nat Clin Pract Gastroenterol Hepatol 4:215-228</p><p>6 Chernyak V, Fowler KJ, Kamaya A et al (2018) Liver Imaging Reporting and Data System (LI-RADS) Version 2018: Imaging of Hepatocellular Carcinoma in At-Risk Patients. Radiology 289:816-830</p><p>7 Bousabarah K, Letzen B, Tefera J et al (2020) Automated detection and delineation of hepatocellular carcinoma on multiphasic contrast-enhanced MRI using deep learning. Abdom Radiol. 10.1007/s00261-020-02604-5</p><p>8 Gross M, Spektor M, Jaffe A et al (2021) Improved performance and consistency of deep learning 3D liver segmentation with heterogeneous cancer stages in magnetic resonance imaging. PLoS One 16:e0260630</p><p>9 Erickson BJ, Kirk S, Lee Y et al (2016) Radiology Data from The Cancer Genome Atlas Liver Hepatocellular Carcinoma [TCGA-LIHC] collection. The Cancer Imaging Archive. 10.7937/K9/TCIA.2016.IMMQW8UQ</p><p>10 dcm2nii DICOM to NIfTI converter. <a href="https://github.com/rordenlab/dcm2niix">https://github.com/rordenlab/dcm2niix</a> Accessed: 2021-12-07.</p><p>11 Mason D, scaramallion;, rhaxton; et al (2020) pydicom/pydicom: pydicom 2.1.2, v2.1.2. Zenodo</p><p>12 X. Papademetris MJ, N. Rajeevan, H. Okuda, R.T. Constable, L.H Staib BioImage Suite: An integrated medical image analysis suite, Section of Bioimaging Sciences, Dept. of Diagnostic Radiology, Yale School of Medicine. <a href="http://www.bioimagesuite.org">http://www.bioimagesuite.org</a>.</p><p>13 Rueckert D, Sonoda LI, Hayes C, Hill DLG, Leach MO, Hawkes DJ (1999) Nonrigid Registration Using Free-Form Deformations: Application to Breast MR Images. IEEE Trans Med Imaging 18:712–721</p><p>14 Studholme C, Hill DL, Hawkes DJ (1999) An overlap invariant entropy measure of 3D medical image alignment. Pattern Recognition 32:71-86</p><p>15 Fedorov A., Beichel R., Kalpathy-Cramer J. et al (2012) 3D Slicer as an Image Computing Platform for the Quantitative Imaging Network. Magn Reson Imaging 30:1323-1341</p><p>16 Ordgod (2020) Ordgod/segmentation_metrics: seg-metrics, v1.0.0. Zenodo</p><p> </p><p>- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -</p><p><strong>ChangeLog</strong></p><p><strong>Version 1.1:</strong> Fixed incorrect liver segmentation mask.</p><p> </p><p> </p><p> </p><p> </p>
Dataset: Preliminary analysis of open data pertaining to the services available through the Health Insurance Institute of Slovenia and provided by family medicine
<p>BACKGROUND: The Health Insurance Institute of Slovenia (ZZZS) began publishing service-related data in May 2023, following a directive from the Ministry of Health (MoH). The ZZZS website provides easily accessible information about the services provided by individual doctors, including their names. The user is provided relevant information about the doctor's employer, including whether it is a public or private institution. The data provided is useful for studying the public system's operations and identifying any errors or anomalies. </p> <p>METHODS: The data for services provided in May 2023 was downloaded and analysed. The published data were cross-referenced using the provider's RIZDDZ number with the daily updated data on ambulatory workload from June 9, 2023, published by ZZZS. The data mentioned earlier were found to be inaccurate and were improved using alerts from the zdravniki.sledilnik.org portal. Therefore, they currently provide an accurate representation of the current situation. The total number of services provided by each provider in a given month was determined by adding up the individual services and then assigning them to the corresponding provider. </p> <p>RESULTS: A pivot table was created to identify 307 unique operators, with 15 operators not appearing in both lists. There are 66 public providers, which make up about 72% of the contractual programme in the public system. There are 241 private providers, accounting for about 28% of the contractual programme. In May 2023, public providers accounted for 69% (n=646,236) of services in the family medicine system, while private providers contributed 31% (n=291,660). The total number of services provided by public and private providers was 937,896. Three linear correlations were analysed. The initial analysis of the entire sample yielded a high R-squared value of .998 (adjusted R-squared value of .996) and a significant level below 0.001. The second analysis of the data from private providers showed a high R Squared value of .904 (Adjusted R Squared = .886), indicating a strong correlation between the variables. Furthermore, the significance level was < 0.001, providing additional support for the statistical significance of the results. The third analysis used data from public providers and showed a strong level of explanatory power, with a R Squared value of 1.000 (Adjusted R Squared = 1.000). Furthermore, the statistical significance of the findings was established with a p-value < 0.001. </p> <p>CONCLUSION: Our analysis shows a strong linear correlation between contract size of the program signed and number services rendered by family medicine providers. A stronger linear correlation is observed among providers in the public system compared to those in the private system. Our study found that private providers generally offer more services than public providers. However, it is important to acknowledge that the evaluation framework for assessing services may have inherent flaws when examining the data. Prescribing a prescription and resuscitating a patient are both assigned a rating of one service. It is crucial to closely monitor trends and identify comparable databases for pairing at the secondary and tertiary levels.</p>
Environmental Performance Assessment of a Novel Process Concept for Propanol Production from Widely Available and Wasted Methane Sources
<p>Dataset for associated publication.</p>
Data from: The impacts of climate change, energy policy, and traditional ecological practices on future firewood availability for Diné (Navajo) People
<p>These data are part of a data portal that accompanies the special issue 'Climate change adaptation needs a science of culture,' published in Philosophical Transactions of the Royal Society B in 2023. To access the data portal, please visit <a href="https://doi.org/10.5061/dryad.bnzs7h4h4"><strong>https://doi.org/10.5061/dryad.bnzs7h4h4</strong></a>.</p> <p>The files consist of the code of an agent-based model (ABM) in a NetLogo, detailed documentation of the ABM in a standard format, and a table of data exported from the simulation experiment reported on in the paper. By downloading the Netlogo file, one could not only rerun the experiment we report on and recreate the data table but toggle parameters or edit the model to explore other dynamics.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.