Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,609
datasets available to search
ShareScore release 0.9.0
Dataset results
2,609 results for “web”
Unveiling Web Fingerprinting in the Wild Via Code Mining and Machine Learning
<p>Dataset of Javascripts used for training and testing the fingerprinting algorithms described in </p> <p>Rizzo, Valentino, Stefano Traverso, and Marco Mellia. "Unveiling Web Fingerprinting in the Wild Via Code Mining and Machine Learning." <em>Proceedings on Privacy Enhancing Technologies</em> 2021.1 (2021): 43-63.</p>
Results of the Web-Delphi process to INAMI stakeholders (three rounds)
<p>IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 3 (Testing the framework with empirical applications), Results of the Web-Delphi process to INAMI stakeholders (three rounds), about the views of stakeholders regarding “The relevance of the following value aspects for the evaluation of new medicines in this disease context is” (2020)</p> <p>For details on the Web-Delphi process, see: IMPACT HTA, Work Package 7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 3, Deliverable 7.3 (Testing the framework with empirical applications), Testing the IMPACT-HTA Value Framework in collaboration with HTA agencies: Case studies on Non-Small Cell Lung Cancer and Spinal Muscular Atrophy (2021), Aris Angelis (LSHTM, LSE), Mónica Oliveira (IST), Teresa Rodrigues (IST), Liliana Freitas (IST), Carlos Bana e Costa (IST), Panos Kanavos (LSE)</p>
IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2 (Multi-criteria evaluation framework), Results of the 2nd Web-Delphi process to HTA stakeholders, organized in a single panel
<p>IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2 (Multi-criteria evaluation framework), Results of the 2<sup>nd</sup> Web-Delphi process to HTA stakeholders, organized in a single panel (all stakeholder groups in a single panel, 2 rounds), about the views of stakeholders regarding “This aspect should be considered in the evaluation of new medicines on a common basis” (2019)</p> <p>For details on the Web-Delphi process, see: IMPACT HTA, Work Package 7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2, Deliverable 7.2 (Multi-criteria evaluation framework), Advancing knowledge and MCDA tools to assist HTA agencies in evaluating medicines on a common basis (2021) Oliveira, M.D. (IST), Panos Kanavos (LSE), Bana e Costa, C. (IST)</p>
IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2 (Multi-criteria evaluation framework), Results of the 1st Web-Delphi process to HTA stakeholders, organized into 6 separate parallel panels
<p>IMPACT HTA, WP7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2 (Multi-criteria evaluation framework), Results of the 1<sup>st</sup> Web-Delphi process to HTA stakeholders, organized into 6 separate parallel panels (one panel per stakeholder group, 2 rounds), about the views of stakeholders regarding “This aspect should be considered in the evaluation of new medicines on a common basis” (2019)</p> <p>For details on the Web-Delphi process, see: IMPACT HTA, Work Package 7 (Methodological tools using multi-criteria value methods for HTA decision-making), Task 2, Deliverable 7.2 (Multi-criteria evaluation framework), Advancing knowledge and MCDA tools to assist HTA agencies in evaluating medicines on a common basis (2021) Oliveira, M.D. (IST), Panos Kanavos (LSE), Bana e Costa, C. (IST)</p>
Extended Wikipedia Web Traffic Daily Dataset (without Missing Values)
<p>This dataset contains 145063 time series representing the number of hits or web traffic for a set of Wikipedia pages from 2015-07-01 to 2022-06-30. This is an extended version of the dataset that was used in the Kaggle Wikipedia Web Traffic forecasting competition. For consistency, the same Wikipedia pages that were used in the competition have been used in this dataset as well. The colons (:) in article names have been replaced by dashes (-) to make the .tsf file readable using our <a href="https://github.com/rakshitha123/TSForecasting/tree/master/utils">data loaders</a>.</p> <p>The original dataset contains missing values. They have been simply replaced by zeros.</p> <p>The data were downloaded from the <a href="https://wikimedia.org/api/rest_v1/#/Pageviews%20data/get_metrics_pageviews_per_article__project___access___agent___article___granularity___start___end_">Wikimedia REST API</a>. According to the conditions of the API, this dataset is licensed under <a href="https://creativecommons.org/licenses/by-sa/3.0/">CC-BY-SA 3.0</a> and <a href="https://www.gnu.org/licenses/fdl-1.3.html">GFDL</a> licenses.</p>
Extended Wikipedia Web Traffic Daily Dataset (with Missing Values)
<p>This dataset contains 145063 time series representing the number of hits or web traffic for a set of Wikipedia pages from 2015-07-01 to 2022-06-30. This is an extended version of the dataset that was used in the Kaggle Wikipedia Web Traffic forecasting competition. For consistency, the same Wikipedia pages that were used in the competition have been used in this dataset as well. The colons (:) in article names have been replaced by dashes (-) to make the .tsf file readable using our <a href="https://github.com/rakshitha123/TSForecasting/tree/master/utils">data loaders</a>.</p> <p><br> The data were downloaded from the <a href="https://wikimedia.org/api/rest_v1/#/Pageviews%20data/get_metrics_pageviews_per_article__project___access___agent___article___granularity___start___end_">Wikimedia REST API</a>. According to the conditions of the API, this dataset is licensed under <a href="https://creativecommons.org/licenses/by-sa/3.0/">CC-BY-SA 3.0</a> and <a href="https://www.gnu.org/licenses/fdl-1.3.html">GFDL</a> licenses.</p>
zdravniki.sledilnik.org web page statistics
<p>The data uploaded here were used in a conference (MI'22) submission: Kaj se skriva v ozadju zdravniki.sledilnik.org created by the same authors.</p> <p>Data is automatically created by Google and Github actions. </p> <p>Github repositories linked to the project are:</p> <ul> <li><a href="https://github.com/sledilnik/zdravniki">sledilnik/zdravniki</a></li> <li><a href="https://github.com/sledilnik/zdravniki-data">sledilnik/zdravniki-data</a></li> </ul> <p>You can read more about the project on <a href="https://zdravniki.sledilnik.org/en/">zdravniki.sledilnik.org</a>.</p> <p>Authors would like to thank all of the <a href="https://covid-19.sledilnik.org/en/stats">COVID-19 tracker</a> project members for making this possible.</p>
Energy-Saving Strategies for Mobile Web Apps and their Measurement: Results from a Decade of Research - Dataset
<p>In 2022, over half of the web traffic was accessed through mobile devices. By reducing the energy consumption of mobile web apps, we can not only extend the battery life of our devices, but also make a significant contribution to energy conservation efforts. For example, if we could save only 5% of the energy used by web apps, we estimate that it would be enough to shut down one of the nuclear reactors in Fukushima. This paper presents a comprehensive overview of energy-saving experiments and related approaches for mobile web apps, relevant for researchers and practitioners. To achieve this objective, we conducted a systematic literature review and identified 44 primary studies for inclusion. Through the mapping and analysis of scientific papers, this work contributes: (1) an overview of the energy-draining aspects of mobile web apps, (2) a comprehensive description of the methodology used for the energy-saving experiments, and (3) a categorization and synthesis of various energy-saving approaches.</p>
Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis
<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>
Temporal variation in spider trophic interactions is explained by the influence of weather on prey communities, web building and prey choice
<p>Materials and Methods</p> <p><em>Fieldwork</em><em> and sample processing</em></p> <p>Field collection and sample processing has been described previously by Cuff, Tercel, et al., (2022), but is briefly described in Supplementary Information 1. In short, money spiders (Araneae: Linyphiidae) and wolf spiders (Araneae: Lycosidae) were collected from occupied webs and the ground in barley fields between April and September 2018. Linyphiids occupying webs (n = 78) were prioritised for collection, but ground-active linyphiid and lycosid spiders were also collected. For each linyphiid taken from a web, the height of the web from the ground (mm) and its approximate dimensions were recorded, the latter calculated as approximate web area (mm<sup>2</sup>). To obtain data on local prey density, ground and crop stems were suction sampled using a ‘G-vac’ for approximately 30 seconds at each 4 m<sup>2</sup> quadrat from which spiders were collected. Extraction, amplification and sequencing of DNA, and bioinformatic analysis is described by Cuff, Tercel, et al. (2022) and Drake et al. (2022), and is also detailed in Supplementary Information 2. Amplification was carried out using two complementary PCR primer pairs: one targeting invertebrates generally, and one intended to exclude amplification of spider DNA to reduce the prevalence of ‘host’ reads in the data output (Cuff et al. 2023). Amplicons were sequenced via Illumina MiSeq V3 with 2x300 bp paired-end reads. The resultant sequencing read counts were converted to presence-absence data of each detected prey taxon in each individual spider. Given the prevalence of sequencing reads associated with each spider analysed and the impossibility of disentangling these from detections of intraspecific predation (i.e., cannibalism), all such reads were removed (Cuff et al. 2023), although intrageneric and intrafamilial predation were still detected.</p> <p> </p> <p><em>Weather data</em></p> <p>Weather data were taken from publicly available reports from the Cardiff Airport weather station (6.6 km from the study site) via “Wunderground” (Wunderground, 2020), to represent local weather conditions. This does not necessarily reflect smaller-scale effects (e.g., microclimate-scale; Bell, 2014; Holtzer et al., 1988), but the timescale of detection for dietary metabarcoding reduces the value of that resolution given that spiders may forage across multiple microclimates. We collated data from 1<sup>st</sup> January 2018 to 17<sup>th</sup> September 2018 (the last field collection). Weather data were also separately extracted for the week preceding each of the two 2017 collection dates (3<sup>rd</sup> to 9<sup>th</sup> August and 29<sup>th</sup> August to 4<sup>th</sup> September 2017). Specifically, daily average temperatures (°C), daily average dew point (°C), maximum daily wind speed (km h<sup>-1</sup>), daily sea level pressure (hPa) and day length (min; sunrise to sunset) were recorded. Precipitation data were downloaded via the UK Met Office Hadley Centre Observation Data (UK Met Office, 2020) as regional precipitation (mm) for South West England & Wales. Weather data were converted to mean values for seven days preceding the collection of spider samples to correspond with the longevity of DNA in the guts of spiders (Greenstone et al., 2014).</p> <p> </p> <p><em>Statistical Analysis</em></p> <p>All analyses were conducted in R v4.0.3 (R Core Team, 2020). To assess how weather affects spider trophic interactions over time, we analysed dietary changes across weather gradients using multivariate models. To identify whether this was likely to be driven by changes in prey abundance, we assessed the corresponding changes in the prey communities and then used null models to ascertain whether spiders were responding to prey abundance changes through prey choice. Given the dependence of linyphiid spiders on webs for foraging, we also compared web height and area over weather gradients to assess whether this may be a component of adaptive foraging. To assess the inter-annual consistency of prey choices in response to weather conditions, we also assessed whether prey preference data could be used to improve the predictive power of prey choice models. For this, we generated null models for 2017 data with prey abundance weighted by prey preferences estimated with the 2018 data. This allowed us to assess the consistency of prey choice under similar conditions, but also provides insight as to whether this framework can be used to predict predator responses to diverse prey communities under dynamic conditions. We detail the specific stages of this analytical framework in the below sections.</p> <p> </p> <p><em>Sampling completeness and diversity assessment</em></p> <p>To assess the diversity represented by the dietary analysis and the invertebrate community sampling, and the completeness of those datasets, coverage-based rarefaction and extrapolation were carried out, and Hill diversity calculated (Chao et al., 2014; Roswell, Dushoff, & Winfree, 2021). This was performed using the ‘iNEXT’ package with species represented by frequency-of-occurrence across samples (Chao et al., 2014; Hsieh et al., 2016; Figures S4 & S6).</p> <p> </p> <p><em>Relationships between weather, spider trophic interactions and prey community composition</em></p> <p>Prey species that occurred in only one spider individual were removed before further analyses to prevent outliers skewing the results. Spider trophic interactions were related to temporal and weather variables in multivariate generalized linear models (MGLMs) with a binomial error family (Wang, Naumann, Wright, & Warton, 2012). Trophic interactions were related to temporal variables and their pairwise interactions (including spider genus to account for any confounding effect), weather variables and their pairwise interactions, and weather variables and their interactions with spider genus and time (to account for any confounding effects) in three separate MGLMs. These variables were separated into different models (Temporal model, Weather interaction model and Confounding effects model) to improve model fit and reduce singularity. Invertebrate communities from suction sampling were related to temporal and weather variables in identically structured MGLMs (excluding the spider genus variable) with a Poisson error family.</p> <p>All MGLMs were fitted using the ‘manyglm’ function in the ‘mvabund’ package (Wang et al., 2012). ‘Temporal model’ independent variables were calendar day (<em>day</em>), mean day length in minutes for the preceding week (<em>day length</em>), spider genus (for dietary models only, to ascertain any effect of spider taxonomic differences on dietary differences over time and day lengths) and all two-way interactions between these variables. ‘Weather interaction model’ independent variables were mean temperature, precipitation, dewpoint, wind speed and pressure for the preceding week, and pairwise interactions between weather variables. ‘Confounding effects model’ independent variables were day (to investigate the interaction between time and weather), spider genus (for dietary models only, to ascertain any effect of spider taxonomic differences on dietary differences over time and day lengths), mean temperature, precipitation, dewpoint, wind speed and pressure for the preceding week, and two-way interactions of each weather variable with day and genus.</p> <p>Trophic interaction and community differences were visualised by non-metric multidimensional scaling (NMDS) using the ‘metaMDS’ function in the ‘vegan’ package (Oksanen et al., 2016) in two dimensions and 999 simulations, with Jaccard distance for spider diets and Bray-Curtis distance for invertebrate communities. For the dietary NMDS, outliers (n = 21; samples containing rare taxa) obscured variation on one axis and were thus removed to facilitate separation of samples and achieve minimum stress. For visualization of the effect of continuous variables against the NMDS, surf plots were created with scaled coloured contours using the ‘ordisurf’ function in the ‘ggplot’ package (Wickham, 2016).</p> <p> </p> <p><em>Relationships between web characteristics and weather variables</em></p> <p>Web area and height were compared against weather and temporal variables using a multivariate linear model (MLM) with the ‘manylm’ command in ‘mvabund’ (Wang et al., 2012). Log-transformed web area and height comprised the multivariate dependent variable, and day, spider genus, temperature, precipitation, dewpoint, wind, pressure and two-way interactions between each of these and day and genus comprised the independent variables.</p> <p> </p> <p><em>Variation in spider prey choice across weather conditions</em></p> <p>To separately represent spiders from different weather conditions in prey choice analyses, sample dates for every spider were clustered based on the mean weather conditions (temperature, precipitation, dewpoint, wind and pressure) of the week before collection (7 days, to align approximately with spider gut DNA half-life; Greenstone et al., 2014). Alongside data from 2018 (n = 24 collection dates), two sampling periods from 2017 were included in the clustering to ascertain similarity of weather conditions for additional inter-annual prey choice analyses described below. The clustering process is described in Supplementary Information 3. Five clusters were generated: High Pressure (HPR), Hot (HOT), Wet Low Dewpoint (WLD), Dry Windy (DWI), Wet Moderate Dewpoint (WMD), and 2017 (2017 sampling periods).</p> <p>Prey preferences of spiders in each of the weather clusters was analysed using network-based null models in the ‘econullnetr’ package (Vaughan et al., 2018) with the ‘generate_null_net’ command. Consumer nodes in this case represented spiders belonging to each of the weather clusters. Econullnetr generates null models based on prey abundance, represented here by suction sample data, to predict how consumers will forage if based on the abundance of resources alone. These null models are then compared against the observed interactions of consumers (i.e., interactions of spiders within each weather cluster with their prey) to ascertain the extent to which resource choice deviated from random (i.e., density dependence). The trophic network was visualised with the associated prey choice effect sizes using ‘igraph’ (Csardi & Nepusz, 2006) with a circular layout, and as a bipartite network using ‘ggnetwork’ (Briatte, 2021; Wickham, 2016). The normalised degree of each weather cluster node was generated using the ‘bipartite’ package (Dormann, Gruber, & Fruend, 2008) and compared against the normalised degree of the same node in the null network to determine whether spiders were more or less generalist than expected by random. Prior to the prey choice analysis, an hemipteran prey identified no further than order level through dietary analysis was removed due to the inability to pair it to any present prey taxa with certainty.</p> <p> </p> <p><em>Validating and predicting relationships between years</em></p> <p>To test how generalisable the results are and the extent to which weather drives prey preferences, we used a measure of prey preference (observed/expected values; observed interaction frequencies divided by interaction frequencies expected by null models) from the above prey choice analysis to assess whether we could more accurately predict observed trophic interactions under similar weather conditions for data from a linked study at the same location in 2017. These additional data represent a subset of the spider taxa analysed above (<em>Tenuiphantes tenuis</em> and <em>Erigone</em> spp.) collected using the same methods by the same researchers and in the same locality (Cuff, Drake, et al., 2021).</p> <p>The similarity in weather conditions between the 2017 study period and each of the five 2018 weather clusters was determined via NMDS of the weather data in two dimensions with Euclidean distance. Centroid coordinates for each 2018 weather cluster and the 2017 data were extracted and pairwise distances calculated between weather clusters:</p> <p> </p> <p>In order, the most proximate weather clusters to the 2017 weather data were HPR (mean Euclidean distance = 8.845), HOT (9.290), WMD (13.626), DWI (13.817) and WLD (18.682; Figure S3).</p> <p>To facilitate comparison between the two years, observed/expected values from the 2018 prey choice models were extracted separately for each of the weather clusters and scaled between 0.1 and 1. For this, 0.1 was used as a minimum since 0 would result in interactions being excluded altogether in the null models, and one as a maximum given the limits of econullnetr but also because this is a multiplier applied to the prey abundances, so greater values would skew prey abundances beyond realistic proportions. Scaling was achieved by the following equation:</p> <p> </p> <p>Missing values (e.g., prey that were absent in certain weather conditions) were represented as 1 to prevent transformation of their abundances in the null models; this treats prey for which data were absent naively, but could increase perceived preferences for them. The scaled values were used to weight the abundance of prey available to the spiders in the 2017 data using the weighting option in econullnetr, whereby values less than 1 proportionally reduce the probability of that taxon being predated in the null models. This effectively redistributes the 2017 relative prey abundance data according to the preference effect sizes generated for each of the 2018 weather clusters. If prey preferences are similar between the 2017 spiders and those from the weather cluster being used to weight the model, the composition of simulated diets should more closely resemble observed diets and fewer significant deviations from the null model should be found.</p> <p>Null models were generated as above (<em>Variation in spider prey choice across weather conditions</em>) but based on the prey availability and trophic interactions from 2017 samples. Three types of model were run: i) a conventional model based on observed prey abundances; ii) a model with prey abundances set to be equal across all prey taxa; and iii) observed prey abundances weighted by prey preferences determined for each of the weather clusters in the 2018 prey choice analysis. A separate model was run for each 2018 weather cluster with abundances weighted by the corresponding scaled observed/expected values. The unweighted conventional model was compared against weighted models to ascertain whether the prey preference weightings from 2018 improved the predictive power of the null models. To compare effect sizes between the unweighted and each other null model for each resource taxon, mean standardised effect size (SES) values were calculated from the paired ‘pre-harvest’ and ‘post-harvest’ data from each model, and paired <em>t</em>-tests were carried out with these between the unweighted and each weighted model. The SES values were plotted for each model and joined between taxa to visualise these paired differences using ‘ggplot’ (Wickham, 2016). Null model-predicted trophic interactions were generated via a modified ‘econullnetr’ function (generate_null_net_indiv) which produces outputs at the individual level to generate simulated diets for individual spiders to compare dietary composition between null model predictions and observed data. These models were run with 2300 simulations to represent 50 simulations per individual spider in the 2017 dataset (n = 46). Null diets were associated with sample IDs by aggregating the 50 simulations per sample and retaining a mean incidence of prey (i.e., mean occurrence across all 50 simulations). A visualisation of the per-sample differences in null model and observed data was generated via NMDS. Mean centroid coordinates for the observed 2017 data and the predicted diets of each model were extracted and the Euclidean distance between the observed data centroid and that of each model was calculated (as above for weather conditions).</p> <p> </p> <p><strong>Supplementary Information 1: Field collection and sample processing</strong></p> <p>Money spiders (Araneae: Linyphiidae) and wolf spiders (Araneae: Lycosidae) were visually located along transects in two adjacent barley fields at Burdons Farm, Wenvoe in South Wales (51°26'24.8"N, 3°16'17.9"W) and collected from occupied webs and the ground, between April and September 2018 (five visits per week of which spiders from 24 collection dates were used). Transects were randomly distributed across the entire field. Along these transects, 64 separate 4 m<sup>2</sup> quadrats, at least 10 m apart, were searched and all observed linyphiids and lycosids were collected. Spiders were placed in 100 % ethanol using an aspirator, regularly changing meshing to limit potential cross-contamination. Linyphiids occupying webs were prioritised for collection, but ground-active linyphiid spiders were also collected. For each spider taken from a web, the height of the web from the ground (mm) and its approximate dimensions were recorded, the latter calculated as approximate web area (mm<sup>2</sup>). Spiders were taken to Cardiff University, transferred to fresh ethanol, adults identified to species-level and juveniles to genus, and stored at -80 °C in 100 % ethanol until DNA extraction.</p> <p>To obtain data on local prey density, ground and crop stems were suction sampled using a ‘G-vac’ for approximately 30 seconds at each 4 m<sup>2</sup> quadrat (n = 64) from which spiders were collected. The collected material was emptied into a bag, any organisms immediately killed with ethyl-acetate and material frozen for storage before sorting into 70 % ethanol in the lab. All invertebrates were identified to family level to match the resolution of the least resolved of the metabarcoding-derived trophic interaction data, and due to difficulties associated with identification to finer taxonomic resolution for many taxa. Exceptions included springtails of the superfamily Sminthuroidea (Sminthuridae and Bourletiellidae were often indistinguishable following suction sampling and preservation due to the fine features necessary to distinguish them) which were left at super-family, mites (many of which were immature or in poor condition) which were identified to order level, and wasps of the superfamily Ichneumonoidea which were identified no further due to obscurity of wing venation due to damage.</p> <p> </p> <p><strong>Supplementary Information 2: Molecular analysis and bioinformatics</strong></p> <p><em>Extraction and high-throughput sequencing of spider gut DNA</em></p> <p>Given their prevalence in field collections, dietary analysis was carried out for the linyphiid genera <em>Erigone</em>, <em>Tenuiphantes</em>, <em>Bathyphantes</em> and <em>Microlinyphia </em>(Araneae: Linyphiidae), and the Lycosidae genus <em>Pardosa</em>. Spiders were transferred to and washed in fresh 100 % ethanol to reduce external contaminants prior to identification via morphological key (Roberts, 1993). Abdomens were removed from spiders and again transferred to and washed in fresh 100 % ethanol. DNA was extracted from the abdomens via Qiagen TissueLyser II and DNeasy Blood & Tissue Kit (Qiagen) as per the manufacturer protocol, but with an extended lysis time of 12 hours to account for the complex and branched gut system in spider abdomens (Krehenwinkel et al., 2017).</p> <p>For amplification of DNA, two primer pairs were used. BerenF-LuthienR (Cuff et al., 2021) amplified a broad range of invertebrates including spiders, and TelperionF-LaureR (Cuff et al., 2022), amplified a range of invertebrates but fewer spiders. Primers were labelled with unique 10 bp molecular identifier tags (MID-tags) so that each individual had a unique pairing of forward and reverse tags for identification of each spider post-sequencing. PCR reactions of 25 µl contained 12.5 µl Qiagen PCR Multiplex kit, 0.2 µmol (2.5 µl of 2 µM) of each primer and 5 µl template DNA. Reactions were carried out in the same thermocycler, optimised via temperature gradient, with an initial 15 minutes at 95 °C, 35 cycles of 95 °C for 30 seconds, the primer-specific annealing temperature for 90 seconds and 72 °C for 90 seconds, respectively, followed by a final extension at 72 °C for 10 minutes. BerenF-LuthienR and TelperionF-LaureR used annealing temperatures of 52 °C and 42 °C, respectively.</p> <p>Within each PCR 96-well plate, 12 negative controls (extraction and PCR), 2 blank controls and 2 positive controls were included (i.e. 80 samples per plate), based on Taberlet <em>et al. </em>(2018). Positive controls were mixtures of invertebrate DNA comprised of non-native Asiatic species in four different proportions and blanks were empty wells within each plate to identify tag-jumping into unused MID-tag combinations. PCR negative controls were DNase-free water treated identically to DNA samples. A negative control was present for each MID-tag to identify any contamination of primers. All PCR products were visualised in a 2 % agarose gel with SYBRSafe (Thermo Fisher Scientific, Paisley, UK) and placed in categories based on their relative brightness. The concentration of these brightness categories was quantified via Qubit dsDNA High-sensitivity Assay Kits (Thermo Fisher Scientific, Waltham, MA, USA) with at least three representatives of each category per plate. The PCR products were then proportionally pooled according to these concentrations. Each pool was cleaned via SPRIselect beads (Beckman Coulter, Brea, USA), with a left-side size selection using a 1:1 ratio (retaining ~300-1000 bp fragments). The concentration of the pooled DNA was then determined via Qubit dsDNA High-sensitivity Assay Kits and pooled together into one library per primer pair. Library preparation for Illumina sequencing was carried out on the cleaned libraries via NEXTflex Rapid DNA-Seq Kit (Bioo Scientific, Austin, USA) and samples were sequenced on an Illumina MiSeq via a V3 chip with 300-bp paired-end reads (expected capacity ≤25,000,000 reads).</p> <p><em>Bioinformatic analysis</em></p> <p>Bioinformatic analysis followed Drake et al., (2022). The Illumina run generated 11,165,405 and 10,959,010 reads for BerenF-LuthienR and TelperionF-LaureR, respectively, which were quality-checked and paired via FastP (Chen et al., 2018) to retain only sequences of at least 200 bp with a quality threshold of 33, resulting in 10,561,874 and 9,355,112 paired reads. The paired reads were demultiplexed and assigned to their respective spider sample according to their MID-tags via the “trim.seqs” command in Mothur v1.39.5 (Schloss et al., 2009), leaving 7,854,610 and 7,437,929 reads with exact matches to the primer and MID-tags.</p> <p>Replicates were removed, and denoising and clustering to zero-radius operational taxonomic units (ZOTUs; clustered without % identity to avoid multiple species represented within a single operational taxonomic unit (OTU)) completed via Unoise3 in Usearch11 (Edgar, 2010). The resultant sequences were assigned a taxonomic identity from GenBank via BLASTn v2.7.1 (Camacho et al., 2009) using a 97 % identity threshold (Alberdi et al., 2017). The BLAST output was analysed in MEGAN v6.15.2 (Huson et al., 2016). Where the top BLAST hit, determined by lowest e-value, was resolved at a higher taxonomic level than species-level, the results were checked; where possibly erroneous entries were preventing species-level assignment (e.g., poorly resolved identifications on GenBank), finer resolution was assigned based on the next-closest match. Where ZOTUs were assigned the same taxon, these were aggregated.</p> <p>Data clean-up used the optimal minimum sequence copy thresholds identified by Drake et al. (2022). The maximum value for a ZOTU present in blank or negative controls was identified and subtracted from all read counts for that ZOTU to remove background contaminants. Simultaneously, known lab contaminants (e.g., German cockroach <em>Blattella germanica</em>), artefacts and errors of the sequencing process, unexpected reads in positive controls and positive control taxon reads in dietary samples were identified. These were calculated as a percentage of their respective sample’s read count and any read counts lower than the highest of these percentages for their respective sample were removed to eliminate additional instances of contamination. These thresholds were defined as 0.38 % and 0.39 % for BerenF-LuthienR and TelperionF-LaureR, respectively. The data from the two libraries (i.e., from each primer pair) were then aggregated together by sample and aggregated again by taxon. Non-target taxa (e.g., fungi) and instances in which predator DNA was amplified (i.e., ZOTUs with high read counts matching the individual’s morphological identity) were removed. All remaining read counts were converted to presence-absence.</p> <p> </p> <p><strong>Supplementary Information 3: Cluster analysis</strong></p> <p>Prior to clustering, weather variables were scaled by subtracting the mean and dividing by the standard deviation. A Euclidean distance matrix was calculated using the ‘dist’ function, and this scaled distance matrix was hierarchically clustered using the ‘hclust’ function. Optimal clustering solutions were determined by comparison of Dunn’s index between methods and <em>k</em> values; this was calculated using the ‘dunn’ function in the “clValid” package (Brock et al., 2008) for each cluster <em>k</em> value above five until the Dunn index decreased. The <em>k</em> value after which Dunn’s index decreased was deemed the optimal solution for each clustering method. Clustering methods based on ‘average’, ‘complete’, ‘single’, ‘median’, ‘centroid’ and ‘mcquitty’ linkages were compared, and the ‘complete’ method selected for subsequent analysis as it resulted in the smallest number of clusters (6; thus, the most efficient simplification of the data; Figure S1). Different clustering methods altered the composition of some clusters, but most sampling dates showed consistent clustering between methods.</p> <p>A heatmap dendrogram was produced using the ‘heatmap.2’ function in the ‘gplots’ package (Warnes et al., 2020), with cluster colours assigned with the ‘Accent’ palette of ‘RColorBrewer’ (Neuwirth, 2014) and relative weather value colour scaling generated using the ‘viridis’ package (Garnier 2018; Figure S2). Weather clusters were named according to unique characteristics relative to the other clusters. These names comprise: High Pressure (HPR; days 142, 253, 256, 250, 173), Hot (HOT; days 162, 204, 205, 208, 201, 187, 198, 197, 183, 184, 194, 190 and 191), Wet Low Dewpoint (WLD; day 121), Dry Windy (DWI; days 131, 169 and 170), Wet Moderate Dewpoint (WMD; days 149 and 152), and 2017 (pre- and post-harvest 2017 sampling periods).</p> <p> </p> <p> </p> <p><strong>References</strong></p> <p>Alberdi, A., Aizpurua, O., Gilbert, M. T. P., & Bohmann, K. (2017). Scrutinizing key steps for reliable metabarcoding of environmental samples. <em>Methods in Ecology and Evolution</em>, <em>9</em>(1), 1–14. https://doi.org/10.1111/2041-210X.12849</p> <p>Brock, G., Pihur, V., Datta, S., & Datta, S. (2008). clValid: an R package for cluster validation. <em>Journal of Statistical Software</em>, <em>25</em>(4), 1–22.</p> <p>Camacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., & Madden, T. L. (2009). BLAST+: architecture and applications. <em>BMC Bioinformatics</em>, <em>10</em>, 1–9. https://doi.org/10.1186/1471-2105-10-421</p> <p>Chen, S., Zhou, Y., Chen, Y., & Gu, J. (2018). Fastp: An ultra-fast all-in-one FASTQ preprocessor. <em>Bioinformatics</em>, <em>34</em>(17), i884–i890. https://doi.org/10.1093/bioinformatics/bty560</p> <p>Cuff, J. P., Drake, L. E., Tercel, M. P. T. G., Stockdale, J. E., Orozco-terWengel, P., Bell, J. R., Vaughan, I. P., Müller, C. T., & Symondson, W. O. C. (2021). Money spider dietary choice in pre- and post-harvest cereal crops using metabarcoding. <em>Ecological Entomology</em>, <em>46</em>(2), 249–261.</p> <p>Cuff, J. P., Tercel, M. P. T. G., Drake, L. E., Vaughan, I. P., Bell, J. R., Orozco-terWengel, P., Müller, C. T., & Symondson, W. O. C. (2022). Density-independent prey choice, taxonomy, life history and web characteristics determine the diet and biocontrol potential of spiders (Linyphiidae and Lycosidae) in cereal crops. <em>Environmental DNA</em>, <em>4</em>(3), 549–564.</p> <p>Drake, L. E., Cuff, J. P., Young, R. E., Marchbank, A., Chadwick, E. A., & Symondson, W. O. C. (2022). An assessment of minimum sequence copy thresholds for identifying and reducing the prevalence of artefacts in dietary metabarcoding data. <em>Methods in Ecology and Evolution</em>, <em>13</em>(3), 694–710.</p> <p>Edgar, R. C. (2010). Search and clustering orders of magnitude faster than BLAST. <em>Bioinformatics</em>, <em>26</em>(19), 2460–2461. https://doi.org/10.1093/bioinformatics/btq461</p> <p>Garnier, S. (2018). <em>viridis: default color maps from ‘matplotlib’</em> (0.5.1). https://cran.r-project.org/package=viridis</p> <p>Huson, D. H., Beier, S., Flade, I., Górska, A., El-Hadidi, M., Mitra, S., Ruscheweyh, H. J., & Tappu, R. (2016). MEGAN Community Edition - interactive exploration and analysis of large-scale microbiome sequencing data. <em>PLoS Computational Biology</em>, <em>12</em>(6), 1–12. https://doi.org/10.1371/journal.pcbi.1004957</p> <p>Krehenwinkel, H., Kennedy, S., Pekár, S., & Gillespie, R. G. (2017). A cost-efficient and simple protocol to enrich prey DNA from extractions of predatory arthropods for large-scale gut content analysis by Illumina sequencing. <em>Methods in Ecology and Evolution</em>, <em>8</em>, 126–134. https://doi.org/10.1111/2041-210X.12647</p> <p>Neuwirth, E. (2014). <em>RColorBrewer: ColorBrewer palettes</em> (1.1-2). https://cran.r-project.org/package=RColorBrewer</p> <p>Roberts, M. J. (1993). <em>The Spiders of Great Britain and Ireland (Compact Edition)</em> (3rd ed.). Harley Books.</p> <p>Schloss, P. D., Westcott, S. L., Ryabin, T., Hall, J. R., Hartmann, M., Hollister, E. B., Lesniewski, R. A., Oakley, B. B., Parks, D. H., Robinson, C. J., Sahl, J. W., Stres, B., Thallinger, G. G., Van Horn, D. J., & Weber, C. F. (2009). Introducing mothur: open-source, platform-independent, community-supported software for describing and comparing microbial communities. <em>Applied and Environmental Microbiology</em>, <em>75</em>(23), 7537–7541. https://doi.org/10.1128/AEM.01541-09</p> <p>Taberlet, P., Bonin, A., Zinger, L., & Coissac, E. (2018). <em>Environmental DNA</em>. Oxford University Press.</p> <p>Warnes, G. R., Bolker, B., Bonebakker, L., Gentleman, R., Huber, W., Liaw, A., Lumley, T., Maechler, M., Magnusson, A., Moeller, S., Schwartz, M., & Venables, B. (2020). <em>gplots: Various R programming tools for plotting data</em> (R package version 3.1.0). https://cran.r-project.org/package=gplots</p>
Universal Chalcidoidea Database World Wide Web electronic publication. http://www.nhm.ac.uk/chalcidoids. hash://sha256/562fc5f7ac62b0dba9952e267dc839ab16be3efcf149b98a8d77f6e88bce4f53 hash://md5/de2883bc9f8b1e79c96f5d00d650f596
<p>This repository contains an archival copy of the Universal Chalcidoidea Database by J.S. Noyes in their original Paradox Database (https://en.wikipedia.org/wiki/Paradox_%28database%29) file format. </p> <p>Files were provided by J.S. Noyes in period 2023-03/2023-04 and gave consent to publish the data under CC0.</p> <p>For getting started, please read the 00-Instructions.pdf first. Then, suggest to take a look at a flowchart in 01-Flowchart.pdf and the Table Structure in 02-TableStructure.pdf.</p> <p><strong>Citation</strong><br> On use of this data, please follow academic tradition and cite the data using: </p> <p>Noyes, J.S. March 2019. Universal Chalcidoidea Database. World Wide Web electronic publication. http://www.nhm.ac.uk/chalcidoids. hash://sha256/562fc5f7ac62b0dba9952e267dc839ab16be3efcf149b98a8d77f6e88bce4f53 hash://md5/de2883bc9f8b1e79c96f5d00d650f596 .</p> <p><strong>Signed Content </strong></p> <p>Using Preston [1,2], the UCD content was packaged and their provenance was signed.</p> <pre><code>preston history\ --anchor hash://sha256/562fc5f7ac62b0dba9952e267dc839ab16be3efcf149b98a8d77f6e88bce4f53\ --remote https://raw.githubusercontent.com/jhpoelen/ucd/main/data\ --remote https://zenodo.org/record/7864604/files\ --remote https://softwareheritage.org\ --remote https://linker.bio </code></pre> <p>yielded:</p> <pre><code><hash://sha256/562fc5f7ac62b0dba9952e267dc839ab16be3efcf149b98a8d77f6e88bce4f53> <http://www.w3.org/ns/prov#wasDerivedFrom> <hash://sha256/dc3f137ed7e456dd964545527cfff3acc3c1655baeaebecb6d07cdf3e1bbd549> . <hash://sha256/dc3f137ed7e456dd964545527cfff3acc3c1655baeaebecb6d07cdf3e1bbd549> <http://www.w3.org/ns/prov#wasDerivedFrom> <hash://sha256/298581b34133b518f251f4321f1920488afd923f3308e45b9c1d169da0e16b5b> . <hash://sha256/298581b34133b518f251f4321f1920488afd923f3308e45b9c1d169da0e16b5b> <http://www.w3.org/ns/prov#wasDerivedFrom> <hash://sha256/ec1760dc83dfb17df003ef5e626b965dd4403850bc58286ac59c7ef3a447e063> . <hash://sha256/ec1760dc83dfb17df003ef5e626b965dd4403850bc58286ac59c7ef3a447e063> <http://www.w3.org/ns/prov#wasDerivedFrom> <hash://sha256/eb416c97bf52a36b31ece2b47431a6a4a9bda7f52b9bc8ccb92f91f5c1bdf268> . <urn:uuid:0659a54f-b713-4f86-a917-5be166a14110> <http://purl.org/pav/hasVersion> <hash://sha256/eb416c97bf52a36b31ece2b47431a6a4a9bda7f52b9bc8ccb92f91f5c1bdf268> . </code></pre> <p>This archive can be cloned using:</p> <pre><code>preston clone\ --anchor hash://sha256/562fc5f7ac62b0dba9952e267dc839ab16be3efcf149b98a8d77f6e88bce4f53\ --remote https://raw.githubusercontent.com/jhpoelen/ucd/main/data\ --remote https://zenodo.org/record/7864604/files\ --remote https://softwareheritage.org\ --remote https://linker.bio </code></pre> <p>or alternatively, by downloading a zip archive from :</p> <p>https://github.com/jhpoelen/ucd/archive/42a5815a2c50e396f1e34e9f73e0b13fcc67c44f.zip</p> <p><br> <strong>References</strong></p> <p>[1] MJ Elliott, JH Poelen, JAB Fortes (2020). Toward Reliable Biodiversity Dataset References. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2020.101132</p> <p>[2] Elliott, M. J., Poelen, J. H., & Fortes, J. (2023, accepted with minor revisions). Signed Citations: Making Persistent and Verifiable Citations of Digital Scientific Content. https://doi.org/10.31222/osf.io/wycjn<br> </p>
Direct and indirect effects of climate and land use change on food webs in lakes and streams
<p>Here, we provide the data and code necessary to reproduce the workflow and analysis in: Barbosa and Siqueira. Direct and indirect effects of climate and land use change on food webs in lakes and streams. A preprint is available at https://doi.org/10.1101/2022.04.18.488700</p> <p>We compiled multicontinental data to investigate how climate and land use change are related to the structure of freshwater food webs, considering the inherent differences in lentic and lotic ecosystems. We analyzed the direct and indirect relationships between land use intensity, and temperature and precipitation changes, and food webs using multi-group structural equation modeling. Freshwater food webs were obtained from three sources: the Mangal interaction database, using the rmangal package in R, the GlobAl databasE of traits and food Web Architecture (GATEWAY) version 1.0, and the Interaction Web Data Base (IWDB). We also included food webs acquired from a search in the Web of Science Core Collection. Land use data was compiled from the global ESA CCI database, an annually generated land cover product at 300 m resolution for the period 1992 – 2015. Climate data was compiled from the TerraClimate database, a monthly generated product for climate and climatic water balance for global terrestrial surfaces at ~ 4 km for the period 1958 – 2015. </p>
Data set and classification method for low quality web traffic identification in video marketing campaigns
<p>Final outcomes of the InPreVi (AI4Media) project developed in 2022. </p> <p>1. Data set describing the statistics of the video ad marketing campaigns</p> <p>2. Script for web traffic classification</p>
RDF dataset produced in the work "Exploring Adverse Outcome Pathways for Nanomaterials with semantic web technologies"
<p>Adverse Outcome Pathways (AOPs) have been proposed to facilitate mechanistic understanding of interactions of chemicals/materials with biological systems. Each AOP starts with a molecular initiating event (MIE) and possibly ends with adverse outcome(s) (AOs) via a series of key events (KEs). So far, the interaction of engineered nanomaterials (ENMs) with biomolecules, biomembranes, cells, and biological structures, in general, is not yet fully elucidated. There is also a huge lack of information on which AOPs are ENMs-relevant or -specific, despite numerous published data on toxicological endpoints they trigger, such as oxidative stress and inflammation. We propose to integrate related data and knowledge recently collected. Our approach combines the annotation of nanomaterials and their MIEs with ontology annotation to demonstrate how we can then query AOPs and biological pathway information for these materials. We conclude that a FAIR (Findable, Accessible, Interoperable, Reusable) representation of the ENM-MIE knowledge simplifies integration with other knowledge.</p>
Ports, Past and Present web archive - All Stories
<p>This .wacz file is the web archive for the story collection at https://portspastpresent.eu/, completed on the 13th of July 2023. It captures the navigation options, user experience and story structure of the Omeka story collection at that time for posterity. For more information, see the attached README.</p>
A core ontology for modeling life cycle sustainability assessment on the Semantic Web with Accompanying Database
<p>To enable and support the uptake of semantic ontologies, we present a core ontology developed specifically to capture the data relevant for life cycle sustainability assessment. We further demonstrate the utility of the ontology by using it to integrate data relevant to sustainability assessments, such as EXIOBASE and the Yale Stocks and Flow Database to the Semantic Web. These datasets can be accessed by the machine-readable endpoint using SPARQL, a semantic query language.</p>
Accompanying Dataset migr_asyappctzm for Efficient Analytical Queries on Semantic Web Data Cubes
<p>This dataset shows how the Eurostat data cube in the orginal publicatin is modelled in QB4OLAP.</p> <p>This data is based on statistical data about asylum applications to the European Union, provided by Eurostat on</p> <p><a href="http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm">http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm</a></p> <p>Further data has been integrated from: https://github.com/lorenae/qb4olap/tree/master/examples</p>
PLAE web app enables powerful searching and multiple visualizations across one million unified single-cell ocular transcriptomes
<p>Supplementary Data for "PLAE web app enables powerful searching and multiple visualizations across one million unified single-cell ocular transcriptomes"</p> <p> </p>
Annotated Web Tables
<p>Data sets used for experimental evaluation in the related publication:</p> <p><em>Matching Web Tables with Knowledge Base Entities: From Entity Lookups to Entity Embeddings<br> International Semantic Web Conference (1) 2017: 260-277</em><br> <em>Vasilis Efthymiou Oktie Hassanzadeh Mariano Rodríguez-Muro Vassilis Christophides</em></p> <p>The gold standard data sets are collections of web tables:</p> <p><strong>T2D</strong> (<strong>v1</strong>) consists of a schema-level gold standard of 1,748 Web tables, manually annotated with class- and property-mappings, as well as an entity-level gold standard of 233 Web tables.</p> <p><strong>Limaye</strong> consists of 400 manually annotated Web tables with entity-, class-, and property-level correspondences, where single cells (not rows) are mapped to entities. The corrected version of this gold standard is adapted to annotate rows with entities, from the annotations of the label column cells.</p> <p><strong>WikipediaGS </strong>is an instance-level gold standard developed from 485K Wikipedia tables, in which links in the label column are used to infer the annotation of a row to a DBpedia entity.</p> <p> </p> <p>Note on license: please refer to the README.txt. Data is derived from Wikipedia and other sources may have different licenses.</p> <p>Wikipedia contents can be shared under the terms of Creative Commons Attribution-ShareAlike License<br> as outlined on Wikipedia: <a href="https://en.wikipedia.org/wiki/Wikipedia:Reusing_Wikipedia_content">https://en.wikipedia.org/wiki/Wikipedia:Reusing_Wikipedia_content</a></p> <p>The correspondences of the T2D Gold standard is provided under the terms of the Apache license. The Web tables are provided according the same terms of use, disclaimer of warranties and limitation of liabilities that apply to the Common Crawl corpus. The DBpedia subset is licensed under the terms of the Creative Commons Attribution-ShareAlike License and the GNU Free Documentation License that applies to DBpedia.<br> Limaye gold standard is downloaded from: <a href="http://websail-fe.cs.northwestern.edu/TabEL/">http://websail-fe.cs.northwestern.edu/TabEL/</a> (download date: August 25, 2016). Please refer to the original website and the following paper for more details and citation information:<br> G. Limaye, S. Sarawagi, and S. Chakrabarti. Annotating and Searching Web Tables Using Entities, Types and Relationships. PVLDB, 3(1):1338–1347, 2010.</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
University of Michigan Biological Station cumulative food web data for terrestrial habitats, 1909-2023.
Here, we present species and interaction lists for a food web of the aboveground terrestrial habitats at the University of Michigan Biological Station (UMBS). The site is composed predominantly of dry-mesic, northern hardwood forests with patches of wooded wetlands (hardwood conifer swamp). Taxa were sourced from lists provided by UMBS, from resident biologists’ personal observations, museum specimens, online databases, historical censuses, and BioBlitz events. Only those that could be resolved to species-level or were genera with < 20 species in the Nearctic were included. We also excluded species that do not have a significant lifestage or feeding behavior in aboveground terrestrial habitats. Our focal taxonomic groups include vascular plants, arthropods, birds, mammals, reptiles, amphibians. The majority of arthropods are insects; non-insect arthropods were highly underrepresented in our lists. Interactions were sourced from online databases, naturalist observations, and field guides and accumulated into a “metaweb” of all potential interactions between local species. Interactions were checked by experts to plausibly occur in the aboveground terrestrial environments at UMBS, given species’ phenology, traits, and habitat usage. Interactions at any taxonomic level were included, so long as they were approved to potentially occur between all species by our experts. To study the effect of taxonomic resolution on food web structure, in this dataset, we retained records at coarser taxonomic groupings even if more highly resolved records were also approved. We included all direct interactions among species in our system with a bioenergetic flow (i.e., one species consuming another), differentiated by their focal resource. We broadly categorized the resources as animal tissues, either (1) live tissues and as prey, or (2) scavenged as carrion, carcasses, or other decaying animal remains, or as plant tissues, grouped as (3) leaves and stems, including grasses, exudates, et
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.