Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
115
datasets available to search
ShareScore release 0.9.0
Dataset results
115 results for “abundance estimation”
Species-level estimated abundances and zero counts of nighttime collected female mosquitoes 2014 - 2022 (Derived from NEON Mosquitoes sampled from CO2 traps (DP1.10043.001, RELEASE-2024))
This Level 2 data package contains species level estimated abundances, including zero counts, and estimated mean number of female mosquitoes per trap derived from the NEON Mosquitoes sampled from CO2 traps (DP1.10043.001), RELEASE-2024 Level 0 data (https://doi.org/10.48443/3cyq-6v47). The data set includes mosquito records of traps collecting mosquito samples at night, for up to 24 trap hours, across a total of 20 terrestrial core and 27 terrestrial gradient sites from 2014 to 2022. To ensure high confidence in abundance estimates, records were only included when at least 90% of collected individuals were identified to sex, and 90% of female specimens were identified to species. Information across multiple QC/QA fields within the NEON mosquito data was evaluated to identify and exclude records where confidence in estimated abundances may have been compromised. Species level zero counts were added for all species collected at least once within the sampling year and trap location. Additionally, species level zero counts were included for trap events where only male mosquitoes had been collected or where QC/QA remarks indicated traps were inactive due to cold temperatures. The data set provides an analysis ready time series of estimated abundances across NEON sites and plots. An R Markdown file that contains descriptions of the QC/QA and data filtering steps along with annotated code, as well as data tables used to filter active and inactive trap events based on QC/QA fields, are published with the data package. Any questions about this data package should be directed to Amely Bauer listed under contacts.
Estimation of Abundance and Distribution of Salt Marsh Plants from Images Using Deep Learning
Recent advances in computer vision and machine learning, most notably deep convolutional neural networks (CNNs), are exploited to identify and localize various plant species in salt marsh images. Three different approaches are explored that provide estimations of abundance and spatial distribution at varying levels of granularity in terms of spatial resolution. In the coarsest-grained approach, CNNs are tasked with identifying which of six plant species are present/absent in large patches within the salt marsh images. CNNs with diverse topological properties and attention mechanisms are shown capable of providing accurate estimations with > 90% precision and recall in the case of the more abundant plant species whereas the performance of the CNNs is observed to decline in the case of less common plant species. Estimation of percent cover of each plant species is performed at a finer spatial resolution, where smaller image patches are extracted and the CNNs tasked with identifying the plant species or substrate at the center of the image patch. In an ecological setting, several image patches (~100) are extracted and classified using this approach to estimate the percent cover of the various plant species in the image. For the percent cover estimation task, the CNNs are observed to exhibit a performance profile similar to that for the presence/absence estimation task, but with an ~ 5–10% reduction in precision and recall. Finally, estimation of the spatial distribution of the various plant species is performed via semantic segmentation of the input images at the finest level of granularity in terms of spatial resolution. The Deeplab-V3 semantic segmentation architecture is observed to provide very accurate estimations for abundant plant species; however, a significant degradation in performance is observed in the case of less abundant plant species and, in extreme cases, rare plant classes are seen to be ignored entirely. Overall, a clear trade-off is observed between
Finite amplitude sound propagation effects in volume backscattering measurements for fish abundance estimation
<p>The upload contains measurement and simulation data for finite-amplitude sound propagation effects in volume backscattering measurements. The experimental data are from a trawl survey conducted in the North Sea with R/V "G. O. Sars", 6-7 November 2004, passing several times over a group of Atlantic mackerel schools. The measurements are of the relative area backscattering coefficient, relative to 38 kHz, 2000 W power setting, at</p> <p>(1) 120 kHz with 250 W transmit power setting, 200 kHz with 120 W transmit power setting<br> (2) 120 kHz with 1000 W power setting, 200 kHz with 1000 W power setting.</p> <p>A Simrad EK60 echosounder system was used, alternating between the low (1) and high (2) power settings through the measurement series.</p> <p>The corresponding simulation data are calculated using the Bergen Code numerical solver of the KZK Equation. The medium parameters input to the simulations are based on CTD data from the field survey . The transducer and amplitude data were found by laboratory measurements on echo sounders of the same type as used in the survey.</p> <p>.m files are included for both .mat data files, with details on how to read the data.</p> <p>An article describing the data has been submitted by the authors to Acta Acustica, 2022.</p>
Data from: Estimation of regional annual abundance and evidence for increasing numbers of white sharks off central California
<p>Raw data consisting of individual identification photographs of white sharks (Carcharodon carcharias) and data table with corresponding metadata. These data support PhD thesis of Paul E. Kanive entitled "VITAL RATES, ANNUAL ABUNDANCE, AND MOVEMENT OF WHITE SHARKS IN THE NORTHEASTERN PACIFIC" and peer-reviewed manuscript "Estimation of regional annual abundance and evidence for increasing numbers of white sharks off central California."</p>
F I G U R E 3 in Estimating multiple years, tributary-specific, and overall Atlantic salmon smolt abundance in a large Canadian catchment using capture-mark-recapture experiments
F I G U R E 3 (a) Posterior distribution of the difference between the true value and posterior distribution estimates (in percentage of the true value) of annual total smolt abundance obtained from DM (Dirichlet-multinomial) model M5 (simulation exercise). Each line is a year, and the 10 (one for each replicate) distributions are overlaid. The overall average difference is indicated in the left corner of the panel and represented by a red dashed vertical line. (b) C.V. of the posterior distributions of total abundance obtained from DM model M5. Each blue dot corresponds to the C.V. of a unique year/replicate, and each dashed blue horizontal line corresponds to the average C.V. of a replicate across years. The average C.V. across years and replicates (red dashed line) is also indicated in the top-right corner of the panel.
F I G U R E 4 in Estimating multiple years, tributary-specific, and overall Atlantic salmon smolt abundance in a large Canadian catchment using capture-mark-recapture experiments
F I G U R E 4 (a–c) Posterior distribution of the relative contribution of the smolt abundance associated with each upstream rotary screw trap (RST) and (d) posterior distribution of the relative contribution of the "rest" of the smolt abundance in relation to the total abundance for the DM (Dirichletmultinomial) model M5 (simulation exercise). The red dots correspond to the true values in the simulated dataset.
F I G U R E 1 in Estimating multiple years, tributary-specific, and overall Atlantic salmon smolt abundance in a large Canadian catchment using capture-mark-recapture experiments
F I G U R E 1 Restigouche River catchment map with main streams and lakes. Red dots indicate the location of the rotary screw traps (RST). Black polygons indicate subcatchments draining into upstream RSTs. Top-right panels: zoom on the location of the two downstream RSTs and example of a smolt with a streamer tag. Photo credit: Marie-Camille St-Amour.
F I G U R E 5 in Estimating multiple years, tributary-specific, and overall Atlantic salmon smolt abundance in a large Canadian catchment using capture-mark-recapture experiments
F I G U R E 5 (left panels, a–e) Posterior distributions of the estimated annual catchability θt,k at each rotary screw trap (RST), (middle panels, f– h) posterior distributions of the annual smolt abundance estimates associated with each upstream RST (Nmt,i), (i) the total smolt abundance for the Restigouche catchment (Nmtot), and (right panels, j–m) posterior distributions of the annual relative proportions of the total smolt abundance t of the Restigouche catchment estimated for each upstream RST and for the unsampled areas (rest, bottom-right panel) using the Restigouche dataset. In the catchability panels (a–e), the dashed gray line is the median, and the light and dark gray areas indicate the 2.5th–97.5th and 25th– 75th interquantile ranges, respectively, of the hyperparameter μθ, that is, the average catchability of each RST of the time series. In the relative k proportions panels (f–i) the thick dashed colored line in the top three panels indicates the average proportion of the total catchment abundance estimated at the upstream RST. The thick dashed gray lines indicate the proportions of the Restigouche catchment wetted areas of the subbasins sampled at the upstream RST(j–l) and the wetted area for the remainder of the Restigouche catchment not sampled by any upstream RST (m); the proportions for the rest of the catchment area change over years based on the specific upstream RSTs that operated in a given year. For all panels, the colored dot is the median, and the thin and thick vertical segments indicate the 2.5th–97.5th and 25th–75th interquantile ranges, respectively.
A hierarchical dependent double-observer method for estimating waterfowl breeding pairs abundance from helicopters
<p>We applied a dependent double-observer method for helicopter surveys and developed a hierarchical Bayesian model as a means to adjust counts of waterfowl for incomplete detection. We conducted our study using 52 plots in Labrador, Canada. A designated pair of primary observers reported counts and location of all waterfowl flocks that they detected to a pair of secondary observers, including details regarding the species, age and sex of observed birds. Secondary observers then reported any additional flocks observed by them but missed by the primary observers. The pairs of observers alternated between primary and secondary roles during the course of the survey, as well as position (front or back) within the helicopter. We used hierarchical Bayesian models to estimate detection probabilities of waterfowl flocks, as well as derive species-specific detection-corrected abundance and sex composition estimates of flocks. The hierarchical model output allowed us to derive estimates of indicated breeding pairs for each species in the survey area corrected for incomplete detection. Observers seated in the back of the helicopter had higher detection probabilities (0.89; 90% Bayesian Credible Intervals [BCI] = 0.82 – 0.95) than those in the front (0.74; 90% BCI = 0.66 – 0.83), and observer experience had a limited effect on detection. Total crew detection probabilities ranged between 0.99 (90% BCI = 0.97 – 1.00) and 0.97(90% BCI = 0.94 – 0.99), depending on the individual observers' position and role in the helicopter. Detection probabilities were higher for sea ducks and diving ducks and lower for dabbling ducks. Observers generally missed less than 5% of the total indicated pairs for all species. We recommend that detection in helicopter surveys be measured to control for observer turnover, observer experience, and aircraft-related differences in visibility.</p>
Copper River, Alaska, Chinook Salmon Inriver Abundance Estimate 2018-2021 DATA ARCHIVE
<p>Long-term monitoring of returning adult Chinook salmon (<em>Oncorhynchus tshawytscha</em>) abundance on the Copper River, AK, has been conducted using fishwheels and two-sample mark-recapture methods since 2003. This data archive is from from the 2018-2021 field seasons. The annual objective was to estimate the inriver abundance of Copper River Chinook salmon such that the estimate was within 25% of the true abundance 95% of the time. This data represents annual catch, bycatch, tagging site data, recapture site data, session data, QC check tables, CPUE, mark-recapture matrix, mark-recapture stratification tables, effort, and daily catch matrix. </p> <p>See annual report for methodology, analyses and results @ http://akssf.org/default.aspx?id=3477 or contact the Alaska Sustainable Salmon Fund or U.S. Fish and Wildlife Service Office of Subsistence Management Fisheries Resource Monitoring Program or Native Village of Eyak DENR. </p> <p> </p>
Estimating the abundance of the critically endangered Baltic Proper harbour porpoise (Phocoena phocoena) population using passive acoustic monitoring
<p>Knowing the abundance of a population is a crucial component to assess its conservation status and develop effective conservation plans. For most cetaceans, abundance estimation is difficult given their cryptic and mobile nature, especially when the population is small and has a transnational distribution. In the Baltic Sea, the number of harbour porpoises (<i>Phocoena phocoena</i>) has collapsed since the mid-20<sup>th</sup> century and the Baltic Proper harbour porpoise is listed as Critically Endangered by the IUCN and HELCOM; however, its abundance remains unknown. Here, one of the largest ever passive acoustic monitoring studies was carried out by eight Baltic Sea nations to estimate the abundance of the Baltic Proper harbour porpoise for the first time. By logging porpoise echolocation signals at 298 stations during May 2011-April 2013, calibrating the loggers' spatial detection performance at sea, and measuring the click rate of tagged individuals, we estimated an abundance of 71-1,105 individuals (95% CI, point estimate 491) during May-October within the population's proposed management border. The small abundance estimate strongly supports that the Baltic Proper harbour porpoise is facing an extremely high risk of extinction, and highlights the need for immediate and efficient conservation actions through international cooperation. It also provides a starting point in monitoring the trend of the population abundance to evaluate the effectiveness of management measures and determine its interactions with the larger neighbouring Belt Sea population. Further, we offer evidence that design-based passive acoustic monitoring can generate reliable estimates of the abundance of rare and cryptic animal populations across large spatial scales.</p>
Рис. 1. Вероятность обнаружения меченых животных (среΑнее ± ошибка) при пяти- и Αесятиметровых интерваΛах межΑу прикормочными станциями в Αвух экспериментах. По второму эксперименту расчеты сΑеΛаны ΑΛя резуΛьтатов отΛова в течение первых трех и поΛных Αесяти Αней. Значение «p» отражает уровень статистической значимости разΛичий межΑу ΑоΛями животных с меткой при Αвух интерваΛах Fig. 1. Probability of finding marked animals (average±standard error) between feeding stations placed at intervals of five and ten meters in the two experiments. In the second experiment, calculations were made for the results of trapping during the first three days and during the whole period of ten days. The p value reflects the statistical significance of differences between the fractions of animals with a mark for two types of intervals in Verification of the bottle-based method for estimating abundance of small mammals using biomarkers
Рис. 1. Вероятность обнаружения меченых животных (среΑнее ± ошибка) при пяти- и Αесятиметровых интерваΛах межΑу прикормочными станциями в Αвух экспериментах. По второму эксперименту расчеты сΑеΛаны ΑΛя резуΛьтатов отΛова в течение первых трех и поΛных Αесяти Αней. Значение «p» отражает уровень статистической значимости разΛичий межΑу ΑоΛями животных с меткой при Αвух интерваΛах Fig. 1. Probability of finding marked animals (average±standard error) between feeding stations placed at intervals of five and ten meters in the two experiments. In the second experiment, calculations were made for the results of trapping during the first three days and during the whole period of ten days. The p value reflects the statistical significance of differences between the fractions of animals with a mark for two types of intervals
Coccolithophore Abundance, Size, Carbon And Distribution Estimates (CASCADE)
<div>CASCADE is a global dataset for 139 extant coccolithophore taxonomic units. CASCADE includes a trait database (size and cellular organic and inorganic carbon contents) and taxonomic-specific global spatiotemporal distributions (Lat/Lon/Depth/Month/Year) of coccolithophore abundance and organic and inorganic carbon stocks. CASCADE covers all ocean basins over the upper 275 meters, spans the years 1964-2019 and includes 33,119 taxonomic-specific abundance observations. Within CASCADE, we characterise the underlying uncertainties due to measurement errors by propagating error estimates between the different studies.</div> <div> </div> <div>Full details of the data set are provided in the associated Scientific Data manuscript. The repository contains five main folders: 1) "Classification", which contains YAML files with synonyms, family-level classifications, and life cycle phase associations and definitions; 2) "Concatenated literature", which contains the merged datasets of size, PIC and POC and which were corrected for taxonomic unit synonyms; 3) "Resampled cellular datasets", which contains the resampled datasets of size, PIC and POC in long format as well as a summary table; 4) "Gridded data sets", which contains gridded datasets of abundance, PIC and POC; 5) "Species lists", which contains spreadsheets of the "common" (>20 obs) and "rare" (<20 obs) species and their number of observations.</div> <div> </div> <div>The CASCADE data set can be easily reproduced using the scripts and data provided in the associated github repository: <a title="Opens in new tab" href="https://github.com/nanophyto/CASCADE/tree/v0.1.1" target="_blank" rel="noopener">https://github.com/nanophyto/CASCADE/</a> (<a href="../doi/10.5281/zenodo.12797197">zenodo.12797197</a>)</div> <div> <p>Correspondence to: Joost de Vries, joost.devries@bristol.ac.uk</p> <p>v.3 fixes:</p> <ol> <li><em>actually</em> fixes Guerreiro et al., 2023 longitude issue</li> <li>replaces negative PIC predictions from GLM with NA (an issue with some small HOL species)</li> </ol> <p>v.2 fixes:</p> <p>1. The wrongly specified <em>S. neapolitana</em> was removed from synonyms.yml (this species is now<em> S. nana</em>)<br>2. Longitudes were corrected for Guerreiro et al., 2023<br>3. A double entry for Dimizia et al., 2015 was fixed<br>4. Units in Sal et al., 2013 were correct to cells/L (previously cells/ml)<br>5. Data from Sal et al., 2013 was re-done, as some species were missing<br>6. Duplicate entries from Baumann et al., 2000 were dropped</p> </div>
Simulated reads for benchmarking SARS-CoV-2 lineage abundance estimation
<p>To evaluate the accuracy of lineage abundance estimates from amplicon-based and whole genome-based sequencing, we simulated paired-end reads from amplicons determined by AmpliDiff, and reads spanning full genomes. Abundances of lineages are based on the relative abundance of a lineage within the dataset. The data consists of the following 8 independent datasets:</p> <ul> <li>200 bp reads from the Netherlands based on AmpliDiff amplicons (1, 2, 5 or 10 amplicons) at 1000x coverage,</li> <li>400 bp reads from the Netherlands based on AmpliDiff amplicons (1, 2, 5 or 10 amplicons) at 1000x coverage,</li> <li>200 bp reads from the Netherlands based on whole genome sequencing at 100x coverage,</li> <li>400 bp reads from the Netherlands based on whole genome sequencing at 100x coverage,</li> <li>200 bp reads from Texas based on AmpliDiff amplicons (1, 2, 5 or 10 amplicons) at 1000x coverage,</li> <li>400 bp reads from Texas based on AmpliDiff amplicons (1, 2, 5 or 10 amplicons) at 1000x coverage,</li> <li>200 bp reads from Texas based on whole genome sequencing at 100x coverage,</li> <li>400 bp reads from Texas based on whole genome sequencing at 100x coverage.</li> </ul> <p>Every independent dataset contains 20 sets of reads (generated with different random seeds). The genomes used for the Netherlands-based simulations can be obtained via GISAID through accession id <a href="https://doi.org/10.55876/gis8.230825fe">EPI_SET_230825fe</a>, and the genomes used for the Texas-based simulations can be obtained via GISAID through accession id <a href="https://doi.org/10.55876/gis8.230825pe">EPI_SET_230825pe</a>.</p>
A hierarchical dependent double-observer method for estimating waterfowl breeding pairs abundance from helicopters
Open the record for dataset details and reuse information.
Integrating presence-only and detection/non-detection data to estimate distributions and expected abundance of difficult-to-monitor species on a landscape-scale
Open the record for dataset details and reuse information.
Genotyping by sequencing for estimating relative abundances of diatom taxa in mock communities
Open the record for dataset details and reuse information.
Estimating the abundance of the critically endangered Baltic Proper harbour porpoise (Phocoena phocoena) population using passive acoustic monitoring
Open the record for dataset details and reuse information.
Improving bird abundance estimates in harvested forests with retention by limiting detection radius through sound truncation
Open the record for dataset details and reuse information.
A new Double Observer based census framework to improve abundance estimations in mountain ungulates and other gregarious species with a reduced effort
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.