Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
210
datasets available to search
ShareScore release 0.9.0
Dataset results
210 results for “citizen data”
Detecting the effect of intensive agriculture on Odonata diversity using citizen science data
<p>These datasets are used in the following article :</p> <p><strong><span>Detecting the effect of intensive agriculture on Odonata diversity using citizen science data</span></strong></p> <p><strong><span><span>By </span></span></strong><span>Renaud Baeta<sup>1</sup>, Justine Léauté<sup>1,2</sup>, Éric Sansault<sup>1</sup> and Sylvain Pincebourde<sup>2*</sup></span></p> <p><span><strong>In </strong></span><em><span>Ecological Applications </span></em><span>(Research article)</span></p> <p><strong><span>Abstract</span></strong></p> <p><span><span><span>Agricultural areas represent one of the major ecosystems of the world. Intensification of agricultural practices produced openfields characterized by low biological diversity. Nevertheless, the distance up to which intensive agricultural fields alter surrounding natural systems is rarely quantified. We determined the spatial scale at which agricultural landscapes alter the diversity of Odonates, a key taxon in wetland ponds, and we tested to what extent citizen-science data can be used reliably for this purpose. We compiled 7,731 observations made in a portion of the region Centre-Val-de-Loire (France) over 10 years by naturalists on 729 water bodies to analyze the effect of agricultural landscapes (mainly wheat, rapeseed, sunflower) on the species richness of both damselflies and dragonflies in lentic systems. Sixty species were reported over the 10-years period. For dragonflies, intensive agricultural landscapes best explained their richness at the scales of 800m and 1,600m for overall and autochthonous species, respectively, when using the full dataset. The spatial scale was smaller for damselflies, at 200m for both overall and autochthonous species. These distances were not severely impacted when constraining the data to consider several biases. Multi-model averaging showed that the proportion of intensive agriculture decreased species richness, despite the potential biases inherent to an imperfect database acquired by citizens. This imperfect citizen dataset allows to infer the lowest effect size of agriculture on species richness. Quantitatively, this effect was more important for autochthonous species. Interestingly, both relatively rare taxa and common or generalist species can be under threat in intensive agricultural landscapes, calling for more ecotoxicological studies. The influence of agricultural practices from a distance implies that conservation and management plans of wetland ponds should consider the landscape ecological characteristics and not only the pond features. Conservation efforts focusing too locally on a site may be undermined because intensive agriculture from a distance limits the potential for the site to recover highly diverse communities. These distant effects should be integrated by policy-makers when deciding which wetland pond should benefit from a conservation plan or which conservation action may be planned, implementing for instance buffer zones and/or ecological corridors composed of natural vegetation. </span></span></span></p>
Data from: Processing citizen science- and machine-annotated time-lapse imagery for biologically meaningful metrics
Time-lapse cameras facilitate remote and high-resolution monitoring of wild animal and plant communities, but the image data produced require further processing to be useful. Here we publish pipelines to process raw time-lapse imagery, resulting in count data (number of penguins per image) and 'nearest neighbour distance' measurements. The latter provide useful summaries of colony spatial structure (which can indicate phenological stage) and can be used to detect movement – metrics which could be valuable for a number of different monitoring scenarios, including image capture during aerial surveys. We present two alternative pathways for producing counts: 1) via the Zooniverse citizen science project Penguin Watch and 2) via a computer vision algorithm (Pengbot), and share a comparison of citizen science-, machine learning-, and expert- derived counts. We provide example files for 14 Penguin Watch cameras, generated from 63,070 raw images annotated by 50,445 volunteers. We encourage the use of this large open-source dataset, and the associated processing methodologies, for both ecological studies and continued machine learning and computer vision development.
GLOBE Mosquito Habitat Mapper Citizen Science Data 2017-2020
<p>Three Cases: Metadata and Procedures</p> <p>The data sets described here were used in an article submitted to the journal GeoHealth in 2021. The data files and further supplemental links (including general information about GLOBE data) can be accessed at https://observer.globe.gov/get-data/mosquito-habitat-data.</p> <p>Case 1: Removal of records with suspect geolocation data. A Python script was applied to remove records where the measured position (in decimal degrees) was identical to the GLOBE MGRS site position. GPS-obtained latitude and longitude coordinates are reported in decimal degrees, so records identified by whole numbers were also removed. This procedure removed 5704 (23%) of the 24983 records in the Mosquito Habitat Mapper database, with 19,279 records remaining. The secondary data sets cleaned only for geolocation anomalies were labeled Case 1.</p> <p>Case 2: Identifying suspected training events. For this test, we sought to identify groups of data that exceeded 10 records sharing these characteristics. Another Python script was employed to extract the photos for ease of visual inspection. Because we needed to manually review the photo records, we set the threshold for groups at >10, so that the analysis could be completed in the time allotted. Groups identified thought this procedure were outputted as case 2: groups. The resulting data set cleaned of groups >10 was labeled Case 2. The resulting data set included 20,006 records and identified 2,447 records found in clusters we postulated were training events.</p> <p>Case 3: The Case 3 secondary dataset result from applying the Python scripts used to create Cases 1 and 2. We used the Case 3 data sets, with improved geolocation and large groups eliminated, in the following analysis.</p> <p>Acknowledgments: These data were obtained from NASA and the GLOBE Program and are freely available for use in research, publications and commercial applications. When data from GLOBE are used in a publication, we request this acknowledgment be included: "These data were obtained from the GLOBE Program." Please include such statements, either where the use of the data or other resource is described, or within the acknowledgments section of the publication.</p>
Data from: Assessing the usefulness of Citizen Science Data for habitat suitability modelling: opportunistic reporting versus sampling based on a systematic protocol
<p><strong>Aim:</strong> To evaluate the potential of models based on opportunistic reporting (OR) compared to models based on data from a systematic protocol (SP) for modelling species distributions. We compared model performance for eight forest bird species with contrasting spatial distributions, habitat requirements, and rarity. Differences in the reporting of species were also assessed. Finally, we tested potential improvement of models when inferring high quality absences from OR based on questionnaires sent to observers.</p> <p><strong>Location:</strong> Both datasets cover the same large area (Sweden) and time period (2000 -2013).</p> <p><strong>Methods:</strong> Species distributions were modelled using logistic regression. Predictive performance of OR models to predict SP data were assessed based on AUC. We quantified the congruence in spatial predictions using Spearman's rank correlation coefficient. We related these results to species characteristics and reporting behaviour of observers. We also assessed the gain in predictive performance of OR models by adding inferred absences. Finally, we investigated the potential impact of sampling bias in OR.</p> <p><strong>Results:</strong> For all species, and despite the sampling biases, results from OR overall agreed well with those of SP, for the nationwide spatial congruence of habitat suitability maps and the selection and directions of species-environment relationships. The OR models also performed well in predicting the SP data. The predictive performance of the OR models increased with species rarity and even outperformed the SP model for the rarest species. No significant impact of observer behaviour was found.</p> <p><strong>Main Conclusions:</strong> Relatively simple analyses with inferred absences could produce reliable spatial predictions of habitat suitability. This was especially true for rare species. OR data should be seen as a complement to SP, as the weakness of one is the strength of the other, and OR may be especially useful at large spatial scales or where no systematic data collection protocols exist.</p>
Data for Ferretto et al, 2021 "Naturally-detached fragments of the endangered seagrass Posidonia australis collected by citizen scientists can be used to successfully restore fragmented meadows"
<p>Please find attached the data for the manuscript "Ferretto et al, 2021" and a brief description of each file.</p>
Students as citizen scientists: project-based learning through the iNaturalist platform could provide useful biodiversity data
<p>This is a compilation of data from iNaturalist platform based on the geographic limits of brazilian semiarid until 28 october, 2022.</p>
Data from: Hazard and catch composition of ghost fishing gear revealed by a citizen science clean-up initiative
<p><span>Ghost fishing, the continued catch of fishes and invertebrates by lost fishing gear, represents an animal welfare issue as well as a waste of both potential food and ecosystem resources. Fishing gear is lost by both commercial and recreational fishers, and management authorities often lack an overview of gear loss and subsequently potential impact on coastal populations. </span><span>To investigate the hazard and catch composition of lost fishing gear along the Norwegian coast</span><span>, recreational divers in collaboration with scientists conducted systematic reporting of retrieved lost fishing gear. </span><span>Through this citizen science project,</span><span> a total of 12,101 gear items were retrieved and reported, including traps, gillnets and fyke nets. Combining both data on the catch ratio of the gear and its relative quantity, we identified the five most hazardous gear types to be parlor traps, gillnets, fyke nets, wrasse traps and square collapsible traps. The parlour trap was the most hazardous trap, due to high catchability and quantity. The correct classification of gear type could not be confirmed in 2.8 – 6.1 % of the pictures taken by divers, depending on reporting format, and divers reported the wrong gear type in 1.4 % of the reports. Brown crab (<em>Cancer</em> <em>pagurus</em>) was the species most often found in retrieved gear. Furthermore, the vulnerable species European lobster (<em>Homarus</em> <em>gammarus</em>) and Atlantic cod (<em>Gadus</em> <em>morhua</em>) were also common. These results can inform future clean up-initiatives and management responses to ghost fishing, including preventive measures against gear loss and gear restrictions and customization. </span></p>
Citizen science data on the presence of invasive mosquitoes in Hungary
<ol> <li><span>Climate change, intensified tourism and trade activity result in several exotic mosquito species invading the temperate zone, which has considerable ecological and economic consequences and threatens human health due to the pathogen-transmitting role of these organisms. Accordingly, three invasive mosquito species (<em>Aedes</em> <em>albopictus</em>, <em>Ae. japonicus</em>, and <em>Ae. koreicus</em>) have been described in the last decade in Hungary, a Central European country. It is crucial to understand how invasive species are introduced and their distribution is expanded at the country-level, for which intense surveillance programs are needed.</span></li> <li><span>We have established a citizen science program, in which we asked the public to submit reports on their observations of invasive mosquitoes. During a three-year campaign, we have collected and taxonomically validated about 3,000 reports that can be arranged along both the temporal and spatial scales. We aggregated these observations into 35 km<sup>2</sup> quadrats and examined if these can be reliably used for scientific inferences.</span></li> <li><span>We first found that the number of validated reports in a quadrat depends on the underlying sampling effort (i.e., number of total reports), but this relationship varies among species and study years. Second, after controlling for study effort, we showed that the prevalence and presence/absence of invasive mosquitoes within quadrats are significantly repeatable among years, but this consistency varies in a species-specific way. Third, we demonstrated that conclusions about the local presence/absence of focal species based on citizen reports corroborate well the results of direct field sampling with conventional trapping protocols. </span></li> <li> <em><span>Synthesis and applications.</span></em><span> We suggest that if the reporting intensity is appropriate (i.e., the number of reports reaches a species-specific threshold), citizen science results can be used to derive biologically meaningful conclusions about the distribution of invasive mosquitoes in a country. Distribution maps of the three invasive species in Hungary can be used to identify ecological predictors that determine such spatial patterns and also to develop a mosquito control program and assess epidemiological risk. </span> </li> </ol>
Baseline data for Citizen Energy Responsible Behaviour in 27 Member States
<p>This dataset contains all the information on the levels that each country in Europe should have in the labelling system to assess citizens energy behaviours fully described in the document AURORA D1.1 Near-Zero Emissions Citizens Label https://doi.org/10.5281/zenodo.7594879</p> <p>The zip file contains one document for each member state. Reference values are extracted from the same datasources individuated in D1.1 for the 5 selected countries fully described in the document.</p>
Source data for the scientometric analysis of citizen science research publications
<p>Source data for the scientometric comparison of citizen science research (n=5749 documents) and a semi-random sample of publications (n=5734) retrieved from the Web of Science Core Collection and published between 1997-2021. The data include information on: author(s) full name(s) ("AF"); year of publication ("PY"); title ("TI"); Digital Object Identifier ("DI"); abstract ("AB"); open access indicator ("OA"); author(s) affiliation(s) ("C1"); document type ("DT"); funding entity ("FU"); funding text ("FX"); total number of citations ("TC"); and identifier of the collection ("group"): CS for citizen science and SRS for the semi-random sample collection respectively.</p>
Data from: Responses of sympatric canids to human development revealed through citizen science
Open the record for dataset details and reuse information.
Data for: Citizen science as an ecosystem of engagement: Implications for learning and broadening participation
Open the record for dataset details and reuse information.
Supporting data from: When birding hotspots get too hot: A geographic evaluation of wildfire-related disturbance on spatiotemporal biases in citizen science data
Open the record for dataset details and reuse information.
Louse flies (Diptera: Hippoboscidae): United Kingdom, Republic of Ireland, and Isle of Man: Citizen science data: Part 2, host associations
Open the record for dataset details and reuse information.
Data from: Seasonal macro-demography of North American bird populations revealed through citizen science monitoring
Open the record for dataset details and reuse information.
Data from CoFish: co-designing citizen science between fishers and scientists to monitor the phosphorus distribution across two Lake Geneva basins
Open the record for dataset details and reuse information.
Data from: Hazard and catch composition of ghost fishing gear revealed by a citizen science clean-up initiative
Open the record for dataset details and reuse information.
Data from: Assessing the usefulness of Citizen Science Data for habitat suitability modelling: opportunistic reporting versus sampling based on a systematic protocol
Open the record for dataset details and reuse information.
Data from: Processing citizen science- and machine-annotated time-lapse imagery for biologically meaningful metrics
Open the record for dataset details and reuse information.
Data from:Using a large citizen science dataset to uncover diverse patterns of elevational migration in Himalayan birds
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.