Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
87
datasets available to search
ShareScore release 0.9.0
Dataset results
87 results for “community science”
Data for: Community science reveals high diversity of nectaring plants visited by painted lady butterflies (Lepidoptera: Nymphalidae) in California sage scrub
Open the record for dataset details and reuse information.
Community science validates climate suitability projections from ecological niche modeling
<p><span>Climate change poses an intensifying threat to many bird species, and projections of future climate suitability provide insight into how species may shift their distributions in response. Climate suitability is characterized using ecological niche models (ENMs), which correlate species occurrence data with current environmental covariates and project future distributions using the modeled relationships together with climate predictions. Despite their widespread adoption, ENMs rely on several assumptions that are rarely validated <i>in situ </i>and can be highly sensitive to modeling decisions, precluding their reliability in conservation decision-making. Using data from a novel, large-scale community science program, we developed dynamic occupancy models to validate near-term climate suitability projections for bluebirds and nuthatches in summer and winter. We estimated occupancy, colonization, and extinction dynamics across species' ranges in the United States in relation to projected climate suitability in the 2020s, and used a Gibbs variable selection approach to quantify evidence of species-climate relationships. We also included a Bird Conservation Region strata-level random effect to examine among-strata variation in occupancy that may be attributable to land-use and ecoregional differences. Across species and seasons, we found strong evidence that initial occupancy and colonization were positively related to 2020 climate suitability, illustrating an independent validation of projections from ENMs across a large geographic area. </span><span>Random strata effects revealed that occupancy probabilities were generally higher than average in core areas and lower than average in peripheral areas of species' ranges, and served as a first step in identifying spatial patterns of occupancy from these community science data. </span><span>Our findings lend much-needed support to the use of ENM projections for addressing questions about potential climate-induced changes in species' occupancy dynamics. More broadly, </span>our work highlights the value of community scientist observations for ground-truthing projections from statistical models and for refining our understanding of the processes shaping species' distributions under a changing climate.</p>
Quantifying how natural history traits contribute to bias in community science engagement: a case study using native and introduced orb weaver spiders in North America
<p>Raw data and R scripts for manuscript.</p>
Butterflies at porch lights: exploring nocturnal light visitation in butterflies using community science data
<p>Data and R code for manuscript.</p>
Smell Pittsburgh: Engaging Community Citizen Science for Air Quality
<p>Link to the files and description of the Smell Pittsburgh Dataset –<br> <a href="https://eur04.safelinks.protection.outlook.com/?url=https%3A%2F%2Fgithub.com%2FCMU-CREATE-Lab%2Fsmell-pittsburgh-prediction%2Ftree%2Fmaster%2Fdataset%2Fv2&data=05%7C01%7Cy.c.hsu%40uva.nl%7C89562067341d40d0bad308da2c475652%7Ca0f1cacd618c4403b94576fb3d6874e5%7C0%7C0%7C637870982141190827%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=gWn0nGRUl5EAHvDfJwTlTNJN%2BWFH62NX6Mw%2B2web6XE%3D&reserved=0">https://github.com/CMU-CREATE-Lab/smell-pittsburgh-prediction/tree/master/dataset/v2</a></p> <p>Smell Pittsburgh (<a href="https://smellpgh.org">https://smellpgh.org</a>) is a mobile application for crowdsourcing reports of bad odors, such as those generated from air pollution. The data is used to train a machine learning model to predict the presence of bad smell and create push notifications to inform citizens about the bad smell. The motivation, background, and design of the Smell Pittsburgh application is described in the following paper.</p> <ul> <li>Yen-Chia Hsu, Jennifer Cross, Paul Dille, Michael Tasota, Beatrice Dias, Randy Sargent, Ting-Hao (Kenneth) Huang, and Illah Nourbakhsh. 2020. Smell Pittsburgh: Engaging Community Citizen Science for Air Quality. ACM Transactions on Interactive Intelligent Systems. 10, 4, Article 32. DOI:<a href="https://doi.org/10.1145/3369397">https://doi.org/10.1145/3369397</a>. Preprint:<a href="https://arxiv.org/pdf/1912.11936.pdf">https://arxiv.org/pdf/1912.11936.pdf</a>.</li> </ul>
Fig. 5. Termite SDM predictions for species located within Southern Australia for A in Utilization of Community Science Data to Explore Habitat Suitability of Basal Termite Genera
Fig. 5. Termite SDM predictions for species located within Southern Australia for A. Porotermes adamsoni (Stolotermitidae)(left) and B. Stolotermes victoriensis (right), and C. Mastotermes darwiniensis (Mastotermitidae). Final model predictions were generated using our thinned occurrence dataset and final set of uncorrelated environmental variables for each species, with 10 bootstrap replicates with 'cloglog' outputs in which raw values are converted to a range of 0 - 1 to approximate a probability of occurrence (Cobos et al. 2018). Brighter colors indicate areas of higher suitability (higher probability of occurrence), while darker colors indicate areas of lower suitability (lower probability of occurrence)..
Fig. 4. Termite SDM predictions for species located within Eastern United States and Canada for A in Utilization of Community Science Data to Explore Habitat Suitability of Basal Termite Genera
Fig. 4. Termite SDM predictions for species located within Eastern United States and Canada for A. Zootermopsis nevadensis (Stolotermitidae)(left) B. Zootermopsis angusticollis (right), and C. Zootermopsis laticeps (bottom). Final model predictions were generated using our thinned occurrence dataset and final set of uncorrelated environmental variables for each species, with 10 bootstrap replicates with 'cloglog' outputs in which raw values are converted to a range of 0 - 1 to approximate a probability of occurrence (Cobos et al. 2019). Brighter colors indicate areas of higher suitability (higher probability of occurrence), while darker colors indicate areas of lower suitability (lower probability of occurrence).
Fig. 3 in Utilization of Community Science Data to Explore Habitat Suitability of Basal Termite Genera
Fig. 3. Termite SDM predictions for species located within Argentina and Chile (top left), New Zealand (top right), and Tazmania (bottom) for A. Porotermes quadricollis (Stolotermitidae) and B. Stolotermes ruficeps, and C. Stolotermes brunneicornis. Final model predictions were generated using our thinned occurrence dataset and final set of uncorrelated environmental variables for each species, with 10 bootstrap replicates with 'cloglog' outputs in which raw values are converted to a range of 0–1 to approximate a probability of occurrence (Cobos et al. 2019). Brighter colors indicate areas of higher suitability (higher probability of occurrence), while darker colors indicate areas of lower suitability (lower probability of occurrence).
Fig. 2 in Utilization of Community Science Data to Explore Habitat Suitability of Basal Termite Genera
Fig. 2. Termite SDM predictions for species located within South Africa (top and bottom), and Argentina (right) for A. Porotermes planiceps (Stolotermitidae) B. Microhodotermes viator (Hodotermitidae), and C. Porotermes quadricollis. Final model predictions were generated using our thinned occurrence dataset and final set of uncorrelated environmental variables for each species, with 10 bootstrap replicates with 'cloglog' outputs in which raw values are converted to a range of 0–1 to approximate a probability of occurrence (Cobos et al. 2019). Brighter shades indicate areas of higher suitability (higher probability of occurrence), while darker shades indicate areas of lower suitability (lower probability of occurrence).
Fig. 1 in Utilization of Community Science Data to Explore Habitat Suitability of Basal Termite Genera
Fig. 1. Summary tree showing current state of termite phylogeny. Simplified schematic based on familial termite relationships recovered with high support in (Engel et al. 2009, Legendre et al. 2015). Stylotermitidae is placed in its current position based on (Bucek et al. 2019). Branches representing unresolved relationships (bootstrap values <75) are indicated by *. The cockroach family Cryptocercidae was used as an outgroup, and soldier illustrations correlate with families used in summary tree. A. Mastotermes (Froggatt 1897, Blattodea Mastotermitidae), B. Zootermopsis angusticollis (Hagen 1858, Blattodea, Archotermopsidae), C. Hodotermopsis sjostedi (Holmgren 1911, Blattodea, Archotermopsidae), D. Anacanothermoes ochraceus (Burmeister 1839, Blattodea, Hodotermitidae), E. Porotermes adamsoni (Froggat 1897, Blattodea, Stolotermitidae), F. Cryptotermes brevis (Walker 1853, Blattodea, Kalotermitidae), G. Stylotermes halumicus (Liang, et al. 2017, Blattodea, Stylotermitidae), H. Coptotermes formosanus (Shiraki 1909, Blattodea, Rhinotermitidae), I. Serritermes serrifer (Hagen and Bates, Blattodea, Serritermitidae), J. Neocapritermes taraqua (Krishna and Araujo 1968, Blattodea,Termitidae), K. Nasutitermes corniger (Motschulsky 1855, Blattodea,Termitidae)
Data from: Bringing ecology blogging into the scientific fold: measuring reach and impact of science community blogs
The popularity of science blogging has increased in recent years, but the number of academic scientists who maintain regular blogs is limited. The role and impact of science communication blogs aimed at general audiences is often discussed, but the value of science community blogs aimed at the academic community has largely been overlooked. Here, we focus on our own experiences as bloggers to argue that science community blogs are valuable to the academic community. We use data from our own blogs (n = 7) to illustrate some of the factors influencing reach and impact of science community blogs. We then discuss the value of blogs as a standalone medium, where rapid communication of scholarly ideas, opinions, and short observational notes can enhance scientific discourse, and discussion of personal experiences can provide indirect mentorship for junior researchers and scientists from underrepresented groups. Finally, we argue that science community blogs can be treated as a primary source and provide some key points to consider when citing blogs in peer-reviewed literature.
Community science butterfly data Northwest Arkansas
<p>This data set contains the butterfly behavior observations of community members who visited the Botanical Gardens of the Ozarks, as well as students in the Principles of Zoology and Animal Behavior courses at the University of Arkansas from spring 2017 to fall 2020. This data was used to assess the relationship between butterfly color and butterfly flower color choice, as well as the relationship between butterfly color, behavior, abundance, and cloud cover in Northwest Arkansas. It corresponds to "Engaging the community in pollinator research: the effect of wing pattern and weather on butterfly behavior" in Integrative and Comparative Biology. This data set includes the original data reported by community observers as well as the cleaned values used in the manuscript.</p>
The hidden influence of communities in collaborative funding of clinical science
<p>Every year the National Institutes of Health allocates $10.7 billion (one-third of its funds) for clinical science research while the pharmaceutical companies spend $52.9 billion (90% of its annual budget). However, we know little about funder collaborations and the impact of collaboratively funded projects. As an initial effort towards this, we examine the cofunding network, where a funder represents a node and an edge signifies collaboration. Our core data include all papers that cite and receive citations by the Cochrane Database of Systemic Reviews, a prominent clinical review journal. We find that 65% of clinical papers have multiple funders and discover communities of funders that are formed by national boundaries and funding objectives. To quantify success in funding, we use a g-index metric that indicates efficiency of funders in supporting clinically relevant research. After controlling for authorship, we find that funders generally achieve higher success when collaborating than when solofunding. We also find that as a funder, seeking multiple, direct connections with various disconnected funders may be more beneficial than being part of a densely interconnected network of co-funders. The results of this paper indicate that collaborations can potentially accelerate innovation, not only among authors but also funders.</p>
The hidden influence of communities in collaborative funding of clinical science
Open the record for dataset details and reuse information.
Data from: Urban environmental predictors of group size in cliff swallows (Petrochelidon pyrrhonota): A test using community-science eBird data
Open the record for dataset details and reuse information.
Data from: Bringing ecology blogging into the scientific fold: measuring reach and impact of science community blogs
Open the record for dataset details and reuse information.
Community science butterfly data Northwest Arkansas
Open the record for dataset details and reuse information.
Does the munch affect the bunch? Using community science to explore insect herbivory and fruit production in an understory plant
Open the record for dataset details and reuse information.
Data from: Community science validates climate suitability projections from ecological niche modeling
Open the record for dataset details and reuse information.
MeadoWatch: a long-term community-science database of wildflower phenology in Mount Rainier National Park
<p>We present a long-term and high-resolution phenological dataset from 17 wildflower species collected in Mt. Rainier National Park, as part of the MeadoWatch (MW) community science project. Since 2013, 500+ unique volunteers and scientists have gathered data on the timing of four key reproductive phenophases (budding, flowering, fruiting, and seeding) in 28 plots over two elevational gradients alongside popular park trails. Trained volunteers (87.2%) and UW scientists (12.8%) collected data 3-9 times/week during the growing season, using a standardized method. Taxonomic assessments were highly consistent between scientists and volunteers, with high accuracy and specificity across phenophases and species. Sensitivity, on the other hand, was lower than accuracy and specificity, suggesting that a few species might be challenging to reliably identify in community-science projects. Up to date, the MW database includes 42,000+ individual phenological observations from 17 species, between 2013 and 2019. However, MW is a living dataset that will be updated through continued contributions by volunteers, and made available for its use by the wider ecological community.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.