Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
50
datasets available to search
ShareScore release 0.9.0
Dataset results
50 results for “data issues”
Replication files for "The role of actors' issue and sector specialization for policy integration in the parliamentary arena: An analysis of Swiss biodiversity policy using text as data"
<p>The ZIP file contains all data and code to replicate the analyses reported in the following paper.</p> <p>Reber, U., Ingold, K., & Fischer, M. (2023). The role of actors' issue and sector specialization for policy integration in the parliamentary arena: An analysis of Swiss biodiversity policy using text as data. <em>Policy Sciences</em>. <a href="https://doi.org/10.1007/s11077-022-09490-2">https://doi.org/10.1007/s11077-022-09490-2</a></p> <p>If you use any of the material included in this repository, please refer to the paper.</p>
Over and Under Sampled Data-sets of Code Issues in Java Open-Source Projects
<p>The dataset comprises code changes made to 15 Java Open-Source projects, classified with sentiment values (0 for negative and 1 for positive) based on developer reviews during various revision submissions. The dataset is available in 8 versions, each containing a sampled dataset using an over or under-sampling technique.</p>
Connecting digital citizen science data quality issue to solution mechanism table
<p>A table explaining how to solve data quality issues in digital citizen science. A total of 35 issues and 64 mechanisms to solve them are proposed</p>
Historical Issue Data of Projects on Jira
<p>This dataset contains all the historical issues from sixty-six Apache Open Source projects collected in February 2018. It was used in the paper “Traceability Network Analysis: A Case Study of Links In Issue Tracking Systems” published in the Seventh International Workshop on Artificial Intelligence and Requirements Engineering (AIRE'20) to study issue traceability as a network structure.</p>
Data for "A Phaseless Auxiliary-Field Quantum Monte Carlo Perspective on the UniformElectron Gas at Finite Temperatures: Issues, Observations, and Benchmark Study"
<p>Phaseless AFQMC data (input and output) for "A Phaseless Auxiliary-Field Quantum Monte Carlo Perspective on the UniformElectron Gas at Finite Temperatures: Issues, Observations, and Benchmark Study" </p> <p> </p> <p>data: contains raw qmc data</p> <p>figures: contains analysed data + plotting scripts.</p>
Data from: A burning issue: Savanna fire management can generate enough carbon revenue to help restore Africa's rangelands and fill protected area funding gaps
<p>Many savanna-dependent species in Africa including large herbivores and apex predators are at increasing risk of extinction. Achieving effective management of protected areas (PAs) in Africa where lions live will cost an estimated USD >$1-2 B/year in new funding. We explored the potential for fire management-based carbon-financing programs to fill this funding gap and benefit degrading savanna ecosystems. We demonstrated how introducing early dry season fire management programs could produce potential carbon revenues (PCR) from either a single carbon-financing method (avoided emissions) or from multiple sequestration methods ranging from USD $59.6-$655.9 M/year (at USD $5/ton) or USD $155.0 M–$1.7 B/year (at USD $13/ton). We highlighted variable but significant PCR for savanna PAs from USD $1.5–$44.4 M/year per PA. We suggest investing in fire management programs to jump-start the United Nations Decade of Ecological Restoration to help restore degraded African savannas and conserve imperiled keystone herbivores and apex predators. </p>
Supplementary data for the publication of Characterization of Emissions in Fab Labs: an Additive Manu-facturing Environment Issue
<p>Datasets for the publication of the article "Characterization of Emissions in Fab Labs: an Additive Manufacturing Environment Issue":</p> <p>- Ultrafine Particles: UFP per Zone and mode;</p> <p>- VOC emissions: VOC per Zone and mode.</p>
Detection of the Fire Drill anti-pattern: 15 real-world projects with ground truth, issue-tracking data, source code density, models and code
<p>This package contains artifacts for <strong>15</strong> real-world software projects. The data is supposed to aid the detection of the presence of the Fire Drill anti-pattern. We include original data, ground truth, code (experimental setups and models), and notebooks. The data supports two distinct methods of detecting the AP: a) through issue-tracking data, and b) through the underlying source code. This version of the dataset corresponds to <strong>v8</strong> of the <a href="https://arxiv.org/abs/2104.15090v8">technical report</a> and the <a href="https://github.com/MrShoenel/anti-pattern-models/releases/tag/arxiv-v8">GitHub repository</a>. The package includes the following:</p> <p>Original data:</p> <ul> <li>For each project, its <strong>original</strong> artifacts (e.g., wikis, meeting minutes, mentor's notes, etc.)</li> <li>Evaluation of raters' notes by the assessor</li> </ul> <p>Fire Drill in issue-tracking data:</p> <ul> <li><strong>Ground truth</strong> for whether and how strong each project exhibits the Fire Drill AP, on a scale from [0,10]. This was determined by two individual raters, who also reached a consensus.</li> <li>Coefficients for indicators for the first method, per project.</li> <li>Detailed issue-tracing data for each project: what occurred and when.</li> <li>Time logs for each project.</li> </ul> <p>Fire Drill in source-code data:</p> <ul> <li><strong>Four</strong> technical reports that document the developed method of how to translate a description into a detectable pattern, and to use the pattern to detect the presence and to score it (similar to the rating). Also includes a report for how activities were assigned to individual commits.</li> <li>Source code density data (metrics) for each commit in each of the nine projects as a separate dataset.</li> <li>Code: a snapshot of the repository that holds all code, models, notebooks, and pre-computed results, for utmost reproducibility (the code is written in R).</li> </ul>
Data for articles in "Historical Network Analysis in the Study of Chinese Religion" (Special issue Religions 2023)
<p>This is the data for the seven articles collected in the special issue of <i>Religions</i> (2023) "Historical Network Analysis in the Study of Chinese Religion":</p><p> - Bingenheimer, Marcus. 2023. "Miyun Yuanwu 密雲圓悟 (1567–1642) and His Impact on 17th-Century Buddhism" Religions 14, no. 2: 248. https://doi.org/10.3390/rel14020248</p><p>- Chen, Song. 2023. "Patterns of Integration: A Network Perspective on Popular Religious Connections in China's Lower Yangzi, 1150–1350" Religions 14, no. 5: 577. https://doi.org/10.3390/rel14050577</p><p>- Chu, Ming-Kin. 2023. "Realizing the "Outwardly Regal" Vision in the Midst of Political Inactivity: A Study of the Epistolary Networks of Li Gang 李綱 (1083–1140) and Sun Di 孫覿 (1081–1169)" Religions 14, no. 3: 389. https://doi.org/10.3390/rel14030389</p><p>- Goossaert, Vincent. 2023. "The Social Networks of Gods in Late Imperial Spirit-Writing Altars" Religions 14, no. 2: 217. https://doi.org/10.3390/rel14020217</p><p>- Nehrdich, Sebastian. 2023. "Observations on the Intertextuality of Selected Abhidharma Texts Preserved in Chinese Translation" Religions 14, no. 7: 911. https://doi.org/10.3390/rel14070911</p><p>- Sokolova, Anna. 2023. "Regional Buddhist Communities in Tang China and Their Social Networks: The Network of Master Fayun (?–766)" Religions 14, no. 3: 335. https://doi.org/10.3390/rel14030335</p><p>- Van Cutsem, Laurent. 2023. "Lineages as Network: A Study of Chan Genealogy in the Zutang ji 祖堂集 Using Social Network Analysis" Religions 14, no. 2: 205. https://doi.org/10.3390/rel14020205</p>
Data from: A burning issue: Savanna fire management can generate enough carbon revenue to help restore Africa’s rangelands and fill protected area funding gaps
Open the record for dataset details and reuse information.
Data from: A thorny issue: woody plant defence and growth in an East African savanna
Open the record for dataset details and reuse information.
Data from: Increasing belief but issue fatigue: changes in Australian Household Climate Change Segments between 2011 and 2016
Open the record for dataset details and reuse information.
Data from: A qualitative study exploring the health literacy issues in the care of Chinese American immigrants with diabetes
Objectives: To investigate why first-generation Chinese immigrants with diabetes have difficulty obtaining, processing and understanding diabetes related information despite the existence of translated materials and translators. Design: This qualitative study employed purposive sampling. Six focus groups and two individual interviews were conducted. Each group discussion lasted approximately 90 min and was guided by semistructured and open-ended questions. Setting: Data were collected in two community health centres and one elderly retirement village in Los Angeles, California. Participants: 29 Chinese immigrants aged ≥45 years and diagnosed with type 2 diabetes for at least 1 year. Results: Eight key themes were found to potentially affect Chinese immigrants' capacity to obtain, communicate, process and understand diabetes related health information and consequently alter their decision making in self-care. Among the themes, three major categories emerged: cultural factors, structural barriers, and personal barriers. Conclusions: Findings highlight the importance of cultural sensitivity when working with first-generation Chinese immigrants with diabetes. Implications for health professionals, local community centres and other potential service providers are discussed.
Data from: Interpreting ELISA analyses from wild animal samples: some recurrent issues and solutions
1. Many studies in disease and immunological ecology rely on the use of assays that quantify the amount of specific antibodies (immunoglobulin) in samples. Enzyme-Linked Immuno Sorbent Assays (ELISAs) are increasingly used in ecology due to their availability for a broad array of antigens and the limited amount of sampling material they require. Two recurrent methodological issues are nevertheless faced by researchers: (i) the limited availability of immunological assays and reagents developed for non-model species, and (ii) the statistical determination of the cut-off threshold used to distinguish individual samples that are likely to have or not to have antibodies against a specific antigen. 2. Here, we outline two solutions to deal with these issues. First, we show that implementing two assays with differing detection methods can help validate the use of reagents, such as antibodies, in species different from their intended target. We illustrate this by comparing the quantification of specific vaccinal antibodies against Newcastle Disease Virus (NDV) using two ELISA approaches in four seabird species (Cory's shearwater, European shag, European storm petrel, and Southern rockhopper penguin). 3. Second, we provide a simple way to determine from the distribution of ELISA values whether the assayed samples are likely to be made of a single group of individuals (likely negative) or of two groups of individuals (negative and positive). We illustrate the use of this approach with two independent datasets: NDV antibody levels following vaccination and anti-Borrelia antibody levels following natural exposure. 4. The practical implementation of these methodological approaches could provide a way to efficiently apply ELISAs and other immune-based assays to address questions in the growing fields of ecological immunology and disease ecology.
Data from: Sharing is caring? measurement error and the issues arising from combining 3D morphometric datasets
Geometric morphometrics is routinely used in ecology and evolution and morphometric datasets are increasingly shared among researchers, allowing for more comprehensive studies and higher statistical power (as a consequence of increased sample size). However, sharing of morphometric data opens up the question of how much nonbiologically relevant variation (i.e., measurement error) is introduced in the resulting datasets and how this variation affects analyses. We perform a set of analyses based on an empirical 3D geometric morphometric dataset. In particular, we quantify the amount of error associated with combining data from multiple devices and digitized by multiple operators and test for the presence of bias. We also extend these analyses to a dataset obtained with a recently developed automated method, which does not require human-digitized landmarks. Further, we analyze how measurement error affects estimates of phylogenetic signal and how its effect compares with the effect of phylogenetic uncertainty. We show that measurement error can be substantial when combining surface models produced by different devices and even more among landmarks digitized by different operators. We also document the presence of small, but significant, amounts of nonrandom error (i.e., bias). Measurement error is heavily reduced by excluding landmarks that are difficult to digitize. The automated method we tested had low levels of error, if used in combination with a procedure for dimensionality reduction. Estimates of phylogenetic signal can be more affected by measurement error than by phylogenetic uncertainty. Our results generally highlight the importance of landmark choice and the usefulness of estimating measurement error. Further, measurement error may limit comparisons of estimates of phylogenetic signal across studies if these have been performed using different devices or by different operators. Finally, we also show how widely held assumptions do not always hold true, particularly that measurement error affects inference more at a shallower phylogenetic scale and that automated methods perform worse than human digitization.
Figure 3 from: Thessen A, Patterson D (2011) Data issues in the life sciences. ZooKeys 150: 15-51. https://doi.org/10.3897/zookeys.150.1766
Figure 3 - Technical infrastructure needed for Big New Biology to fully emerge (based on Sinha et al. 2010).
Figure 2 from: Thessen A, Patterson D (2011) Data issues in the life sciences. ZooKeys 150: 15-51. https://doi.org/10.3897/zookeys.150.1766
Figure 2 - A Big New Biology can only emerge with a framework that optimizes reuse. Ideally, data should be in forms that can flow from source into a common pool and can flow back out to consumers, be subject to quality control, or be enhanced through analysis to rejoin the pool as processed data.
Figure 1 from: Thessen A, Patterson D (2011) Data issues in the life sciences. ZooKeys 150: 15-51. https://doi.org/10.3897/zookeys.150.1766
Figure 1 - Rogers adoption curve describes the acceptance of a new technology. Life Sciences is still in the Early Adopters phase for accepting principles of data readiness.
Data from: Interpreting ELISA analyses from wild animal samples: some recurrent issues and solutions
Open the record for dataset details and reuse information.
Data from: A qualitative study exploring the health literacy issues in the care of Chinese American immigrants with diabetes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.