Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.9.0
Dataset results
56 results for “false positive”
SQLite database to accompany the paper, "Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring"
<p>This dataset is a SQLite database that accompanies methods and analysis described in the paper, "Statistical learning mitigation of false positives from template-detected data in automated acoustic wildlife monitoring" (Balantic & Donovan 2019, Bioacoustics, https://www.tandfonline.com/doi/full/10.1080/09524622.2019.1605309). </p> <p>A Github repository containing code for using the SQLite database also accompanies this paper at: <a href="https://github.com/cbalantic/false-positive-mitigation">http://github.com/cbalantic/false-positive-mitigation</a></p>
Generalized model-based solutions to false positive error in species detection/non-detection data: DataS5.
<p>Data/code associated with empirical case study (Gray fox relative abundance estimation/prediction) in article "Generalized model-based solutions to false positive error in species detection/non-detection data" [doi pending].</p>
Correcting for population stratification reduces false positive and false negative results in joint analyses of host and pathogen genomes [G2G-Simulator: Simulated dataset]
<p>Data associated with the paper 'Correcting for population stratification reduces false positive and false negative results in joint analyses of host and pathogen genomes'.</p> <p>It contains the raw simulated data from the 'G2G-Simulator' program. </p> <p>Those data need to be loaded in a R environment .</p> <p>You can reproduce plots present in the paper by parsing the R object using the script 'parse_paper_data.R' present in the G2G-Simulator GitHub repository (https://github.com/onaret/G2G-Simulator/paper/parse_paper_dataset.R). </p>
◂Fig. 3 Gynoecium of C. crenata %yellow frames), C. cf. grandicalyx %blue frames) and C. sinensis %pink frames; A, B stack shot images; C–K light microscopy; G polarised light; TS in horizontal orientation). A, B Anthetic female flower, calyx and corolla partly removed. B LS of gynoecium. C LS of functionally female flower %note strongly stained peripheral tissue of corolla, anther and gynoecium). D LS of gynoecium. E, F TS of functionally female flower %note strongly stained, peripheral tissue). G TS of functionally female flower %note crystal deposition). H LS of ovule %note stalked embryo sac). J TS of functionally male flower with non-functional ovules. K LS of functionally male flower %style lacking, original position indicated by an asterisk) %LS, longisection; TS, transverse section; a,anther; bs, basal septum; c, calyx; car, carpel; co, corolla; db, dorsal bundles; es, embryo sac; fs, false septum; lb, lateral bundles; o, ovule; stg, stigma; sty, style; t, trichomes; tt, transmission tissue; ut, peripheral, strongly stained tissue; vb, ventral bundles; vs, ventral slit) in Observations on flower and fruit anatomy in dioecious species of Cordia (Cordiaceae, Boraginales) with evolutionary interpretations
◂Fig. 3 Gynoecium of C. crenata %yellow frames), C. cf. grandicalyx %blue frames) and C. sinensis %pink frames; A, B stack shot images; C–K light microscopy; G polarised light; TS in horizontal orientation). A, B Anthetic female flower, calyx and corolla partly removed. B LS of gynoecium. C LS of functionally female flower %note strongly stained peripheral tissue of corolla, anther and gynoecium). D LS of gynoecium. E, F TS of functionally female flower %note strongly stained, peripheral tissue). G TS of functionally female flower %note crystal deposition). H LS of ovule %note stalked embryo sac). J TS of functionally male flower with non-functional ovules. K LS of functionally male flower %style lacking, original position indicated by an asterisk) %LS, longisection; TS, transverse section; a,anther; bs, basal septum; c, calyx; car, carpel; co, corolla; db, dorsal bundles; es, embryo sac; fs, false septum; lb, lateral bundles; o, ovule; stg, stigma; sty, style; t, trichomes; tt, transmission tissue; ut, peripheral, strongly stained tissue; vb, ventral bundles; vs, ventral slit)
Figure 1 in Strategies for false positive reduction and multimodal lesion characterization in computer-aided diagnosis of breast cancer
Figure 1. - Representative ultrasound images at four-month post copulation (8 Sep. 2010) before resorption, five-month post copulation (20 Oct. 2010) during resorption, and six-month post copulation (3 Nov. 2010) after resorption. A: Uterine horn; B: Fetus; C: Ovary; D: Follicle.
Atlas of false-positive bundles for RecobundlesX
<p>This data is made to be used with the <a href="https://zenodo.org/record/4630660">Atlas for RecobundlesX</a>.</p> <p>In the study performed by <em>MaierHein et al. (2017)</em> about the ISMRM 2015 Tractography Challenge, several bundles from the submitted tractograms were labeled as false positives because they were not in the ground-truth tractogram. Quality control was performed on these bundles to select those with the most anatomical implausible trajectory. From these bundles, 44 were selected and refined to obtain smooth bundle models.</p> <p>v2. Streamlines outliers that were not removed with an automatic pruning step were removed manually.</p> <p><em>Maier-Hein KH, Neher PF, Houde J-C, Côté M-A, Garyfallidis E, Zhong J, et al. 2017. The challenge of mapping the human connectome based on diffusion tractography. Nat Commun 8:1349.</em></p>
Supplemental data for: A bibliometric assessment of the incidence of amyloid-Eszett (Aß), a false positive of amyloid-beta (Aβ), in the neurodegenerative disease literature
<p>One claimed reason for the development of Alzheimer’s disease (AD), a prominent neurodegenerative disease, is the extracellular aggregation of amyloid-beta (Aβ). A linguistic or formatting error has resulted in the misrepresentation of the Greek letter β with the German letter Eszett (ß), resulting in the formation of a non-existent compound, amyloid-Eszett (Aß). These datasets offer a quantified appreciation of the AD-related literature, carrying a mention of this false positive in the title, abstract and keywords of papers indexed in the Web of Science Core Collection and Scopus. Also, as a curiosity given the popularity of this large language model, we asked the questions to ChatGPT. This AI chatbot developed by OpenAI was able to recognize Eszett as a linguistic or typographic error, within this context, recognizing Aß and Aβ as equals. This erroneous substitution of a Greek letter (in Aβ) by a German one (Aß), despite giving a non-existent compound, will likely not change the underlying scientific conclusions of the affected papers, although errata might be useful to enlighten others, including metadata managers and journal copyeditors, so as not to repeat the same mistake.</p>
Lung nodule CT false positive reduction
<p>This data set is part of the public development data for the <a href="http://auc23.grand-challenge.org/">2023 Automated Universal Classification Challenge</a> (AUC23). The data set concerns the classification of lung nodules on screening thoracic computed tomography (CT) scans and was derived from the <a href="https://luna16.grand-challenge.org/Data/">2016 Lung Nodule Analysis (LUNA) challenge</a>. The data set was previously introduced and described by <a href="https://www.sciencedirect.com/science/article/pii/S1361841517301020?via%3Dihub">Setio et al. (2017)</a> and no images or patient information were added. Only the "false positive reduction" track was considered, where a provided set of nodule candidates should be classified. Data was restructured in compliance with the <a href="https://auc23.grand-challenge.org/">AUC23</a> challenge format. The data set was collected from the largest publicly available reference database for lung nodules: <a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=1966254">The Lung Image Database Consortium image collection</a>.</p> <p>Images are 3D tensors:</p> <ul> <li>0: 3D axial screening CT scan (cropped to candidate nodule's region of interest)</li> </ul> <p>Classification labels:</p> <ul> <li>0: Is a nodule</li> <li>1: Is not a nodule</li> </ul> <p>Folder structure:</p> <p>imagesTr (root folder with all patients and studies)<br> ├── LUNA16_0000_0000.mha (axial CT imaging for nodule 0000)<br> ├── LUNA16_0000_0002.mha (axial CT imaging for nodule 0002)<br> ├── ...<br> </p> <p>Please cite the following article if you are using the <a href="https://luna16.grand-challenge.org/Data/">2016 Lung Nodule Analysis (LUNA) challenge</a> false positive reduction track data set:</p> <pre><code>Setio AAA, Traverso A, de Bel T, Berens MSN, Bogaard CVD, Cerello P, Chen H, Dou Q, Fantacci ME, Geurts B, Gugten RV, Heng PA, Jansen B, de Kaste MMJ, Kotov V, Lin JY, Manders JTMC, Sóñora-Mengana A, García-Naranjo JC, Papavasileiou E, Prokop M, Saletta M, Schaefer-Prokop CM, Scholten ET, Scholten L, Snoeren MM, Torres EL, Vandemeulebroucke J, Walasek N, Zuidhof GCA, Ginneken BV, Jacobs C. Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: The LUNA16 challenge. Med Image Anal. 2017 Dec;42:1-13. doi: 10.1016/j.media.2017.06.015. Epub 2017 Jul 13. PMID: 28732268.</code></pre> <p> </p>
False and true positives in arthropod thermal adaptation candidate gene lists
<p>Genome-wide studies are prone to false positives due to inherently low priors and statistical power. One approach to ameliorate this problem is to seek validation of reported candidate genes across independent studies: genes with repeatedly discovered effects are less likely to be false positives. Inversely, genes reported only as many times as expected by chance alone, while possibly representing novel discoveries, are also more likely to be false positives. We show that, across over 30 genome-wide studies that reported <i>Drosophila</i> and <i>Daphnia </i>genes with possible roles in thermal adaptation, the combined lists of candidate genes and orthologous groups are rapidly approaching the total number of genes and orthologous groups in the genome, respectively, consistent with the expectation of high frequency of false positives. The majority of these spurious candidates have been identified by one or a few studies, as expected by chance alone. In contrast, a noticeable minority of genes have been identified by numerous studies with the probabilities of such discoveries occurring by chance alone being exceedingly small. For this subset of genes, different studies are in agreement with each other despite differences in the ecological settings, genomic tools and methodology, and reporting thresholds. We provide a reference set of presumed true positives among <i>Drosophila</i> candidate genes and orthologous groups involved in response to changes in temperature, suitable for cross-validation purposes. Despite this approach being prone to false negatives, this list of presumed true positives includes several hundred genes, consistent with the "omnigenic" concept of genetic architecture of complex traits.</p>
False Positives Dataset
<p>This dataset contains false positives identified in our study on evaluating SQL queries through Online Judge Systems (OJSs). False positives refer to incorrect queries that were mistakenly judged as correct by the OJS. These false positives reflect the types of errors students may encounter when submitting SQL queries in OJSs, particularly logical errors. The dataset includes various classifications of these logical errors and also provides code for statistical analysis of these classifications. Notably, a single query may exhibit multiple error types, meaning it can fall into several error categories simultaneously.</p>
False and true positives in arthropod thermal adaptation candidate gene lists
Open the record for dataset details and reuse information.
Data from: Occupancy models for data with false positive and false negative errors and heterogeneity across sites and surveys
False positive detections, such as species misidentifications, occur in ecological data, although many models do not account for them. Consequently, these models are expected to generate biased inference. The main challenge in an analysis of data with false positives is to distinguish false positive and false negative processes while modeling realistic levels of heterogeneity in occupancy and detection probabilities without restrictive assumptions about parameter spaces. Building on previous attempts to account for false positive and false negative detections in occupancy models, we present hierarchical Bayesian models that utilize a subset of data with either confirmed detections of a species' presence (CP model) or both confirmed presences and confirmed absences (CACP model). We demonstrate that our models overcome the challenges associated with false positive data by evaluating model performance in Monte Carlo simulations of a variety of scenarios. Our models also have the ability to improve inference by incorporating previous knowledge through informative priors. We describe an example application of the CP model to quantify the relationship between songbird occupancy and residential development, plus we provide instructions for ecologists to use the CACP and CP models in their own research. Monte Carlo simulation results indicated that, when data contained false positive detections, the CACP and CP models generated more accurate and precise posterior probability distributions than a model that assumed data did not have false positive errors. For the scenarios we expect to be most generally applicable, those with heterogeneity in occupancy and detection, the CACP and CP models generated essentially unbiased posterior occupancy probabilities. The CACP model with vague priors generated unbiased posterior distributions for covariate coefficients. The CP model generated unbiased posterior distributions for covariate coefficients with vague or informative priors, depending on the function relating covariates to occupancy probabilities. We conclude that the CACP and CP models generate accurate inference in situations with false positive data for which previous models were not suitable.
Occupancy in dynamic systems: accounting for multiple scales and false positives using environmental DNA to inform monitoring
<div class="page"> <div class="section"> <p>Occupancy is an important metric to understand current and future trends in populations that have declined globally. In addition, occupancy can be an efficient tool for conducting landscape-scale and long-term monitoring. A challenge for occupancy monitoring programs is to determine the appropriate spatial scale of analysis and to obtain precise occupancy estimates for elusive species. We used a multi-scale occupancy model to assess occupancy of Columbia spotted frogs in the Great Basin, USA, based on environmental DNA (eDNA) detections. We collected three replicate eDNA samples at 220 sites across the Great Basin. We estimated and modeled ecological factors that described watershed and site occupancy at multiple spatial scales simultaneously while accounting for imperfect detection. Additionally, we conducted visual and dipnet surveys at all sites and used our paired detections to estimate the probability of a false positive detection for our eDNA sampling. We applied the estimated false positive rate to our multi-scale occupancy dataset and assessed changes in model selection. We had higher naïve occupancy estimates for eDNA (0.37) than for traditional survey methods (0.20). We estimated our false positive detection rate per qPCR replicate at 0.023 (95% CI: 0.016-0.033). When the false positive rate was applied to the multi-scale dataset, we did not observe substantial changes in model selection or parameter estimates. Conservation and resource managers have an increasing need to understand species occupancy in highly variable landscapes where the spatial distribution of habitat changes significantly over time due to climate change and human impact. A multi-scale occupancy approach can be used to obtain regional occupancy estimates that can account for spatially dynamic differences in availability over time, especially when assessing potential declines. Additionally, this study demonstrates how eDNA can be used as an effective tool for improved occupancy estimates across broad geographic scales for long-term monitoring.</p> </div> </div>
Raw Data for Studies 1 and 2 in False-Positive Psychology
<p>Excel file containing data for experiments 1 and 2 reported http://pss.sagepub.com/content/22/11/1359.short</p> <p> </p>
Kepler photometry used in KOI false positive probability calculations
<p>This is the Kepler photometry that can be used to reproduce the calculations of the Morton+ (2016) KOI FPP analysis. See instructions at the koi-fpp repository for use.</p>
True positive rate and false positive rate used to generate S2_Fig3A and S2_Table6
<p>In the attached file, first column corresponds to true positive (TP), second column corresponds to false positive (FP), third column corresponds to false negative (FN), fourth column corresponds to true negative (TN), fifth column corresponds to true positive rate (TPR) and sixth column corresponds to false positive rate (FPR), while each row corresponds to the session number. The data was used to generate patient's B, "Receiver operating characteristic (ROC) curve of the binary support vector machine (SVM) classifier" and "Contingency table" during training session, S2_Fig3A and S2_Table6, respectively.</p>
True positive rate and false positive rate used to generate S2_Fig3B and S2_Table7
<p>In the attached file, first column corresponds to true positive (TP), second column corresponds to false positive (FP), third column corresponds to false negative (FN), fourth column corresponds to true negative (TN), fifth column corresponds to true positive rate (TPR) and sixth column corresponds to false positive rate (FPR), while each row corresponds to the session number. The data was used to generate patient's B, "Receiver operating characteristic (ROC) curve of the binary support vector machine (SVM) classifier" and "Contingency table" during feedback session, S2_Fig3B and S2_Table7, respectively.</p>
True positive rate and false positive rate used to generate S2_Fig4A and S2_Table8
<p>In the attached file, first column corresponds to true positive (TP), second column corresponds to false positive (FP), third column corresponds to false negative (FN), fourth column corresponds to true negative (TN), fifth column corresponds to true positive rate (TPR) and sixth column corresponds to false positive rate (FPR), while each row corresponds to the session number. The data was used to generate patient's W, "Receiver operating characteristic (ROC) curve of the binary support vector machine (SVM) classifier" and "Contingency table" during training session, S2_Fig4A and S2_Table8, respectively.</p>
True positive rate and false positive rate used to generate S2_Fig4B and S2_Table9
<p>In the attached file, first column corresponds to true positive (TP), second column corresponds to false positive (FP), third column corresponds to false negative (FN), fourth column corresponds to true negative (TN), fifth column corresponds to true positive rate (TPR) and sixth column corresponds to false positive rate (FPR), while each row corresponds to the session number. The data was used to generate patient's W, "Receiver operating characteristic (ROC) curve of the binary support vector machine (SVM) classifier" and "Contingency table" during training session, S2_Fig4B and S2_Table9, respectively.</p>
True positive rate and false positive rate used to generate S2_Fig1A and S2_Table2
<p>In the attached file, first column corresponds to true positive (TP), second column corresponds to false positive (FP), third column corresponds to false negative (FN), fourth column corresponds to true negative (TN), fifth column corresponds to true positive rate (TPR) and sixth column corresponds to false positive rate (FPR), while each row corresponds to the session number. The data was used to generate patient's F, "Receiver operating characteristic (ROC) curve of the binary support vector machine (SVM) classifier" and "Contingency table" during training session, S2_Fig1A and S2_Table2, respectively.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.