Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
226
datasets available to search
ShareScore release 0.9.0
Dataset results
226 results for “confidence”
Confidence in Detection and Discrimination
Open the record for dataset details and reuse information.
Disentangling the origins of confidence in speeded perceptual judgments through multimodal imaging
Open the record for dataset details and reuse information.
Water Dawgs STEM Confidence Survey Results, 2023
This dataset originates from a STEM confidence survey conducted during the Water Dawgs program—a paid summer initiative hosted at the University of Georgia (UGA) in Athens, Georgia, USA. Held over 10 days in Summer 2023, the program engaged 16 high school students in a hands-on experience in freshwater science. Water Dawgs was designed to support students’ academic and professional development, with an emphasis on increasing access for individuals from populations historically excluded from STEM fields. The initiative was part of the broader impacts of two National Science Foundation-funded research projects and was shaped by three main objectives: (1) to expand access to university-led STEM opportunities by collaborating with a local public high school to recruit students from underrepresented backgrounds; (2) to highlight the connections between environmental science and students’ everyday lives and future career options, including non-STEM pathways; and (3) to foster greater self-efficacy in engaging with STEM subjects, particularly environmental science. The survey was administered at both the beginning and conclusion of the 10-day program to assess changes in participants’ confidence related to STEM. The survey measured STEM confidence using a Likert scale ranging from 1 to 5, with 1 indicating "not at all confident" and 5 indicating "totally confident." The included R code contains a script to generate a figure and summary statistics for questions related to general STEM confidence, specifically Questions 2, 6, 7, 9, and 12. One participant who only submitted a post-program survey was excluded from the dataset and analysis to ensure consistency across responses.
Simultaneous EEG-fMRI - Confidence in perceptual decisions
Open the record for dataset details and reuse information.
Improving the accuracy of automated labeling of specimen images datasets via a confidence-based process - Datasets
<p>This dataset contains supporting data for a research project aimed at analysing herbarium samples from the New England area at a large scale with deep learning techniques. Details on the methodology are shared in the acompanying paper (to be published).</p> <p>Content:</p> <ul> <li>dataset600k_withAI.csv : A dataset of over 600.000 herbarium samples with its record metadata and a corresponding AI phenological annotations with matching confidence scores. The entirety of the record headers are provided, extracted directly from the NEVP portal. In addition, the AI labels are defined by the following headers. These 8 columns represent 4 binary classifiers with the Presence/Absence of each 4 traits and corresponding confidence (as a percentage - presence/absence percentages sum to 1).<br> <ul> <li> <table> <tbody> <tr> <td>Flowering</td> <td>Not Flowering</td> <td>Budding</td> <td>Not Budding</td> <td>Fruiting</td> <td>Not Fruiting</td> <td>Reproductive</td> <td>Not Reproductive</td> </tr> </tbody> </table> </li> </ul> </li> </ul> <ul> <li>data_species_with_statuses.csv: A processed dataset summarizing flowering period shift at a species level. Two types of headers are provided. <ul> <li>First metadata concerning the flowering shift and the data used to compute that value: <ul> <li> <table> <tbody> <tr> <td>genus</td> <td>genus_species</td> <td>slope</td> <td>nb_specimens</td> <td>p_value_significance</td> <td>trend_category</td> </tr> <tr> <td>Genus of the species</td> <td>Binomial name of the species</td> <td>Regression slope defining the flowering shift as a slope</td> <td>Number of herbarium specimens used to compute the shift</td> <td>P-value significance of the slope being non-zero. ('Non Significant'/'Significant')</td> <td>Summary of the shift as a binary characteristic ('Earlier'/'Later')</td> </tr> </tbody> </table> </li> </ul> </li> <li>Second, metadata summarizing various traits associated to each species: <ul> <li> <table> <tbody> <tr> <td>lifeform_status</td> <td>native_introduced_status</td> <td>wetland_status</td> <td>seasonality_average</td> <td>seasonality_spread</td> </tr> <tr> <td>Growth form from the USDA PLANTS Database. 'Forb_Herb', 'Shrub_Tree' or 'Vine'</td> <td>'Native'/'Introduced' status from the USDA PLANTS Database.</td> <td> <p>National Wetland Plant List (NWPL) Wetland Indicator Status within the Northcentral and Northeast Region</p> <p>'OBL'/'FACW'/'FAC'/'FACU'/'UPL'</p> </td> <td>A characteristic of the flowering season of the species based on the mean Day of Year of the analysed specimens: if <=180: 'Early', else 'Late'</td> <td>A characteristic of the flowering season of the species based on the spread of the flowering season. Less than 28 days: 'Narrow', larger: 'Large'.</td> </tr> </tbody> </table> <p> </p> </li> </ul> </li> </ul> </li> <li>phylogenetic_tree.tre: The raw data used to generate the visualization of the flowering seasonality character and the detected flowering shift foreach species on a phylogenetic tree.</li> <li>phylogenetic_processed_dataset.csv: The processed dataset resuting from the phylogenetic signal analysis. For each trait, an associated significance binary value is provided.</li> </ul>
ESM Atlas v0 random sample of high confidence predicted protein structures
<p>A random sample out of the 225M high confidence predictions in the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> High confidence is defined as mean pLDDT > 0.7 and pTM > 0.7 and corresponds to ∼36% of the total 617M proteins folded.<br> This is the random sample used for analysis in the paper as well as visualization on the <a href="http://esmatlas.com/">esmatlas.com</a> Explore page.<br> Sample size: 999,520 based on 999,996 unique randomly sampled IDs and 0.05% missing data in the processing pipeline.</p>
Bioactive compounds with no structural analogs (high-confidence activity data)
<p>A set of 52,815 unique bioactive compounds (human targets, high-confidence activity data) with no structural analogs with high-confidence activity data was extracted from ChEMBL. For each compound the ChEMBL compound ID (CHEMBLID_Compound) and high-confidence target annotation(s) (CHEMBLID_Targets) are provided. The data set was generated as a part of an analysis to be published in 'Medicinal Chemistry Communications'. </p>
Data for High-confidence 3D template matching for cryo-electron tomography
<p>This repository contains supporting data to the manuscript: "High-confidence 3D template matching for cryo-electron tomography" by Sergio Cruz-León, et al. </p> <p>It contains an example to run high-confidence template matching with GAPSTOP-TM, supporting raw data to the manuscript and a jupyter notebook for data visualization. </p> <p> </p> <p>Contact information:<br>Name: Sergio Cruz-León, PhD<br>Institution: Department of Theoretical Biophysics, Max Planck Institute of Biophysics<br>Address: Max-von-Laue-Str. 3, 60438 Frankfurt am Main, Germany<br>Email: sergio.cruz@biophys.mpg.de</p>
Model Zoo for Robust Models are less Over-Confident
<p><strong>Model Zoo (PyTorch) of non-adversarially trained models for Robust Models are less Over-Confident (NeurIPS'22)</strong></p> <p>Abstract: <em>"Regardless of the success of convolutional neural networks (CNNs) in many academic benchmarks of computer vision tasks, their application in real-world is still facing fundamental challenges, like the inherent lack of robustness as unveiled by adversarial attacks. These attacks target to manipulate the network's prediction by adding a small amount of noise onto the input. In turn, adversarial training (AT) aims to achieve robustness against such attacks by including adversarial samples in the trainingset. However, a general analysis of the reliability and model calibration of these robust models beyond adversarial robustness is still pending. In this paper, we analyze a variety of adversarially trained models that achieve high robust accuracies when facing state-of-the-art attacks and we show that AT has an interesting side-effect: it leads to models that are significantly less overconfident with their decisions even on clean data than non-robust models. Further, our analysis of robust models shows that not only AT but also the model's building blocks (activation functions and pooling) have a strong influence on the models' confidence."</em></p>
Figure S1. Mediation analysis on the effect of insulin resistance on intraocular pressure. Figure S2. Forest plot showing the OR (95% CI) for EIOP of ALD versus NAFLD and the OR (95% CI) for EIOP of drinkers versus non-drinkers. Abbreviations: OR, odds ratio; CI, confidence interval; ALD, alcoholic liver disease; NAFLD, non-alcoholic fatty liver disease.
<p>Figure S1. Mediation analysis on the effect of insulin resistance on intraocular pressure.</p> <p>Figure S2. Forest plot showing the OR (95% CI) for EIOP of ALD versus NAFLD and the OR (95% CI) for EIOP of drinkers versus non-drinkers. Abbreviations: OR, odds ratio; CI, confidence interval; ALD, alcoholic liver disease; NAFLD, non-alcoholic fatty liver disease.</p>
Рис. 5. Коррелограммы покаЗателей обилиЯ наЗемных моллюсков раЗных воЗрастных групп (1 – ювенильные; 2 – вЗрослые; 3 – все вместе): A – H. lucorum, участок № 1, 2010 г.; B – Ch. tridens, участок № 2, 2011 г.; C – Ch. tridens, участок № 4, 2012 г.); D – Ch. tridens, участок № 5, 2012 г. (достоверные оценки индекса Морана отмечены Залитыми Значками). Fig. 5. Spatial correlogram of the land snail different age groups abundance (1 – juvenile; 2 – adult; 3 – total): A – H. lucorum, site 1, 2010; B – Ch. tridens, site 2, 2011; C – Ch. tridens, site 4, 2012; D – Ch. tridens, site 5, 2012 (Moran index confidence value presented by filled signs). in Analysis of the spatial distribution patterns of the land snail populations: a geostatistic method approach
Рис. 5. Коррелограммы покаЗателей обилиЯ наЗемных моллюсков раЗных воЗрастных групп (1 – ювенильные; 2 – вЗрослые; 3 – все вместе): A – H. lucorum, участок № 1, 2010 г.; B – Ch. tridens, участок № 2, 2011 г.; C – Ch. tridens, участок № 4, 2012 г.); D – Ch. tridens, участок № 5, 2012 г. (достоверные оценки индекса Морана отмечены Залитыми Значками). Fig. 5. Spatial correlogram of the land snail different age groups abundance (1 – juvenile; 2 – adult; 3 – total): A – H. lucorum, site 1, 2010; B – Ch. tridens, site 2, 2011; C – Ch. tridens, site 4, 2012; D – Ch. tridens, site 5, 2012 (Moran index confidence value presented by filled signs).
Рис. 4. Коррелограммы покаЗателей обилиЯ наЗемного моллюска M. cartusiana раЗных воЗрастных групп (1 – ювенильные; 2 – вЗрослые; 3 – все вместе): A – участок № 1, 2010 г.; B – участок № 2, 2011 г.; C – участок № 4, 2012 г.); D – участок № 5, 2012 г. (достоверные оценки индекса Морана отмечены Залитыми Значками). Fig. 4. Spatial correlogram of land snail M. cartusiana age groups abundance (1 – juvenile; 2 – adult; 3 – total): A – site 1, 2010; B – site 2, 2011; C – site 4, 2012; D – site 5, 2012 (Moran index confidence value presented by filled sings). in Analysis of the spatial distribution patterns of the land snail populations: a geostatistic method approach
Рис. 4. Коррелограммы покаЗателей обилиЯ наЗемного моллюска M. cartusiana раЗных воЗрастных групп (1 – ювенильные; 2 – вЗрослые; 3 – все вместе): A – участок № 1, 2010 г.; B – участок № 2, 2011 г.; C – участок № 4, 2012 г.); D – участок № 5, 2012 г. (достоверные оценки индекса Морана отмечены Залитыми Значками). Fig. 4. Spatial correlogram of land snail M. cartusiana age groups abundance (1 – juvenile; 2 – adult; 3 – total): A – site 1, 2010; B – site 2, 2011; C – site 4, 2012; D – site 5, 2012 (Moran index confidence value presented by filled sings).
Рис. 3. Коррелограммы покаЗателей обилиЯ наЗемного моллюска B. cylindrica раЗных воЗрастных групп (1 – ювенильные; 2 – вЗрослые; 3 – все вместе): A – участок № 1, 2010 г.; B – участок № 2, 2011 г.; C – участок № 4, 2012 г.); D – участок №5, 2012 г. (достоверные оценки индекса Морана отмечены Залитыми Значками). Fig. 3. Spatial correlogram of the land snail B. cylindrica age groups abundance (1 – juvenile; 2 – adult; 3 – total): A – site 1, 2010; B – site 2, 2011; C – site 4, 2012; D – site 5, 2012 (Moran index confidence value presented by filled signs). in Analysis of the spatial distribution patterns of the land snail populations: a geostatistic method approach
Рис. 3. Коррелограммы покаЗателей обилиЯ наЗемного моллюска B. cylindrica раЗных воЗрастных групп (1 – ювенильные; 2 – вЗрослые; 3 – все вместе): A – участок № 1, 2010 г.; B – участок № 2, 2011 г.; C – участок № 4, 2012 г.); D – участок №5, 2012 г. (достоверные оценки индекса Морана отмечены Залитыми Значками). Fig. 3. Spatial correlogram of the land snail B. cylindrica age groups abundance (1 – juvenile; 2 – adult; 3 – total): A – site 1, 2010; B – site 2, 2011; C – site 4, 2012; D – site 5, 2012 (Moran index confidence value presented by filled signs).
Рис. 2. ΔоΛговременная Αинамика весенней чисΛенности трех виΑов уток (A — трескунка; B — касатки; C — шиΛохвости) на ΑебеΑинском стационаре Хинганского заповеΑника (показаны уровень значимости и 95-процентный ΑоверитеΛьный интерваΛ) Fig. 2. Long-term spring number dynamics of three duck species at the Lebedinsky Station of Khingansky State Nature Reserve with p-values and 0.95 confidence intervals. A — Gargany; B — Falcated Duck; C — Pintail in The results of long-term observation of waterfowl spring migration in Khingan Nature Reserve, Eastern Russia
Рис. 2. ΔоΛговременная Αинамика весенней чисΛенности трех виΑов уток (A — трескунка; B — касатки; C — шиΛохвости) на ΑебеΑинском стационаре Хинганского заповеΑника (показаны уровень значимости и 95-процентный ΑоверитеΛьный интерваΛ) Fig. 2. Long-term spring number dynamics of three duck species at the Lebedinsky Station of Khingansky State Nature Reserve with p-values and 0.95 confidence intervals. A — Gargany; B — Falcated Duck; C — Pintail
Building Public Confidence in Constructed Wetlands for Wastewater Treatment and Reuse: Survey Data and Focus Group Transcripts
<p>Constructed wetlands have been proposed as a cost-effective wastewater treatment, storage, and reuse solution for communities that are considering alternative water supply options to meet essential demands. In 2016, we began exploring the idea of wastewater reuse and the construction of an experimental wetland in Sewanee, located in the southern U.S. state of Tennessee. As a major barrier to water reuse is often public resistance, we conducted a survey and focus groups to determine strategies to develop and initiate a community engagement campaign, aiming to empower residents to form reasoned opinions about local water supply options.</p> <p>This data set includes the survey that was distributed to Sewanee community members between November 2015 and February 2016, as well as protocols for three focus groups that were conducted with K12 teachers and community leaders on February 11 and 12, 2016. The survey results are summarized in a Microsoft Excel file. The three focus groups were transcribed, these transcripts are included here as PDF documents.</p>
How Confidence in Prior Attitudes, Social Tag Popularity, and Source Credibility Shape Confirmation Bias Toward Antidepressants and Psychotherapy in a Representative German Sample: Randomized Controlled Web-Based Study
<p>ABSTRACT</p> <p>Background: In health-related, Web-based information search, people should select information in line with expert (vs nonexpert) information, independent of their prior attitudes and consequent confirmation bias.</p> <p>Objective: This study aimed to investigate confirmation bias in mental health–related information search, particularly (1) if high confidence worsens confirmation bias, (2) if social tags eliminate the influence of prior attitudes, and (3) if people successfully distinguish high and low source credibility.</p> <p>Methods: In total, 520 participants of a representative sample of the German Web-based population were recruited via a panel company. Among them, 48.1% (250/520) participants completed the fully automated study. Participants provided <em>prior attitudes</em> about antidepressants and psychotherapy. We manipulated (1) <em>confidence</em> in prior attitudes when participants searched for blog posts about the treatment of depression, (2) <em>tag popularity</em> —either psychotherapy or antidepressant tags were more popular, and (3) <em>source credibility</em> with banners indicating high or low expertise of the tagging community. We measured <em>tag</em> and <em>blog post</em> selection, and <em>treatment</em><em>efficacy ratings</em> after navigation.</p> <p>Results: Tag popularity predicted the proportion of selected antidepressant tags (beta=.44, SE 0.11; <em>P</em><.001) and blog posts (beta=.46, SE 0.11; <em>P</em><.001). When confidence was low (−1 SD), participants selected more blog posts consistent with prior attitudes (beta=−.26, SE 0.05; <em>P</em><.001). Moreover, when confidence was low (−1 SD) and source credibility was high (+1 SD), the efficacy ratings of attitude-consistent treatments increased (beta=.34, SE 0.13; <em>P</em>=.01).</p> <p>Conclusions: We found correlational support for defense motivation account underlying confirmation bias in the mental health–related search context. That is, participants tended to select information that supported their prior attitudes, which is not in line with the current scientific evidence. Implications for presenting persuasive Web-based information are also discussed.</p> <p>Trial Registration: ClinicalTrials.gov NCT03899168; https://clinicaltrials.gov/ct2/show/NCT03899168 (Archived by WebCite at http://www.webcitation.org/77Nyot3Do)</p> <p>J Med Internet Res 2019;21(4):e11081</p> <p>doi:10.2196/11081</p>
Figure. Mean pre-adult development time (in days) values for all strains. Vertical bars denote 0.95 confidence intervals. in Effects of artificial migration of susceptible individuals on resistance and fitness of a fenitrothion-resistant strain of Musca domestica (L.) Diptera
Figure. Mean pre-adult development time (in days) values for all strains. Vertical bars denote 0.95 confidence intervals.
Text-fig. 6b. Z-score profile for Moča skull. Comparison of Moča skull indices with LUP and recent sample, –1.96 and +1.96: 95% tolerance interval (95% of the LUP and recent variability), 0: LUP and recent sample mean, grey band: 95% confidence interval of mean z-scores. Only those samples having more than n = 5 in the particular group are included. in A Late Upper Palaeolithic Skull From Moča (The Slovak Republic) In The Context Of Central Europe
Text-fig. 6b. Z-score profile for Moča skull. Comparison of Moča skull indices with LUP and recent sample, –1.96 and +1.96: 95% tolerance interval (95% of the LUP and recent variability), 0: LUP and recent sample mean, grey band: 95% confidence interval of mean z-scores. Only those samples having more than n = 5 in the particular group are included.
Text-fig. 6a. Z-score profile for Moča skull. Comparison of Moča skull measurements with LUP and recent sample, –1.96 and +1.96: 95% tolerance interval (95% of the LUP and recent variability), 0: LUP and recent sample mean, gray band: 95% confidence interval of mean z-scores. Only those samples having more than n = 5 in the particular group are included. in A Late Upper Palaeolithic Skull From Moča (The Slovak Republic) In The Context Of Central Europe
Text-fig. 6a. Z-score profile for Moča skull. Comparison of Moča skull measurements with LUP and recent sample, –1.96 and +1.96: 95% tolerance interval (95% of the LUP and recent variability), 0: LUP and recent sample mean, gray band: 95% confidence interval of mean z-scores. Only those samples having more than n = 5 in the particular group are included.
Fig. 1 in Assessing confidence intervals for stratigraphic ranges of higher taxa: The case of Lissamphibia
Fig. 1. Current hypotheses on the origin of the extant amphibians. Extant taxa in bold, paraphyletic taxa in quotation marks. A, B. Monophyletic origin from within the temnospondyls, with lepospondyls at the basalmost part of the amphibian stem (Panchen and Smithson 1988; Trueb and Cloutier 1991; Lombard and Sumida 1992; Ahlberg and Milner 1994). C, D. Monophyletic origin from within the temnospondyls, with the lepospondyls as reptiliomorphs (Ruta and Coates 2007; see also Ruta et al. 2003). E. Monophyletic origin from within the lepospondyls (Vallin and Laurin 2004; see also Laurin and Reisz 1997, 1999). F. Diphyletic origin with the anurans as temnospondyls, caecilians as lepospondyls, and urodeles as either temnospondyls or lepospondyls (Carroll and Currie 1975; Carroll and Holmes 1980; Carroll et al. 2004).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.