Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,474
datasets available to search
ShareScore release 0.7.1
Dataset results
1,474 results for “Reliability”
Experimental dataset: C-SMART for reliable and efficient NN inference in a neutron-irradiated bare-metal system
<p>Will add more information soon.</p>
Supporting data for "Reliable interpretability of biology-inspired deep neural networks"
<p><strong>Contents</strong></p> <p><em>data.tgz</em> contains all data necessary for reproducing the analysis in the manuscript. After cloning the GitHub repository, extract the contents of this file into folder <em>data</em>. The archive contains the following subfolders:</p> <ul> <li><em>dtox</em><br> DTox results, one subfolder per seed <ul> <li><em>module_relevance.tsv</em>: contains node importance scores, with the following columns: <ul> <li>(first, unnamed): compound identifier</li> <li>remaining columns: node identifiers (UniProt and Reactome IDs)</li> </ul> </li> <li><em>test_labels.csv</em>: predictions for the test set, with two columns: <ul> <li>truth: true label (0 or 1)</li> <li>predicted: predicted label (decimal number between 0 and 1)<br> </li> </ul> </li> </ul> </li> <li><em>mskimpact_[cancer type]_[experiment]</em><br> P-NET results using the MSK-IMPACT 2017 dataset, one subfolder per seed<br> [cancer type] is one of bc (breast cancer), cc (colorectal cancer), nsclc (non-small cell lung cancer), or pc (prostate cancer)<br> [experiment] is one of original (original setup) and shuffled (shuffled labels)<br> </li> <li><em>pnet_[experiment]</em><br> P-NET results using the original (prostate cancer) dataset, one subfolder per seed<br> [experiment] is one of deterministic (deterministic input data), original (original setup), and shuffled (shuffled labels) <ul> <li><em>node_importance.csv</em>: contains node importance scores, with the following columns: <ul> <li>(first, unnamed): node name</li> <li>coef: original node importance scores</li> <li>coef_graph: indegree plus outdegree of node</li> <li>coef_combined: adjusted node importance score (= coef / coef_graph if coef_graph > mean(coef_graph) + 5 sd(coef_graph) in the respective layer)</li> <li>coef_combined_zscore: scaled coef_combined</li> <li>coef_combined2: z(z(coef_graph) - z(coef))</li> <li>layer: layer of the node</li> </ul> </li> <li><em>predictions_test.csv</em>: predictions for the test set, with the following columns: <ul> <li>(first, unnamed): sample name</li> <li>pred: predicted class (unfortunately, encoded by a double 1.0 or 0.0)</li> <li>pred_scores: probability of the predicted class</li> <li>y: true class (encoded as integer 1 or 0)</li> </ul> </li> <li><em>predictions_train.csv</em>: predictions for the training set (same columns as above)</li> <li><em>link_weights_[layer].csv</em>: only in subfolder 234_20080808; matrices with edge weights</li> </ul> </li> </ul> <p> </p> <p><strong>Changelog</strong></p> <p><em>v1.1.0 – 2023-06-28</em></p> <ul> <li>added DTox results</li> <li>added results of P-NET experiments with MSK-IMPACT 2017 dataset</li> </ul> <p><em>v1.0.0 – 2023-03-22</em></p> <ul> <li>initial release</li> </ul>
Supplementary codes and datasets for "Efficient numerical method for reliable upper and lower bounds on homogenized parameters"
<p>This repository supports L. Gaynutdinova, M. Ladecký, A. Nekvinda, I. Pultarová, and J. Zeman, <em>Efficient numerical method for reliable upper and lower bounds on homogenized parameters</em> (first announced as the arXiv preprint <a href="http://arxiv.org/abs/2208.09940">2208.09940</a>).</p> <p>In particular, it contains MATLAB source files for reproducing the results presented in Examples 1 and 2 of the manuscript for discretizations with <span class="math-tex">\(N_1 = N_2 = N_3 = 6, 12, 24\)</span>. For other parameters, the code needs to be modified manually.</p> <p>The most recent version of the codes is available in the <a href="https://gitlab.com/ul_bounds_homog_CTU/3d-fem-homogenized-parameters">GitLab repository</a>.</p>
Reliable imputation of spatial transcriptome with uncertainty estimation and spatial regularization
<p>Imputation of missing features in spatial transcriptomics is urgently demanded due to technology limitations, while most existing computational methods suffer from moderate accuracy and cannot estimate the reliability of the imputation. <br> To fill the research gaps, we introduce a computational model, TransImp, that imputes the missing feature modality in spatial transcriptomics by mapping it from single-cell reference. Uniquely, we derived a set of attributes that can accurately predict imputation uncertainty, hence enabling us to select reliably imputed genes. Also, we introduced a spatial auto-correlation metric as a regularization to avoid overestimating spatial patterns. Multiple datasets from various platforms have demonstrated that our approach significantly improves the reliability of downstream analyses in detecting spatial variable genes and interacting ligand-receptor pairs. Therefore, TransImp offers a way towards a reliable spatial analysis of missing features for both matched and unseen modalities, e.g., nascent RNAs.</p>
Replication Package for "A Catch-22--the Test-Retest Method of Reliability Estimation"
<p>Replication package for the paper "A Catch-22--the Test-Retest Method of Reliability Estimation".</p><p>Files included in the replication package and purpose of each file:</p><p>[1]Datafiles containing variables used in the analysis: question content, stability and reliability estimates (in SPSS and Stata format):<br>01_GSS_gammaV26_April2022_extract.sav<br>01_GSS_gammaV26_April2022_extract.dta<br>01_GSS_gammaV26_April2022_long_extract.dta</p><p>[2]SPSS Syntax file for replicating Tables 2,3 and Appendix Table: <br>02_SPSS_syntax_Catch22.sps</p><p>[3]Stata .do file containing the code for the regression models (Table 4):<br>03_Stata_code_Catch22.do</p><p>[4]List of GSS variables, wordings, and responses for variables included in the analysis (Excel file)<br>04_GSS variable wordings.xlsx </p>
16S rRNA phylogeny and clustering is not a reliable proxy for genome-based taxonomy in Streptomyces
<p>This file is intended as supplementary information for a forthcoming publication: 16S rRNA phylogeny and clustering is not a reliable proxy for genome-based taxonomy in <em>Streptomyces</em>. </p>
Codes and catalogs for: Parametric testing of EQTransformer's performance against a high-quality, manually-picked catalog for reliable and accurate seismic phase picking
<p><strong>Codes and Catalogs for:</strong> "Parametric Testing of EQTransformer's Performance Against a High-Quality, Manually-Picked Catalog for Reliable and Accurate Seismic Phase Picking."</p> <p><strong>Codes:</strong></p> <ol> <li><strong>overlap_check.py:</strong> This script tests the overlap parameter of EQTransformer to help minimize detection inconsistencies.</li> <li><strong>picks_comparison_other_networks.py:</strong> This code evaluates the probability threshold of EQTransformer by obtaining the time differences between picks from a catalog and EQTransformer.</li> <li><strong>test_seisbench_eq_eqt.py:</strong> A comparative analysis between the native EQTransformer and its implementation in SeisBench.</li> </ol> <p><strong>Catalogs:</strong></p> <ol> <li><strong>picks_differences_all_years_0.01_mag_cat.csv:</strong> This catalog presents pick differences for the central Alpine Fault using the SAMBA network and manual picks from Michailos et al. (2019).</li> <li><strong>sed_picks.csv:</strong> A catalog that showcases pick differences derived from data obtained from the Swiss Seismological Service (SED).</li> </ol> <p><strong>Note:</strong> Versions <1.0 represent pre-acceptence files and should not be used.</p>
Discovery of sparse, reliable omic biomarkers with Stabl
<p><span>Adoption of high-content omic technologies in clinical studies, coupled with computational </span><span>methods, have yielded an abundance of candidate biomarkers. However, translating such find</span><span>ings into bona fide clinical biomarkers remains challenging.</span> <span>To facilitate this process, we </span><span>introduce Stabl, a general machine learning framework that identifies a sparse, reliable set </span><span>of biomarkers by integrating noise injection and a data-driven signal-to-noise threshold into </span><span>multivariable predictive modeling.</span> <span>Evaluation of Stabl on synthetic datasets and five inde</span><span>pendent clinical studies demonstrates improved biomarker sparsity and reliability compared to </span><span>commonly used sparsity-promoting regularization methods while maintaining predictive per</span><span>formance; it distills datasets containing 1,400 to 35,000 features down to 4 to 34 candidate </span><span>biomarkers. Stabl extends to multi-omic integration tasks, enabling biological interpretation of </span><span>complex predictive models, as it hones in on a shortlist of proteomic, metabolomic, and cyto</span><span>metric events predicting labor onset, microbial biomarkers of preterm birth, and a pre-operative </span><span>immune signature of post-surgical infections.</span></p>
Data For: Herbarium specimens provide reliable estimates of phenological responsiveness to climate at unparalleled taxonomic and spatiotemporal scales
Open the record for dataset details and reuse information.
Discovery of sparse, reliable omic biomarkers with Stabl
Open the record for dataset details and reuse information.
Reliably predicting pollinator abundance: challenges of calibrating process-based ecological models
<p>1. Pollination is a key ecosystem service for global agriculture but evidence of pollinator population declines is growing. Reliable spatial modelling of pollinator abundance is essential if we are to identify areas at risk of pollination service deficit and effectively target resources to support pollinator populations. Many models exist which predict pollinator abundance but few have been calibrated against observational data from multiple habitats to ensure their predictions are accurate.</p> <p>2. We selected the most advanced process-based pollinator abundance model available and calibrated it for bumblebees and solitary bees using survey data collected at 239 sites across Great Britain. We compared three versions of the model: one parameterised using estimates based on expert opinion, one where the parameters are calibrated using a purely data-driven approach and one where we allow the expert opinion estimates to inform the calibration process.</p> <p>3. All three model versions showed significant agreement with the survey data, demonstrating this model's potential to reliably map pollinator abundance. However, there were significant differences between the nesting/floral attractiveness scores obtained by the two calibration methods and from the original expert opinion scores.</p> <p>4. Our results highlight a key universal challenge of calibrating spatially-explicit, process-based ecological models. Notably, the desire to reliably represent complex ecological processes in finely mapped landscapes necessarily generates a large number of parameters, which are challenging to calibrate with ecological and geographical data that is often noisy, biased, asynchronous and sometimes inaccurate. Purely data-driven calibration can therefore result in unrealistic parameter values, despite appearing to improve model-data agreement over initial expert opinion estimates. We therefore advocate a combined approach where data-driven calibration and expert opinion are integrated into an iterative Delphi-like process, which simultaneously combines model calibration and credibility assessment. This may provide the best opportunity to obtain realistic parameter estimates and reliable model predictions for ecological systems with expert knowledge gaps and patchy ecological data.</p>
Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates
DNA metabarcoding is promising for cost-effective biodiversity monitoring, but reliable diversity estimates are difficult to achieve and validate. Here we present and validate a method, called LULU, for removing erroneous molecular operational taxonomic units (OTUs) from community data derived by high-throughput sequencing of amplified marker genes. LULU identifies errors by combining sequence similarity and co-occurrence patterns. To validate the LULU method, we use a unique data set of high quality survey data of vascular plants paired with plant ITS2 metabarcoding data of DNA extracted from soil from 130 sites in Denmark spanning major environmental gradients. OTU tables are produced with several different OTU definition algorithms and subsequently curated with LULU, and validated against field survey data. LULU curation consistently improves α-diversity estimates and other biodiversity metrics, and does not require a sequence reference database; thus, it represents a promising method for reliable biodiversity estimation.
Data from: Females can solve the problem of low signal reliability by assessing multiple male traits
Male signals that provide information to females about mating benefits are often of low reliability. It is thus not clear why females often express strong signal preferences. We tested the hypothesis that females can distinguish between males with preferred signals that provide lower and higher quality direct benefits. In the field cricket, Gryllus lineaticeps, females usually prefer higher male chirp rates, but chirp rate is positively correlated with the fecundity benefits females will receive from males only for males that have experienced low quality diets. We paired females with muted males that were maintained on low or high nutrition diets, during the interactions we broadcast a replacement high chirp rate, and we observed whether females mated with the assigned male. Females were more likely to mate when paired with low nutrition males. These results suggest that females have evolved assessment mechanisms that allow them distinguish between males with preferred signals that provide high quality benefits (low nutrition males with high chirp rates) and males with preferred signals that provide low quality benefits (high nutrition males with high chirp rates).
Psychosocial functioning before and after surgical treatment for morbid obesity: Reliability and validity of the Norwegian version of obesity-related problems scale
<p>This is a dataset (SPSS, sav. file) related to the study "Psychosocial functioning before and after surgical treatment for morbid obesity: Reliability and validity of the Norwegian version of obesity-related problems scale". If anyone wants the file in a different format, contact the uploader (see thelink on the right)</p>
Intersession reliability of population receptive field estimates
<p>Proccessed pRF data comparing parameter estimates in visual regions across two days. Requires Matlab and SamSrf toolbox (https://figshare.com/articles/SamSrf_toolbox_for_pRF_mapping/1344765).</p>
Dataset for Article: "Citizen Science Provides a Reliable and Scalable Tool to Track Disease-Carrying Mosquitoes"
<p>This is the dataset used for the analysis in Citizen Science Provides a Reliable and Scalable Tool to Track Disease-Carrying Mosquitoes, by John R.B. Palmer, Aitana Oltra, Francisco Collantes, Juan Antonio Delgado, Javier Lucientes, Sarah Delacour, Mikel Bengoa, Roger Eritja, and Frederic Bartumeus. </p> <p> </p> <p>Copyright &copy; 2017 John R.B. Palmer, Aitana Oltra, Francisco Collantes, Juan Antonio Delgado, Javier Lucientes, Sarah Delacour, Mikel Bengoa, Roger Eritja, and Frederic Bartumeus.</p> <p><br> Dataset for Article: "Citizen Science Provides a Reliable and Scalable Tool to Track Disease-Carrying Mosquitoes" by John R.B. Palmer, Aitana Oltra, Francisco Collantes, Juan Antonio Delgado, Javier Lucientes, Sarah Delacour, Mikel Bengoa, Roger Eritja, and Frederic Bartumeus is licensed under a Creative Commons Attribution 4.0 International License.</p>
Opening the museum's vault: Historical field records preserve reliable ecological data
<p><span>Museum specimens have long served as foundational data sources for ecological, evolutionary, and environmental research. Continued reimagining of museum collections is now also generating new types of data associated with, but beyond physical specimens, a concept known as "extended specimens". Field notes penned by generations of naturalists contain first-hand ecological observations associated with museum collections and comprise a form of extended specimens with the potential to provide novel ecological data spanning broad geographic and temporal scales. Despite their data-yielding potential, however, field notes remain underutilized in research due to their heterogeneous, unstandardized, and qualitative nature. We introduce an approach for transforming descriptive ecological notes into quantitative data suitable for statistical analysis. Tests with simulated and real-world published data show that field notes and our transformation approach retain reliable quantitative ecological information under a range of sample sizes and evolutionary scenarios. Unlocking the wealth of data contained within field records could facilitate investigations into the ecology of clades whose diversity, distribution, or other demographic features present challenges to traditional ecological studies, improve our understanding of long-term environmental and evolutionary change, and enhance predictions of future change.</span></p>
Shared single copy genes are generally reliable for inferring phylogenetic relationships among polyploid taxa
<p>Polyploidy, or whole-genome duplication, is expected to confound the inference of species trees with phylogenetic methods for two reasons. First, the presence of retained duplicated genes requires the reconciliation of the inferred gene trees to a proposed species tree. Second, even if the analyses are restricted to shared single copy genes, the occurrence of reciprocal gene loss, where the surviving genes in different species are paralogs from the polyploidy rather than orthologs, will mean that such genes will not have evolved under the corresponding species tree and may not have gene trees that allow inference of the species tree. Here we analyze three different ancient polyploidy events, using synteny-based inferences of orthology and paralogy to infer gene trees from more than 17,000 sets of homologous genes. We find that the simple use of single copy genes from polyploid organisms provides reasonably robust phylogenetic signals, despite the presence of reciprocal gene losses. Such gene trees are also most often in accord with the inferred species relationships inferred from maximum likelihood models of gene loss after polyploidy: a completely distinct phylogenetic signal present in these genomes. As seen in other studies, however, we find that methods for inferring phylogenetic confidence yield high support values even in cases where the underlying data suggest meaningful conflict in the phylogenetic signals.</p>
Establishing Fully-Automated Fundus-Controlled Dark Adaptometry: A Validation and Retest-Reliability Study
<p>This is the data presented in our manuscript <i>'Establishing Fully-Automated Fundus-Controlled Dark Adaptometry: A Validation and Retest-Reliability Study'</i> published in Translational Vision Science & Technology.</p>
Figure 3 in Determination of an efficient and reliable method for PCR detection of borrelial DNA from engorged ticks
Figure 3. Ratio of four different PCR types in presented ticks categories and isolation groups.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.