Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,129
datasets available to search
ShareScore release 0.9.0
Dataset results
2,129 results for “scores”
Data for: Advances and critical assessment of machine learning techniques for prediction of docking scores
Open the record for dataset details and reuse information.
Gini index mean score grain size estimates for Murray formation rocks in the Vera Rubin ridge (Gale crater, Mars) from ChemCam LIBS data (sols 1808-2298)
<p>Data in this repository are included as main research results in the Journal of Geophysical Research: Planets manuscripts entitled:</p> <p><em>"</em>Extensive diagenesis revealed by fine-scale features at Vera Rubin ridge, Gale crater, Mars<em>" </em></p> <p>"A lacustrine paleoenvironment recorded at Vera Rubin ridge, Gale crater: Overview of the sedimentology and stratigraphy observed by the Mars Science Laboratory Curiosity rover".</p> <p> </p> <p><strong>Table captions:</strong></p> <p><em>Table 1. </em>Murray formation targets from the Vera Rubin ridge used in the Gini mean index score (GIMS) analysis (sols 1808-2298) with summary information and G<sub>MEAN </sub>values with associated standard deviation errors. The grain size regimes (GSRs) were defined during the calibration procedure in Rivera-Hernández et al. (2020) and are defined as: mud (G<sub>MEAN</sub>=0.00-0.07; GSR1) and coarse silt to very fine sand (G<sub>MEAN</sub>=0.07-0.10; GSR2). Rocks with G<sub>MEAN</sub>=0.07 are exactly at GSR1 and GSR2 boundary and are reported as GSR1/GSR2. Next to the target names, the symbol * denotes that the ChemCam target was imaged by the Mars Hand Lens Imager (MAHLI), the symbol ** denotes that the dust removal tool was used before the MAHLI image was taken, and ~ signifies that a location close to the ChemCam target was imaged by the MAHLI. </p> <p><em>Table 2</em>. The mean, median, minimum and maximum G<sub>MEAN</sub> and the minimum and maximum grain size regime (GSR) for each Murray formation locality in the Vera Rubin ridge.</p> <p> </p> <p>References:</p> <p>Rivera-Hernández, F., Sumner, D. Y., Mangold N., Wiens, R.W., Edgett, K., Fedo, C., Schieber, J., Banham, S.G., Newsom, H., Gupta, S., Heydari, E., Stack, K.M., Nachon, M., Stein, N., & Maurice, S. (2020) Grain Size Variations in the Murray Formation: Stratigraphic Evidence for Changing Depositional Environments in Gale Crater, Mars. <em>Journal of Geophysical Research:</em><em> Planets.</em> doi:10.1029/2019JE006230</p>
BFI Inventory Scores
<p>Big Five Inventory (BFI) data collected from fresher Marketing students in the UK over five consecutive years. Data is as used for calculating group/cohort personality scores - i.e. with all scores for relevant items reversed in accordance with John, O. P., Donahue, E. M., & Kentle, R. L. (1991). The big five inventory: versions 4a and 54. University of California, Berkeley, Institute of Personality and Social Research.</p>
Eigen scores for human genome assembly GRCh38 Part 1 (Chr12 - Chr22)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Eigen scores for human genome assembly GRCh38 Part 2 (Chr6 - Chr11)
<p>Eigen is a spectral approach to the functional annotation of genetic variants in coding and noncoding regions. Eigen makes use of a variety of functional annotations in both coding and noncoding regions (such as protein function scores, evolutionary conservation scores, and epigenetic annotations from ENCODE and Roadmap Epigenomics projects), and combines them into one single measure of functional importance. Eigen is an unsupervised approach, and, unlike many existing methods, is not based on any labelled training data. Eigen produces estimates of predictive accuracy for each functional annotation score, and subsequently uses these estimates of accuracy to derive the aggregate functional score for variants of interest as a weighted linear combination of individual annotations.</p>
Figure 3 in An approach to scoring cursorial limb proportions in carnivorous dinosaurs and an attempt to account for allometry
Figure 3. Theropod phylogeny, with CLP scores reported for individual species and average CLP scores reported for larger clades.
Figure 2 in An approach to scoring cursorial limb proportions in carnivorous dinosaurs and an attempt to account for allometry
Figure 2. Log/log plot of femur vs. lower-leg length for the initial dataset of 53 theropod taxa. The red line denotes the best-fit power curve and the dotted lines denote the confidence interval.
Figure 1 in An approach to scoring cursorial limb proportions in carnivorous dinosaurs and an attempt to account for allometry
Figure 1. The general observation that smaller-bodied non-avian theropods tend to have proportionately longer lower legs holds true across comparisons between distantly related taxa (A), closely related taxa (B), and ontogenetic stages within a single taxon (C). All illustrations scaled to the same proximodistal femur length.
The reliability and validity of an Arabic version of the modified Japanese Orthopedic Association score for cervical myelopathy
<p><b>Purpose: </b>The modified Japanese Orthopedic Association score (mJOA) is considered as one of the most comprehensive scores in assessment of patients with cervical myelopathy. Hence, comes the need for providing a validated, translated and cross-culturally adapted versions in different languages in order to standardize patient's evaluation. This study aims to translate and validate an Arabic version of the mJOA.</p> <p><b>Methods: </b>65 patients of variable age and etiologies for compressive cervical myelopathy were recruited. Forward and backward translations were performed, followed by measurement of intra-observer and inter-observer reliability using intra-class correlation coefficient and Cronbach's alpha coefficient.</p> <p><b>Results: </b>The patients had a mean age of 58.08, predominantly males (69.2%). The intra-observer reliability and inter-observer reliability showed almost perfect agreement for the different sections and the total score, with 96.8 and 97.4, respectively.</p> <p><b>Conclusion: </b>This study provides a reliable, cross-culturally adapted Arabic version of mJOA for cervical myelopathy patients. Although the study was conducted on Egyptian patients, we believe it could be implemented with majority of Arabic speaking population.</p>
Data from: Prevalence of afebrile malaria and development of risk-scores for gradation of villages: a study from a hot-spot in Odisha
Introduction: Malaria is a public health emergency in India and Odisha. The national malaria elimination programme aims to expedite early identification, treatment and follow-up of malaria cases in hot-spots through a robust health system, besides focusing on efficient vector control. This study, a result of mass screening conducted in a hot-spot in Odisha, aimed to assess prevalence, identify and estimate the risks and develop a management tool for malaria elimination. Methods: Through a cross-sectional study and using WHO recommended Rapid Diagnostic Test (RDT), 13221 individuals were screened. Information about age, gender, education and health practices were collected along with blood sample (5 µl) for malaria testing. Altitude, forestation, availability of a village health worker and distance from secondary health center were captured using panel technique. A multi-level poisson regression model was used to analyze association between risk factors and prevalence of malaria, and to estimate risk scores. Results: The prevalence of malaria was 5.8% and afebrile malaria accounted for 79 percent of all confirmed cases. Higher proportion of Pv infections were afebrile (81%). We found the prevalence to be 1.38 (1.1664 - 1.6457) times higher in villages where the Accredited Social Health Activist (ASHA) didn't stay; the risk increased by 1.38 (1.0428 - 1.8272) and 1.92 (1.4428 - 2.5764) times in mid- and high-altitude tertiles. With regard to forest coverage, villages falling under mid- and highest-tertiles were 2.01 times (1.6194 - 2.5129) and 2.03 times (1.5477 - 2.6809), respectively, more likely affected by malaria. Similarly, villages of mid tertile and lowest tertile of education had 1.73 times (1.3392 - 2.2586) and 2.50 times (2.009 - 3.1244) higher prevalence of malaria. Conclusion: Presence of ASHA worker in villages, altitude, forestation, and education emerged as principal predictors of malaria infection in the study area. An easy-to-use risk-scoring system for ranking villages based on these risk factors could facilitate resource prioritization for malaria elimination.
GREEN-VARAN scores resources (GRCh37)
<p>Processed non-coding prediction scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh37 version for the following scores:</p> <ul> <li>ReMM v0.3.1 (<a href="https://charite.github.io/software-remm-score.html">https://charite.github.io/software-remm-score.html</a>)</li> <li>NCBoost v.1 (<a href="https://github.com/RausellLab/NCBoost">https://github.com/RausellLab/NCBoost</a>)</li> <li>ExPECTO (<a href="https://hb.flatironinstitute.org/expecto/">https://hb.flatironinstitute.org/expecto/</a>)</li> <li>LinSight (<a href="https://github.com/CshlSiepelLab/LINSIGHT">https://github.com/CshlSiepelLab/LINSIGHT</a>)</li> <li>GWAVA v1.0 (<a href="https://www.sanger.ac.uk/sanger/StatGen_Gwava">https://www.sanger.ac.uk/sanger/StatGen_Gwava</a>)</li> </ul> <p>If you use any of these score annotations with GREEN-VARAN please cite also the corresponding paper.</p>
GREEN-VARAN scores resources (DANN GRCh38)
<p>Processed DANN scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for DANN.</p> <p>See: <a href="https://academic.oup.com/bioinformatics/article/31/5/761/2748191">https://academic.oup.com/bioinformatics/article/31/5/761/2748191</a></p> <p>If you use DANN score annotations with GREEN-VARAN don't forget to cite also the original DANN paper.</p>
GREEN-VARAN scores resources (FIRE GRCh38)
<p>Processed FIRE scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for FIRE.</p> <p>See: <a href="https://sites.google.com/site/fireregulatoryvariation/">https://sites.google.com/site/fireregulatoryvariation/</a></p> <p>If you use FIRE score annotations with GREEN-VARAN don't forget to cite also the original FIRE paper.</p>
GREEN-VARAN scores resources (GRCh38)
<p>Processed non-coding prediction scores to be used with GREEN-VARAN</p> <p>This dataset contains the GRCh38 version for the following scores. When not available from the original source, the GRCh38 coordinates were obtained by liftover.</p> <ul> <li>ReMM v0.3.1 (<a href="https://charite.github.io/software-remm-score.html">https://charite.github.io/software-remm-score.html</a>)</li> <li>NCBoost v.1 (<a href="https://github.com/RausellLab/NCBoost">https://github.com/RausellLab/NCBoost</a>)</li> <li>ExPECTO (<a href="https://hb.flatironinstitute.org/expecto/">https://hb.flatironinstitute.org/expecto/</a>)</li> <li>LinSight (<a href="https://github.com/CshlSiepelLab/LINSIGHT">https://github.com/CshlSiepelLab/LINSIGHT</a>)</li> <li>GWAVA v1.0 (<a href="https://www.sanger.ac.uk/sanger/StatGen_Gwava">https://www.sanger.ac.uk/sanger/StatGen_Gwava</a>)</li> </ul> <p>If you use any of these score annotations with GREEN-VARAN please cite also the corresponding paper.</p>
Biometric Scores 2014 (BIOSCOTE 2014)
<p><strong>Description</strong><br> <br> This dataset contains raw scores in plain text format of several biometric (face and speaker) recognition systems applied on several datasets such as BANCA, Arface, FRGC, GBU, LFW, Multi-PIE, MOBIO, CAS-PEAL, NIST SRE 2012.</p> <p>The biometric recognition systems are described in the aforementioned manuscript and encompasses Gaussian mixture models, inter-session variability modelling, joint factor analysis and probabilistic linear discriminant analysis.</p> <p>The databases considered are the following ones:</p> <ul> <li><a href="http://www.ee.surrey.ac.uk/CVSSP/banca/">BANCA</a></li> <li><a href="http://www2.ece.ohio-state.edu/~aleix/ARdatabase.html">AR face database</a></li> <li><a href="http://www.nist.gov/itl/iad/ig/frgc.cfm">Face Recognition Grand Challenge version 2</a></li> <li><a href="http://www.nist.gov/itl/iad/ig/focs.cfm">The Good, The Bad and the Ugly</a></li> <li><a href="http://vis-www.cs.umass.edu/lfw">Labeled Faces in the Wild</a></li> <li><a href="http://www.multipie.org">Multi-PIE</a></li> <li><a href="https://www.idiap.ch/dataset/mobio">MOBIO</a></li> <li><a href="http://www.jdl.ac.cn/peal/index.html">CAS-PEAL</a></li> <li><a href="http://www.nist.gov/itl/iad/mig/sre12.cfm">NIST Speaker Recognition Evaluation 2012</a></li> </ul> <p>These scores allow to replicate easily and quickly the plots of the manuscript by using the following package:<br> http://pypi.python.org/pypi/xbob.thesis.elshafey2014</p> <p><br> <strong>Citation</strong></p> <p>If you use this dataset in your publication, we would appreciate that you cite the following thesis:</p> <p>Laurent El Shafey, “Scalable Probabilistic Models for Face and Speaker Recognition”, PhD thesis, 2014.<br> http://publications.idiap.ch/index.php/publications/show/2830</p>
Continental-scale genomic analysis suggests shared post-admixture adaptation in the Americas - Z scores tables
<p><strong>Zscores20Pop. </strong>20pop dataset Z scores by SNP in the three ancestries (A. Africa B. Europe C. America).</p> <p><strong>Zscores1090Ind.</strong> 1090Ind dataset Z scores by SNP in the three ancestries (A. Africa B. Europe C. America).</p>
Data from: INSTRAL: discordance-aware phylogenetic placement using quartet scores
Phylogenomic analyses have increasingly adopted species tree reconstruction using methods that account for gene tree discordance using pipelines that require both human effort and computational resources. As the number of available genomes continues to increase, a new problem is facing researchers. Once more species become available, they have to repeat the whole process from the beginning because updating species trees is currently not possible. However, the de novo inference can be prohibitively costly in human effort or machine time. In this paper, we introduce INSTRAL, a method that extends ASTRAL to enable phylogenetic placement. INSTRAL is designed to place a new species on an existing species tree after sequences from the new species have already been added to gene trees; thus, INSTRAL is complementary to existing placement methods that update gene trees.
Data from: Checkerboard score-area relationships reveal spatial scales of plant community structure
Identifying the spatial scale at which particular mechanisms influence plant community assembly is crucial to understanding the mechanisms structuring communities. It has long been recognized that many elements of community structure are sensitive to area; however the majority of studies examining patterns of community structure use a single relatively small sampling area. As different assembly mechanisms likely cause patterns at different scales we investigate how plant species co-occurrence patterns change with sampling unit scale. We use the checkerboard score as an index of species segregation, and examine species C-score-sampling area patterns in two ways. First, we show via numerical simulation that the C-score-area relationship is necessarily hump shaped with respect to sample plot area. Second we examine empirical C-score-area relationships in arctic tundra, grassland, boreal forest, and tropical forest communities. The minimum sampling scale where species co-occurrence patterns were significantly different from the null model expectation was at 0.1 m2 in the tundra, 0.2 m2 in grassland, and 0.2 Ha in both the boreal and tropical forests. Species were most segregated in their co-occurrence (maximum C-score) at 0.3 m2 in the tundra (0.54 m by 0.54 m quadrats), 1.5 m2 in the grassland (1.2 by 1.2 m quadrats), 0.26 Ha in the tropical forest (71 m by 71 m quadrats), and a maximum was not reached at the largest sampling scale of 1.4 Ha in the boreal forest. The most important finding is that the dominant scales of community structure in these systems are large relative to plant body size, and hence we infer that the dominant mechanisms structuring these communities must be at similarly large scales. This provides a method for identifying the spatial scales at which communities are maximally structured; ecologists can use this information to develop hypotheses and experiments to test scale-specific mechanisms that structure communities.
Data from: Validation of the hospital frailty risk score in a tertiary care hospital in Switzerland: results of a prospective, observational study
Objectives: Recently, the Hospital Frailty Risk Score based on a derivation and validation study in the United Kingdom has been proposed as a low-cost, systematic screening tool to identify older, frail patients who are at greater risk of adverse outcomes and for whom a frailty-attuned approach might be useful. We aimed to validate this Score in an independent cohort in Switzerland. Design: Secondary analysis of a prospective, observational study (TRIAGE study). Setting: One 600-bed tertiary care hospital in Aarau, Switzerland Participants: Consecutive medical inpatients aged 75 years or older that presented to the emergency department or were electively admitted between October 2015 and April 2018. Primary and secondary outcome measures: The primary endpoint was all-cause 30-day mortality. Secondary endpoints were length of hospital stay, hospital readmission, functional impairment, and quality of life measures. We used multivariate regression analyses. Results: Of 4957 included patients, 3150 (63.5%) were classified as low risk, 1663 (33.5%) intermediate risk, and 144 (2.9%) high risk for frailty. Compared to the low-risk group, patients in the moderate risk and high-risk groups had increased risk for 30-day mortality (odds ratio [OR] 2.53, 95%CI 2.09 to 3.06, P<0.001 and OR 4.40, 95%CI 2.94 to 6.57, P<0.001) with overall moderate discrimination (area under the ROC curve 0.66). The results remained robust after adjustment for important confounders. Similarly, we found longer length of hospital stay, more severe functional impairment and a lower quality of life in higher risk group patients. Conclusion: Our data confirms the prognostic value of the Hospital Frailty Risk Score to identify older, frail people at risk for mortality and adverse outcomes in an independent patient population. Trial registration number: ClinicalTrials.gov; Identifier: NCT01768494
Data from: Crowds replicate performance of scientific experts scoring phylogenetic matrices of phenotypes
Scientists building the Tree of Life face an overwhelming challenge to categorize phenotypes (e.g., anatomy, physiology) from millions of living and fossil species. This biodiversity challenge far outstrips the capacities of trained scientific experts. Here we explore whether crowdsourcing can be used to collect matrix data on a large scale with the participation of the non-expert students, or "citizen scientists." Crowdsourcing, or data collection by non-experts, frequently via the internet, has enabled scientists to tackle some large-scale data collection challenges too massive for individuals or scientific teams alone. The quality of work by non-expert crowds is, however, often questioned and little data has been collected on how such crowds perform on complex tasks such as phylogenetic character coding. We studied a crowd of over 600 non-experts, and found that they could use images to identify anatomical similarity (hypotheses of homology) with an average accuracy of 82% compared to scores provided by experts in the field. This performance pattern held across the Tree of Life, from protists to vertebrates. We introduce a procedure that predicts the difficulty of each character and that can be used to assign harder characters to experts and easier characters to a non-expert crowd for scoring. We test this procedure in a controlled experiment comparing crowd scores to those of experts and show that crowds can produce matrices with over 90% of cells scored correctly while reducing the number of cells to be scored by experts by 50%. Preparation time, including image collection and processing, for a crowdsourcing experiment is significant, and does not currently save time of scientific experts overall. However, if innovations in automation or robotics can reduce such effort, then large-scale implementation of our method could greatly increase the collective scientific knowledge of species phenotypes for phylogenetic tree building. For the field of crowdsourcing, we provide a rare study with ground truth, or an experimental control that many studies lack, and contribute new methods on how to coordinate the work of experts and non-experts. We show that there are important instances in which crowd consensus is not a good proxy for correctness.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.