Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
134
datasets available to search
ShareScore release 0.9.0
Dataset results
134 results for “random sample”
Stiffness of randomly sampled stainless steel frames under gravity and gravity plus wind load scenarios
<p>Data was generated using the general purpose finite element software ABAQUS and performing advanced nonlinear analyses. The database is comprised of vertical and lateral system stiffness values corresponding to different random samples of six different nominal stainless steel frames under gravity and gravity plus wind load combinations. The values of the random variable assignments are given for each case. </p> <p>The full details of the finite element model can be found in: Arrayago, I.; Rasmussen, K.J.R. Reliability of stainless steel frames designed using the Direct Design Method in serviceability limit states. Journal of Constructional Steel Research 196, 107425, 2022. DOI: https://doi.org/10.1016/j.jcsr.2022.107425</p> <p>The data included in the dataset corresponds to the vertical & lateral stiffness of each frame under different load conditions.</p> <p>Although the data has been generated using the finite element software ABAQUS, no special software is required to read or interpret the data.</p>
Omnibus Poll Ukraine - August 2023 (Ilko Kucheriv Democratic Initiatives Foundation + Razumkov Centre) – Random-sample questionnaire-based representative poll
This data collection offers a representative omnibus survey of the Ukrainian population, living in territories controlled by the Ukrainian government without ongoing armed hostilities. The survey was conducted by the Ilko Kucheriv Democratic Initiatives Foundation together with the sociological service of the Razumkov Center from 09 to 15 August 2023. The survey was conducted using a stratified multi-stage sample. The structure of the sample reflects the demographic structure of the adult population of the surveyed territories as of the beginning of 2022 (by age, gender, type of settlement). 2019 respondents aged 18 and older were interviewed. The theoretical sampling error does not exceed 2.3%. At the same time, additional systematic sample deviations may be caused by the consequences of Russian aggression, in particular, the forced evacuation of millions of citizens. The survey covers five thematic fields: assessment of the current situation in the country, the Russian war of aggression, energy sector, corruption, volunteering. This data collection contains the original survey data. The SPSS file (.sav) is the original file provided by the Ilko Kucheriv Democratic Initiatives Foundation. It has been exported into an Excel file. The content of the respective xlsx-file should be identical with the original sav-file. The sav-file contains the questions and answer options of the original questionnaire in Ukrainian. The original questionnaire and an English translation are also included in this data collection as separate pdf-file. Additionally, the data collection contains three files with "selected results" which document some major results of the survey in the form of analytical summaries and descriptive statistics: two in English, covering assessment of the current situation in the country + the Russian war of aggression as well as volunteering; one in Ukrainian covering corruption. New in version 1.1: The numbering of questions in the separate questionnaire (file "DIF_CR_0823-questionnaire-revised.pdf") has been adjusted to the numbering in the original data file ("DIF_CR_0823.sav"). A third file with "selected results" has been added. New in version 1.2: An English translation of the questionnaire has been added under "files".
Omnibus Poll Ukraine - July 2023 (Ilko Kucheriv Democratic Initiatives Foundation + Kyiv International Institute of Sociology) – Random-sample questionnaire-based representative poll
This data collection offers a representative omnibus survey of the Ukrainian population, living in territories controlled by the Ukrainian government without ongoing armed hostilities. The survey was conducted by the Ilko Kucheriv Democratic Initiatives Foundation together with the Kyiv International Institute of Sociology from 03 to 17 July 2023. A description of the methodology is given on p.2 of the "selected results" file, which is part of this data collection. The poll covers the following thematic fields: jobs + entrepreneurship, corruption, economic situation, healthcare sector, war, people under Russian occupation. This data collection contains the original survey data. The SPSS file (.sav) is the original file provided by the Ilko Kucheriv Democratic Initiatives Foundation. It has been exported into an Excel file. The content of the respective xlsx-file should be identical with the original sav-file. The sav-file contains the questions and answer options of the original questionnaire in Ukrainian. The original questionnaire and an English translation are also included in this data collection as separate pdf-files. Additionally, the data collection contains one file with "selected results" which document some major results of the survey in the form of a analytical summaries and descriptive statistics and another file with a clarification concerning the interpretation of question 5.24 about the president's "personal responsibility" for corruption in the country. These files are in Ukrainian only. New in version 1.1: An English translation of the questionnaire has been added under "files".
ESM Atlas v0 representative random sample of predicted protein structures
<p>A representative random sample of the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> Sample size: 997,405.</p>
ESM Atlas v0 random sample of high confidence predicted protein structures
<p>A random sample out of the 225M high confidence predictions in the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> High confidence is defined as mean pLDDT > 0.7 and pTM > 0.7 and corresponds to ∼36% of the total 617M proteins folded.<br> This is the random sample used for analysis in the paper as well as visualization on the <a href="http://esmatlas.com/">esmatlas.com</a> Explore page.<br> Sample size: 999,520 based on 999,996 unique randomly sampled IDs and 0.05% missing data in the processing pipeline.</p>
Large Uniform Random SAT Samples
<p>Large Random SAT samples generated with the following <em>samplers</em>:</p> <ul> <li>BDDSampler</li> <li>Spur</li> <li>QuickSampler </li> <li>KUS</li> <li>Unigen2</li> <li>Smarch </li> </ul>
USPTO patent data: 250k random sample and NPEs' patents
<p>This package includes:</p> <ol> <li>Two Stata .dta files consisting of information on patents assigned by the United States Patent and Trademark Office between 1976 and 2014: a random sample of 250 000 US patents, and data on patent owned by Intellectual Ventures, RPX, and several other companies. The variables for example include: grant date, application date, forward and backward citations, renewals, claims and others.</li> <li>Source codes and methods used in generating and analyzing the two data files.</li> </ol> <p>A bachelor thesis with further information will be linked here.</p>
Random Sample of Open Source Ventilators
<p>This is a random sample of references to open source ventilators based on the data published by the PubInv project (https://github.com/PubInv/covid19-vent-list CC0-1.0). It was used to verify technology and documentation readiness scales introduced in a CIRP design conference paper. This upload shall make the data set used for the conference paper publicly accessible.</p>
Random sample of habitat suitability for wolves in Scotland
<p>A rule-based habitat suitability model was created for wolves (<em>Canis lupus</em>) in mainland Scotland. Six variations of the model were run, in order to test sensitivity to changes in input values. In order to test for difference in the outputs of the six models, 500 random points were sampled from all six models, and then a test for statistical difference performed. This dataset constitutes the values of the 500 random points, where 1 indicates complete suitability and 0 indicates complete unsuitability.</p>
Randomly sampled coefficients for synchrotron radiative transfer in the Stokes basis, power law model, computed by rimphony, for consumption by neurosynchro
<p>This directory contains a training set of 22 million randomly-sampled radiative transfer coefficients generated by <a href="https://github.com/pkgw/rimphony/">rimphony</a>, suitable for use with the <a href="https://github.com/pkgw/neurosynchro/">neurosynchro</a> package. These coefficients can be used for numerical radiative transfer of synchrotron emission in the Stokes basis with a package such as <a href="https://github.com/jadexter/grtrans/">grtrans</a>.</p> <p>In this particular dataset, coefficients were computed using a model of a power law electron distribution isotropic in pitch angle. The input parameters, which were sampled randomly in a three-dimensional space, are:</p> <ul> <li><em>s</em>, the harmonic number, dimensionless, sampled logarithmically between 5 and 50,000,000.</li> <li><em>theta</em>, the angle between the ray path and the local magnetic field, measured in radians, sampled linearly between 0.001 and π/2 (namely, 1.5707963267948966).</li> <li><em>p</em>, the power-law index of the energetic electrons, dimensionless, sampled linearly between 1.5 and 7.</li> </ul> <p>The coefficients were computed on Harvard’s Odyssey cluster using Git commit <a href="https://github.com/pkgw/rimphony/commit/772161ebda0217b8c1ccb8ce3801ad9dc3701a4f">772161</a> of rimphony. A total of about 5,000 CPU hours were used, with 500 processes running for about 10 hours each. There are 2,748,835 data rows in total. The data are provided in their original format, split among 500 files, so that smaller subsamples of the data may be loaded easily. A README.md file provides more detailed information.</p>
A Methodology for the Fast Identification and Monitoring of Microplastics in Environmental Samples using Random Decision Forest Classifiers
<p>This short video shows the results of the application of a classifier for microplastics as described by Hufnagl et al. (2019).</p> <p> </p> <p>If you reuse this video please cite</p> <p> </p> <p>Hufnagl, B., Steiner, D., Renner, Löder, M. G. J., Laforsch, C. and Lohninger, H. <em>A Methodology for the Fast Identification and Monitoring of Microplastics in</em><em> Environmental Samples using Random Decision Forest Classifiers,</em> Analytical Methods, 2019, DOI:10.1039/C9AY00252A</p>
How Confidence in Prior Attitudes, Social Tag Popularity, and Source Credibility Shape Confirmation Bias Toward Antidepressants and Psychotherapy in a Representative German Sample: Randomized Controlled Web-Based Study
<p>ABSTRACT</p> <p>Background: In health-related, Web-based information search, people should select information in line with expert (vs nonexpert) information, independent of their prior attitudes and consequent confirmation bias.</p> <p>Objective: This study aimed to investigate confirmation bias in mental health–related information search, particularly (1) if high confidence worsens confirmation bias, (2) if social tags eliminate the influence of prior attitudes, and (3) if people successfully distinguish high and low source credibility.</p> <p>Methods: In total, 520 participants of a representative sample of the German Web-based population were recruited via a panel company. Among them, 48.1% (250/520) participants completed the fully automated study. Participants provided <em>prior attitudes</em> about antidepressants and psychotherapy. We manipulated (1) <em>confidence</em> in prior attitudes when participants searched for blog posts about the treatment of depression, (2) <em>tag popularity</em> —either psychotherapy or antidepressant tags were more popular, and (3) <em>source credibility</em> with banners indicating high or low expertise of the tagging community. We measured <em>tag</em> and <em>blog post</em> selection, and <em>treatment</em><em>efficacy ratings</em> after navigation.</p> <p>Results: Tag popularity predicted the proportion of selected antidepressant tags (beta=.44, SE 0.11; <em>P</em><.001) and blog posts (beta=.46, SE 0.11; <em>P</em><.001). When confidence was low (−1 SD), participants selected more blog posts consistent with prior attitudes (beta=−.26, SE 0.05; <em>P</em><.001). Moreover, when confidence was low (−1 SD) and source credibility was high (+1 SD), the efficacy ratings of attitude-consistent treatments increased (beta=.34, SE 0.13; <em>P</em>=.01).</p> <p>Conclusions: We found correlational support for defense motivation account underlying confirmation bias in the mental health–related search context. That is, participants tended to select information that supported their prior attitudes, which is not in line with the current scientific evidence. Implications for presenting persuasive Web-based information are also discussed.</p> <p>Trial Registration: ClinicalTrials.gov NCT03899168; https://clinicaltrials.gov/ct2/show/NCT03899168 (Archived by WebCite at http://www.webcitation.org/77Nyot3Do)</p> <p>J Med Internet Res 2019;21(4):e11081</p> <p>doi:10.2196/11081</p>
Monthly word embeddings for Twitter random sample (English, 2012-2018)
<p>This dataset contains monthly word embeddings created from the tweets available via the statuses/sample endpoint of the Twitter Streaming API from 2012 to 2018. Full details of the creation of the dataset are given in <a href="https://www.aclweb.org/anthology/D19-1007/">Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings</a>. </p> <p>The md5sum of the gzipped tarball file is a76888ffec8cc7aebba09d365ca55ace .</p>
Vitamin B12 deficiency anaemia and gestational diabetes mellitus: a two-sample Mendelian randomization study
Open the record for dataset details and reuse information.
A new sampling capability for uncertainty quantification in the Ice-sheet and Sea-level System Model v4.19 using Gaussian Markov random fields -- Datasets and results
<p>Data archives for test experiments (Section 3) and Pine Island Glacier application (Section 4) from the manuscript "Kevin Bulthuis and Eric Larour, A new sampling capability for uncertainty quantification in the Ice-sheet and Sea-Level System Model v4.19 using Gaussian Markov random fields"</p> <p>Source code is available at https://doi.org/10.5281/zenodo.5532775.</p>
Uniform Random SAT Samples
<p>Random SAT samples generated with the following <em>samplers</em>:</p> <ul> <li>Spur</li> <li>QuickSampler </li> <li>Unigen2</li> <li>Smarch </li> </ul>
Male-female disparity in clinical features and significance of mild vertebral fractures in community-dwelling residents aged 50 and over: A Japanese cohort survey randomly sampled from a basic resident registry
Open the record for dataset details and reuse information.
Investigating the causal association between immune cell phenotypes and allergic diseases and non-allergic asthma using conventional Two-sample and Bayesian weighted Mendelian randomization
<p>Investigating the causal association between immune cell phenotypes and allergic diseases and non-allergic asthma using conventional Two-sample and Bayesian weighted Mendelian randomization</p>
Ultimate load of randomly sampled stainless steel frames under gravity plus wind loads
<p>Data was generated using the general purpose finite element software ABAQUS and performing advanced nonlinear analyses. The database is comprised of ultimate load factors corresponding to different random samples of six different nominal stainless steel frames under gravity and wind load combinations. The values of the random variable assignments are given for each case.</p> <p>The full details of the finite element model can be found in: Arrayago, I.; Rasmussen, K.J.R.; Zhang, H. System-based reliability analysis of stainless steel frames subjected to wind loads. "Structural Safety", July 2022, vol. 97, art. No. 102211.</p> <p>DOI: https://doi.org/10.1016/j.strusafe.2022.102211</p>
Surrogate-modelling & machine learning dataset : finite element stress analysis of biaxial specimen with random elastic properties - 1000 samples
<p>Dataset finite element stress analysis of biaxial specimen with random elastic properties</p> <p>Unzip and execute dataset.py to visualise data samples. PyVista is needed.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.