Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,243
datasets available to search
ShareScore release 0.7.1
Dataset results
1,243 results for “Statistics”
Artificial Intelligence and the Future of Smart Cities-Figure 7. Q.9.Which of the following job functions will AI impact the most over the next 10 years? (Statistically significant differences by age for 9.3, 9.4, 9.5 and 9.6)
<p>The analysis reveals that people perceive that AI will have a greater impact over the next 10 years on marketing (for example, intelligent customer targeting, planning and executing marketing campaigns) scored significantly higher (M=4.37, SD=.69) than on finance (for example, robotic financial advisors, automated corporate financial analysis) (M=3.87, SD=.60) (Figure 7). For the same question customer services scored significantly higher (M=3.75, SD=.83) than health (e.g. consultation and diagnosis, surgery) (M=3.25, SD=.83). For the same question, the analyses by gender reveals that the majority of female participants scored significantly higher (M=3.40, SD=.81) than male participants (M=3.18, SD=.57) and those aged in the second group.</p>
Turbulence statistics in smooth wall oscillatory boundary layer flow
<p>Experimental and numerical datasets belonging to Van der A, D.A., Scandura, P., O’Donoghue, T. (2018). Turbulence statistics in smooth wall oscillatory boundary layer flow, Journal of Fluid Mechanics, 849, 192-230. </p> <p> </p>
Final Report on Data Management - Raw data of DC-TRNG for D2.4 statistical testing
<p>Collected Raw & Post-processed data from DC-TRNG for both AIS-31 and NIST800-90B tests suites for Final Report on Data Management</p> <p>The purpose of the final report on data management is to provide an update of the analysis of the main elements of the data management policy used by the applications with regards to all the datasets that were generated by the project. Most important aspects regarding data management, like metadata generation, data preservation, and responsibilities, were updated compared to the initial report D5.2 (Data Management Plan) according to the outcome of the project.</p>
Milan Area Statistics
<p>JSON data related to average statistics on children food habits and physical activities (per Area)</p>
Asti Area Statistics
<p>JSON data related to average statistics on children food habits and physical activities (per area)</p>
Supplemental Materials to Accompany: Abowd and Schmutte "An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices"
<p>Materials to supplement "<a href="https://www.aeaweb.org/articles?id=10.1257/aer.20170627">An Economic Analysis of Privacy Protection and Statistical Accuracy as Social Choices</a>" by John Abowd and Ian Schmutte.</p>
Meandering evolution and width variations: a physics-statistics based modeling approach
<p>Coordinates (x,y) of the central bank (field 1 and 2), the distance of the central axis (field 3), coordinates (x,y) of the left bank (fields 4 and 5), coordinates (x,y) of the right bank (fields 6 and 7)</p>
Zonal Statistics
<p>This data set is composed of two parts each having its proper origins, formats and rights. This data set was used to answer a challenge given by a high school to help students to learn the problem of population density and needs.</p>
Statistical analysis of chlorate occurrence data in food
<p>In accordance with Article 29 (1) (a) of Regulation (EC) No 178/2002, the European Commission asked the European Food Safety Authority (EFSA) in 2014 for a scientific opinion on the risks for human health related to the presence of chlorate in food from all sources, taking also into account its presence in drinking water. The opinion found that “Chronic exposure of adolescent and adult age classes did not exceed the TDI. However, at the 95th percentile the TDI was exceeded in all surveys in ‘Infants’ and ‘Toddlers’ and in some surveys in ‘Other children’. Chronic exposures are of concern in particular in younger age groups with mild or moderate iodine deficiency.” Food manufacturers have started to optimise their manufacturing processes to lower chlorate residue level in foods and the European Commission in 2017 provided a revised Guidance document on good hygiene practices regarding the use of chlorinated disinfectant. It can therefore be expected that chlorate levels in foods are now lower compared with the levels found in the samples from 2011 to 2014. The European Commission (EC) requested EFSA in 2018 to provide an updated statistical analysis on chlorate occurrence levels in foods. A set of 14 Excel tables containing the statistical analysis of reported results for the analysis of chlorates from pesticides monitoring and contaminants monitoring programmes have been prepared. The data is presented at three levels of aggregation using FoodEx product categories. Analysis was performed for two time points, the 2011-2017 dataset contained 15,741 valid results and the 2015-2017 dataset contained 28,033 valid results. Mean, median and percentile (75th, 90th, 95th) for lower bound, middle bound and upper bound concentration values were calculated. Caution should be applied to percentile values calculated from a limited number of results.</p>
Annex B to the technical report on the raw primary commodity (RPC) model - Summary statistics of the output data
<p><strong>The raw primary commodity model</strong>:</p> <p>Dietary exposure is typically calculated by combining food consumption data with occurrence data. EFSA’s food consumption data are stored in the Comprehensive European Food Consumption Database (Comprehensive Database). Some of these data, however, cannot be used in exposure assessments when the occurrence data are reported for the raw primary commodities (RPCs). The RPC model aims to bridge this gap by transforming the Comprehensive Database into RPC consumption data. Using the RPC model, EFSA successfully developed a new RPC Consumption Database, which contains 51 dietary surveys from 23 different countries. These surveys cover a total of 94,532 subjects and 26,573,088 RPC consumption records. The consumption data generated by the RPC model were manually checked and validated by means of case studies. These case studies demonstrated that the RPC consumption data are suitable for assessing dietary exposure to chemicals where the occurrence data are predominantly available for RPCs.</p> <p><strong>Annex B to the technical report on the raw primary commodity model:</strong></p> <p>Annex B is an excel file which presents summary statistics of the output data generated by the RPC model. The following tables are included in Annex B:</p> <p>Table B.1 :Summary statistics of chronic RPC consumption expressed in g/kg bw per day (total population)</p> <p>Table B.2 :Summary statistics of chronic RPC consumption expressed in g/day (total population)</p> <p>Table B.3 :Summary statistics of acute RPC consumption expressed in g/kg bw (consumers only)</p> <p>Table B.4 :Summary statistics of acute RPC consumption expressed in g (consumers only)</p> <p>Table B.5 :Comparison of the RPC consumption data with RPC consumption data used in EFSA's Pesticides Residues Intake Model (PRIMo)</p> <p>Table B.6 :Contribution of processed products to the average chronic RPC consumption</p>
Summary statistics - Imputed gene associations identify replicable trans-acting genes enriched in transcription pathways and complex traits
<p>Summary statistics for all trans-acting/target gene pairs tested in our manuscript.</p> <p>Preprint available: <a href="https://doi.org/10.1101/471748">https://doi.org/10.1101/471748</a></p>
3D Taylor-Green vortex Direct Numerical Simulation statistics from Re=1250 to Re=20000
<p>Statistical data for the 3D Taylor Green flow from Re=1250 to Re=20000 obtained with the flow solver <a href="https://www.incompact3d.com/">Incompact3d</a>. </p> <p># ===========================================================================================<br> # When publishing results using this data, the following paper should be cited as the source: <br> # Thibault Dairay, Eric Lamballais, Sylvain Laizet and John Christos Vassilicos<br> # Numerical dissipation vs. subgrid-scale modelling for large eddy simulation<br> # Journal of Computational Physics 337 (2017) 252–274<br> # https://doi.org/10.1016/j.jcp.2017.02.035<br> # ===========================================================================================</p> <p># Column 1 : time t<br> # Column 2 : kinetic energy E_k [=(u^2+v^2+w^2)/2]<br> # Column 3 : dissipation epsilon_t [=-dE_k/dt]<br> # Column 4 : dissipation epsilon [= nu ((du/dx)^2+(du/dy)^2+(du/dz)^2+(dv/dx)^2+(dv/dy)^2+(dv/dz)^2+(dw/dx)^2+(dw/dy)^2+ dw/dz)^2)]<br> # Column 5 : enstrophy Dzeta [=2 nu epsilon]<br> # Column 6 : mean square u^2<br> # Column 7 : mean square v^2<br> # Column 8 : mean square w^2<br> # Column 9 : mean square (du/dx)^2<br> # Column 10 : mean square (du/dy)^2<br> # Column 11 : mean square (du/dz)^2<br> # Column 12 : mean square (dv/dx)^2<br> # Column 13 : mean square (dv/dy)^2<br> # Column 14 : mean square (dv/dz)^2<br> # Column 15 : mean square (dw/dx)^2<br> # Column 16 : mean square (dw/dy)^2<br> # Column 17 : mean square (dw/dz)^2</p>
Supporting data for: "Diaphysator: an online application for the exhaustive cartography and user-friendly statistical analysis of long bone diaphyses"
<p>Example of dataset to be used with the R-shiny application “Diaphysator”, composed of right tibiae and femora.</p> <p>These data file have been published in: Lacoste Jeanson, A., Santos, F., Villa, C., Banner, J., & Bruzek, J. (2018). Architecture of the femoral and tibial diaphyses in relation to body mass and composition: Research from whole-body CT. <em>American Journal of Physical Anthropology</em>, 167, 813– 826. doi: <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/ajpa.23713">10.1002/ajpa.23713</a></p> <p>This zip file contains:</p> <ul> <li>an “Information file” in CSV format</li> <li>various data files for human femora and tibiae in CSV format</li> </ul> <p>For all CSV files, the field separator is the comma “,” and the character used for decimal points is the dot “.”</p>
Hyperparameter tuning and performance assessment of statistical and machine-learning models using spatial data.
<p>This is a research compendium (RC) for the publication "Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data".</p> <p>The code (including figures, appendices and the manuscript) is packed in <strong>pathogen-modeling-3.zip </strong>or can be found directly in the <a href="https://github.com/pat-s/pathogen-modeling">Github repository</a>.</p> <ul> <li><strong>Publication figures</strong>: analysis/paper/submission/3/latex-source-files/</li> <li><strong>Appendices</strong>: analysis/paper/submission/3/</li> </ul> <p>This RC represents a static snapshot at the time of submission. The Github repository will receive changes after the publication was published.</p> <p><strong>Data sources</strong></p> <ul> <li>Atlas Climatico: <a href="http://opengis.uab.es/wms/iberia/index.htm">http://opengis.uab.es/wms/iberia/index.htm</a></li> <li>DEM: ftp://ftp.geo.euskadi.eus/lidar/MDE_LIDAR_2016_ETRS89/</li> <li>Lithology: <a href="http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home">http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home</a></li> <li>pH: <a href="https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0">https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0</a></li> <li>soil: <a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a></li> </ul> <p><strong>Licenses</strong></p> <p>All files are shared via the given license with the exception of "soil.tif" which is shared via the <strong>ODbL </strong>license<strong>.</strong></p>
Casa do Adolescente (SP) - Statistical Indicators (24/02/2017 - 24/03/2017)
<p>Survey conducted between 03 /February / 2017 - 03 /March/ 2017, with the objective of gathering<br> information in order to start a database in order to serve in the construction of statistical indicators related to the mapping<br> socioeconomic mapping and rates of teenage pregnancies of participating adolescents<br> of the Casa do Adolescente-SP project.</p> <p>The data were collected exclusively in the Casa do Adolescente building and<br> always conducted by professionals in the areas of Health and / or Education.<br> The participation of adolescents occurred during the routine activities of the<br> Casa do Adolescente project, the respondents were informed about<br> the nature of the questions, their objectives and then invited to participate.<br> At no time were the responders identified, preserving<br> identity, secrecy and security with respect to the information collected.<br> The units corresponding to the Casa do Adolescente that participated in the survey were:</p> <p>São Paulo (Pinheiros) -SP and São Paulo (Heliópolis) -SP<br> <br> The total of participants was 105 adolescents and the questions and the corresponding answers are in the database.<br> </p>
Social media statistics
<p>This data was collected/ generated through the periodic monitoring of the project’s social media statistics (including Facebook, Twitter and LinkedIn) with a view to measuring and assessing the performance and results of the project’s social media activity in terms of dissemination and communication.</p>
Replication package for "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Replication package for "Evolution of statistical analysis in ESE research"
<p>This is the replication package for the analysis done in the paper "Evolution of statistical analysis in empirical software engineering research: Current state and steps forward" (DOI: <a href="https://doi.org/10.1016/j.jss.2019.07.002">https://doi.org/10.1016/j.jss.2019.07.002</a>, preprint: <a href="https://arxiv.org/abs/1706.00933">https://arxiv.org/abs/1706.00933</a>).</p> <p>The package includes CSV files with data on statistical usage extracted from 5 journals in SE (EMSE, IST, JSS, TOSEM, TSE). The data was extracted from papers between 2001 - 2015. The package also contains forms, scripts and figures (generated using the scripts) used in the paper.</p> <p>The extraction tool mentioned in the paper is available in dockerhub via: <a href="https://hub.docker.com/r/robertfeldt/sept">https://hub.docker.com/r/robertfeldt/sept</a></p>
Statistical Methods for Identifying Sequence Motifs Affecting Point Mutations
<p>Scripts and derived data for the indicated paper. Links to the <a href="https://zenodo.org/record/1204695">original data</a> and <a href="https://zenodo.org/record/3497585">library source code</a> used to generate this material. Inflate the ENU_mutation_classification.tar.gz archive, remove the ENU_mutation_classification/classifier directory (this will be replaced). Move all others archives into the ENU_mutation_classification directory and inflate them.</p>
Trace-Share Dataset for Evaluation of Statistical Characteristics Preservation
<p>The dataset contains all data used during the evaluation of statistical characteristics preservation. Archives are protected by password "<strong>trace-share</strong>" to avoid false detection by antivirus software.</p> <p>For more information, see the project repository at <strong><a href="https://github.com/Trace-Share">https://github.com/Trace-Share</a></strong>.</p> <p> </p> <p><strong>Selected Attack Traces</strong></p> <p>We selected 72 different traces of network attacks obtained from various internet databases. File names refer to common names of contained vulnerabilities, malware, or attack tools.</p> <p> </p> <p><strong>Background Traffic Data</strong></p> <p>Publicly available dataset <a href="https://www.unb.ca/cic/datasets/ids-2018.html">CSE-CIC-IDS-2018</a> was used as a background traffic data. The evaluation uses data from the day Thursday-01-03-2018 containing a sufficient proportion of regular traffic without any statistically significant attacks. Only traffic aimed at victim machines (range 172.31.69.0/24) is used to reduce less significant traffic.</p> <p> </p> <p><strong>Evaluation Results and Dataset Structure</strong></p> <ul> <li>Traces variants (<a href="https://zenodo.org/record/3553063/files/traces-normalized.zip">traces-normalized.zip</a>, <a href="https://zenodo.org/record/3553063/files/traces-adjusted.zip">traces-adjusted.zip</a>) <ul> <li>./traces-normalized/ — normalized PCAP files and details in YAML format;</li> <li>./traces-adjusted/ — configuration files for traces combination in YAML format.</li> </ul> </li> <li>Computed statistics (<a href="https://zenodo.org/record/3553063/files/statistics.zip">statistics.zip</a>) <ul> <li>./statistics-background/ — background traffic statistics computed by ID2T;</li> <li>./statistics-combination/ — combined traces statistics computed by ID2T for all adjust options (selected only combinations where ID2T provided all statistics files);</li> <li>./statistics-difference/ — computed mean and median differences of background and combined traffic traces.</li> </ul> </li> <li>Evaluation results <ul> <li><a href="https://zenodo.org/record/3553063/files/statistics-difference.ipynb">statistics-difference.ipynb</a> — file containing visualization of statistics differences.</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.