Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
20
datasets available to search
ShareScore release 0.9.0
Dataset results
20 results for “Exploratory data analysis”
Handling of Personal Data by Smart Home Equipment: an Exploratory Analysis in the Context of LGPD
<p>This dataset provides data about an exploratory research that analyzed the Privacy and Security Policies and the Instruction Manuals of 59 home automation equipment for Smart Home in order to verify which personal data was handled and how these documents were providing information about processes performed in personal data. The analysis was conducted with a quantitative approach followed by a qualitative analysis, using content analysis.</p>
Data from: Research and exploratory analysis driven - time-data visualization (read-tv) software
<strong><em>read-tv</em></strong> <p>The main paper is about, <em>read-tv</em>, open-source software for longitudinal data visualization. We uploaded sample use case surgical flow disruption data to highlight <em>read-tv</em>'s capabilities. We scrubbed the data of protected health information, and uploaded it as a single CSV file. A description of the original data is described below.</p> Data source <p>Surgical workflow disruptions, defined as "<i>deviations from the natural progression of an operation thereby potentially compromising the efficiency or safety of care", </i>provide a window on the systems of work through which it is possible to analyze <u>mismatches between the work demands and the ability of the people to deliver the work</u>. They have been shown to be sensitive to different intraoperative technologies, surgical errors, surgical experience, room layout, checklist implementation and the effectiveness of the supporting team. The significance of flow disruptions lies in their ability to provide a hitherto unavailable perspective on the quality and efficiency of the system. This allows for a systematic, quantitative and replicable assessment of risks in surgical systems, evaluation of interventions to address them, and assessment of the role that technology plays in exacerbation or mitigation.</p> <p>In 2014, Drs Catchpole and Anger were awarded NIBIB R03 EB017447 to investigate flow disruptions in Robotic Surgery which has resulted in the detailed, multi-level analysis of over 4,000 flow disruptions. Direct observation of 89 RAS (robitic assisted surgery) cases, found a mean of 9.62 flow disruptions per hour, which varies across different surgical phases, predominantly caused by coordination, communication, equipment, and training problems.</p>
Exploratory Landscape Analysis Feature Data of Recombinations and noiseless BBOB Instances 1-15
<p>Data presented in the paper 'Increasing Diversity of Benchmark Functions Sets by Affine Recombination' which was submitted to PPSN 2022.</p>
Funding Covid-19 research: Insights from an exploratory analysis using open data infrastructures - Supplementary material
<p>This dataset contains supplementary material for the paper 'Funding Covid-19 research: Insights from an exploratory analysis using open data infrastructures' by Alexis-Michel Mugabushaka, Nees Jan van Eck, and Ludo Waltman.</p> <ul> <li>supplementary_material_1_dataset.ods: Dataset of Covid-19 publications.</li> <li>supplementary_material_2_sample.ods: Samples of publications used to assess the accuracy of funding data in the different databases.</li> <li>supplementary_material_3_tables_and_figures.ods: Statistics underlying the tables and figures presented in the paper.</li> </ul>
Code and Data Supplement for Using feature importance as exploratory data analysis tool on earth system models
<p>This contains:</p> <ul> <li>Code for all analyses in</li> <li>E3SM data</li> </ul> <p>For the paper Using <em>feature importance as exploratory data analysis tool on earth system models.</em></p>
Data from: Research and exploratory analysis driven - time-data visualization (read-tv) software
Open the record for dataset details and reuse information.
Sociotechnical Dynamics in Open Source Smart Contract Repositories: An Exploratory Data Analysis of Curated High Market Value Projects
<p>This is the replication package for the paper “Sociotechnical Dynamics in Open Source Smart Contract Repositories: An Exploratory Data Analysis of Curated High Market Value Projects”.</p> <p>In project_curation_selection, there is the curation process of the 100 selected projects including the identification of GitHub repositories and classification of evolution scenarios. </p> <p>In distribution_commits_issues_contributors_market_value_before_after_deploy, data collection from GitHub projects includes the distribution of total commits, contributors, and issues before and after deployment of each investigated project. </p> <p>In analysis_commit_messages, there is qualitative analysis of commit message content from all investigated projects. </p> <p>In the analysis_contributors section, the data focuses on analyzing the profiles of each GitHub contributor involved in the investigated projects.</p> <p>In analysis_market_value_by_project, data refers to the market value and volume of each investigated project. </p> <p>In codes, there are scripts used to obtain the analyzed data.</p> <p> </p>
Exploratory Data Analysis of SonarCloud and GitHub Data
Open the record for dataset details and reuse information.
Data and scripts from: Exploratory analysis of multi-trait coadaptations in the light of population history
<p><span>During the process of range expansion, populations encounter a variety of environments. They respond to the local environments by modifying their mutually interacting traits. Common approaches of landscape analysis include first focusing on the genes that undergo diversifying selection or directional selection in response to environmental variation. To understand the whole history of populations, it is ideal to capture the history of their range expansion with reference to the series of surrounding environments and to infer the multi-trait coadaptation. To this end, we propose a complementary approach; it is an exploratory analysis using up-to-date methods that integrates population genetic features and features of selection on multiple traits. First, we conduct correspondence analysis of site frequency spectra, traits and environments with auxiliary information of population-specific fixation index (FST). This visualizes the structure and the ages of populations and helps infer the history of range expansion, encountered environmental changes and selection on multiple traits. Next, we further investigate the inferred history using an admixture graph that describes the population split and admixture. Finally, principal component analysis of the </span><span>selection on edge-by-trait (SET) matrix identifies multi-trait coadaptation and the associated edges of the admixture graph. We introduce a newly defined factor loadings of environmental variables in order to identify the environmental factors that caused the coadaptation. A numerical simulation of one-dimensional stepping-stone population expansion showed that the exploratory analysis reconstructed the pattern of the environmental selection that was missed by analysis of individual traits. Analysis of a public dataset of natural populations of black cottonwood in northwestern America identified the first principal component (PC) coadaptation of photosynthesis- vs growth-related traits responding to the geographical clines of temperature and daylength. The second PC coadaptation of volume-related traits suggested that soil condition was a limiting factor for above-ground environmental selection.</span></p>
Exploratory Data Analysis for Disease Pedigrees and Cancer Genetics
ClinicalTrials.gov study NCT00339508. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Data from: An exploratory association analysis of the insulin gene region with diabetes mellitus in two dog breeds
Open the record for dataset details and reuse information.
Data and scripts from: Exploratory analysis of multi-trait coadaptations in the light of population history
Open the record for dataset details and reuse information.
Experimental Data Set for the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy"
<p>This are the feature values used in the study "Exploratory Landscape Analysis is Strongly Sensitive to the Sampling Strategy".</p> <p>The dataset regroups feature values for every "cheap" features available in the R package <em>flacco </em>and are computed using 5 sampling strategies and in dimension <span class="math-tex">\($d=5$\)</span>:</p> <ol> <li>Random: the classical Mersenne-Twister algorithm;</li> <li>Randu: a random number generator that is notoriously bad;</li> <li>LHS: a centered Latin Hypercube Design;</li> <li>iLHS: an improved Latin Hypercube Design;</li> <li>Sobol: points extracted from a Sobol' low-discrepancy sequence.</li> </ol> <p>The csv file <em>features_summury_dim_5_ppsn.csv </em>regroups 100 values for every features whereas <em>features_summury_dim_5_ppsn_median.csv </em>regroups for every feature the median of the 100 values.</p> <p>In the folder <em>PPSN_feature_plots</em> are the histograms of feature values on the 24 COCO functions for 3 sampling strategies: Random, LHS and Sobol.</p> <p>The Python file <em>sampling_ppsn.py</em> is the code used to generate the sample points from which the feature values are computed.</p> <p>The file <em>stats50_knn_dt.csv</em> provide the raw data of median and IQR (inter quartile interval) for the heatmaps and boxplots available in the paper.</p> <p>Finally, the files <em>results_classif_knn100.csv</em> (resp. dt) provide the accuracy of 100 classifications for every settings.</p> <p> </p>
Data from: Evidence assessing the diagnostic performance of medical smartphone apps: a systematic review and exploratory meta-analysis
Objective: The number of mobile applications addressing health topics is increasing. Whether these apps underwent scientific evaluation is unclear. We comprehensively assessed papers investigating the diagnostic value of available diagnostic health applications using in-built smartphone-sensors. Methods: Systematic Review - Medline, Scopus, Web of Science inclusive Medical Informatics and Business Source Premier (by citation of reference) were searched from inception until December 15th, 2016. Checking of reference lists of review articles and of included articles complemented electronic searches. We included all studies investigating a health application that used in-built sensors of a smartphone for diagnosis of disease. The methodological quality of 11 studies used in an exploratory meta-analysis was assessed with the QUADAS-2 tool and the reporting quality with the STARD statement. Sensitivity and specificity of studies reporting two-by-two tables were calculated and summarized. Results We screened 3'296 references for eligibility. Eleven studies, most of them assessing melanoma screening apps, reported 17 two-by-two tables. Quality assessment revealed high risk of bias in all studies. Included papers studied 1'048 subjects (758 with the target conditions and 290 healthy volunteers). Overall, the summary estimate for sensitivity was 0.82 (95 % confidence interval (CI); 0.56 to 0.94) and 0.89 (95 %CI; 0.70 to 0.97) for specificity. Conclusions The diagnostic evidence of available health apps on Apple's and Google's app stores is scarce. Consumers and healthcare professionals should be aware of this when using or recommending them.
An Additional Analysis of Data From the PARADIGM Exploratory Study (NCT02394834) in Patients With Advanced/Recurrent Colorectal Cancer
ClinicalTrials.gov study NCT05030493. IPD Sharing: YES. Countries: 1. Publications: 0.
Data from: Evidence assessing the diagnostic performance of medical smartphone apps: a systematic review and exploratory meta-analysis
Open the record for dataset details and reuse information.
Targeted proteomics analysis of Cutaneous Lupus Erythematosus patient interstitial skin fluid and plasma (Neuro Exploratory panel data)
GEO Series GSE182300. Homo sapiens. 38 samples. Type: Other.
Data from: Exploratory proteomic analysis implicates the alternative complement cascade in Primary CNS Vasculitis
Objective: To identify molecular correlates of primary angiitis of the central nervous system (PACNS) through proteomic analysis of cerebrospinal fluid (CSF) from a biopsy-proven patient cohort. Methods: Using mass spectrometry, the CSF proteome of biopsy-proven PACNS patients (n=8) was quantitatively compared to CSF from individuals with non-inflammatory conditions (n=11). Significantly enriched molecular pathways were identified using a gene ontology workflow, and high confidence hits within enriched pathways (fold change >1.5 and concordant Benjamini-Hochberg-adjusted p-value <0.05 on DeSeq and t-test) were identified as differentially regulated proteins. Results: Compared to non-inflammatory controls, 283 proteins were differentially expressed in PACNS patient CSF, with significant enrichment of the complement cascade pathway (C4-binding protein, CD55, CD59, properdin, complement C5, complement C8 and complement C9) and neural cell adhesion molecules. A subset of clinically relevant findings was validated by western blot and commercial ELISA. Conclusions: In this exploratory study we found evidence of deregulation of the alternative complement cascade in CSF from biopsy-proven PACNS as compared to non-inflammatory controls. More specifically, several regulators of the C3 and C5 convertases and components of the terminal cascade were significantly altered. These preliminary findings shed light on a previously unappreciated similarity between PACNS and systemic vasculitides, especially Anti-Neutrophil Cytoplasmic Antibody (ANCA)-associated vasculitis. The therapeutic implications of this common biology, and the diagnostic and/or therapeutic utility of individual proteomic findings warrant validation in larger cohorts.
Data from: Exploratory proteomic analysis implicates the alternative complement cascade in Primary CNS Vasculitis
Open the record for dataset details and reuse information.
Data analysis of tools functionalities for an exploratory data webpage
<p>This publication contains the results of the data analysed for the study published by Calvera-Isabal, Santos & Hernández-Leo (2023). It includes the dataset collected from the workshops developed with teachers and TEL experts.</p> <p>This work has been funded by PID2020-112584RB-C33 funded by MCIN/AEI/10.13039/501100011033, the CS Track project, EU Horizon 2020 programme [grant agreement No 872522], grant for activities to increase the social impact of research in 2021 from Universitat Pompeu Fabra (UPF) and H2O Learn project PID2020-112584RB-C33 funded by MCIN/ AEI / 10.13039/501100011033.</p> <p>Please, contact miriam.calvera@upf.edu for data access.</p> <p>Calvera-Isabal, M., Santos, P., & Hernández-Leo, D. (2023). Towards Citizen Science-Inspired Learning Activities: The Co-design of an Exploration Tool for Teachers Following a Human-Centred Design Approach. International Journal of Human–Computer Interaction, 1-22.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.