Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,694

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

4,694 results for “data analysis”

Learn how ShareScore rates datasets ↗
zenodo32/100

Data for the thesis "Exploring Heuristics for Predicting Microbenchmark Stability and Code Coverage using Static Code Analysis"

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Replication data for the paper "Leveraging Large Language Models for Comprehensive Psychological Analysis: Insights from Four Theoretical Frameworks"

<p>This is a replication data for the paper titled "Leveraging Large Language Models for Comprehensive Psychological Analysis: Insights from Four Theoretical Frameworks" submitted for a blind review.</p> <p>Abstract</p> <p>The rapid advancement of generative Artificial Intelligence (AI) has significantly transformed various research domains. This paper introduces a novel, fully automated methodology for applying Large Language Models (LLMs) to psychological text analysis. The approach includes prompt design for zero-shot and few-shot learning, model internal consistency analysis, autonomous machine evaluation, and additional human validation. Applied to four psychological theories&mdash;Self-Determination Theory, the Big Five Personality Traits, Psychological Well-being, and Cognitive Behavioral Therapy&mdash;this methodology is tested on a dataset of 25,780 emails written by a senior executive (called Person X) over 16 years. The analysis involves extracting psychological characteristics from the emails and regressing these characteristics against personal, professional, and environmental factors. The results demonstrate that the methodology provides unique insights into the examined psychological theories, offering a detailed understanding of how various factors influence psychological states and traits over time. This research highlights the potential of LLMs in capturing and analyzing complex psychological patterns in large text corpora, contributing a robust framework for future studies and practical applications in psychological assessment and intervention. The findings underscore the transformative impact of generative AI in psychological research, opening new avenues for understanding human behavior through advanced language models.</p> <p>The zipped file contains five csv files:</p> <ol> <li>Email_classification-csv: LLM (GPT-3.5 Turbo) classification of 25,780 emails for four psychological theories: SDT, Big Five, PWB and CBT.</li> <li>SDT_regression_data.csv</li> <li>Big_Five_regression_data.csv</li> <li>PWB_regression_data.csv</li> <li>CBT_regression_data.csv</li> </ol> <p>For 2-5 files the dependent variable is monthy percentage share of emails the were assigned a given value for categories of one of the four psychological theories analyzed.&nbsp;</p> <p>Linear regression model has been applied, where dependent variable is the percentage of emails in a specified category that assigned a specific value in this category. For example in Big Five Traits Model, for the Openness category, for each month we calculated percentage of emails that exhibit <em>High</em> or <em>Low</em> openness, or <em>None</em> if the content of the email does not provide enough information to assess whether the specific need is relevant. Two dependent variables were created: <em>Openness-high</em> and <em>Openness-low</em> and regressed on all independent variables. Regressions were not run for the <em>None</em> values.</p> <p>Descriptions of independent variables:</p> <p>- <em>income_index</em>: Person X salary income and consulting fees in a given month, normalized to [0,1].</p> <p>- <em>card_spending</em>: Person X credit card expenditures in a given month, normalized to [0,1].</p> <p>- <em>abroad_far</em>: dummy variable set to 1 for months when Person X worked in Central Asia</p> <p>- <em>abroad_near</em>: dummy variable set to 1 when Person X worked in other EU country</p> <p>- <em>death_1_war</em>: variable set to 1 in a month when Person X&rsquo; farther in law passed away. In the same month Russia invaded Ukraine. The variable was set to .75 in the following month, and to .5 in the month after that.</p> <p>- <em>death_2</em>: variable set to 1 in a month when Person X&rsquo; mother passed away. The variable was set to .75 in the following month, and to .5 in the month after that.</p> <p>- <em>court_case</em>: dummy variable set to 1 for months with the emotionally engaging inheritance court case involving other family members.</p> <p>- <em>BIG4_partner</em>: dummy variable set to 1 for months when Person X worked as a partner in BIG4 accounting firm, which resulted in adopting a professional activity sharply different from the usual Person X habits.</p> <p>- <em>AI_company</em>: dummy variable set to 1 for months when Person X worked as C-level executive at a company specializing in artificial intelligence.</p> <p>- <em>elections</em>: dummy variable set to 1 for months when Person X unsuccessfully run in parliamentary elections</p> <p>- <em>covid_lockdown</em>: dummy variable set to 1 for month where Polish government imposed tough measures during two covid lockdowns.</p> <p>- <em>no_receive</em>: number of different email recipients each month, normalized to [0,1].</p> <p>- <em>avg_length</em>: average number of words in emails sent each month, normalized to [0,1].</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; While the email data was collected for January 2008 &ndash; March 2014 period, financial data was available from October 2009. There were some months where no emails with more than 10 words were sent, yielding 166 monthly observations used for regressions, before removing outliers.</p> <p>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Independent variables were tested for multicollinearity, outlier months were removed, regressions were estimated with robust standard errors, and a range of standard tests were conducted for normality and autocorrelation of residuals, confirming good statistical properties of estimated models.</p> <p>Due to privacy concerns, the email texts cannot be publicly shared. However, the classifications of psychological categories derived from the email texts, along with all other relevant data, are made publicly available in this open access repository, with the consent of email author.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Supplementary file 2; Photographs of various camp survey locations that illustrate various characteristics of the study site and points of rationale for the approach taken to data collection, analysis and interpretation

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Density field data, VTK files and Jupyter notebook for dislocation embedding analysis (with periodic boundary condition)

<p><strong>pbc_dataset.zip</strong></p> <p>The density field data is only for 0 degrees of misorientation with loading directions along [100], [110], [111], [234].&nbsp;</p> <p>Folder paths for low density and low resolution simulation are:</p> <ol> <li>0deg/dir100/2.5e+13/config3/10x10x10</li> <li>0deg/dir110/2.5e+13/config3/10x10x10</li> <li>0deg/dir111/2.5e+13/config3/10x10x10</li> <li>0deg/dir234/2.5e+13/config3/10x10x10</li> </ol> <p>Folder paths for high density and high resolutions are:</p> <ol> <li>0deg/dir100/1e+14/config1/20x20x20</li> <li>0deg/dir110/1e+14/config1/20x20x20</li> <li>0deg/dir111/1e+14/config1/20x20x20</li> <li>0deg/dir234/1e+14/config1/20x20x20</li> </ol> <p>Each set of simulation has 2000 density field data files. Total files : 16000</p> <p>Please ensure that above paths are entered in the Jupyter notebook script file.</p> <p>&nbsp;</p> <p><strong>vtk.zip</strong></p> <p>VTK files for each simulation. Each simulation has 2000 files.&nbsp;</p> <p><strong>Dislocation_embeddings_pbc.ipynb</strong></p> <p>Jupyter notebook to generate dislocation embeddings for uploaded dataset.</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Data for "Evaluating Metal-Organic Precursors for Focused Ion Beam Induced Deposition through Solid-Layer Decomposition Analysis"

<p>Experimental Data for "Evaluating Metal-Organic Precursors for Focused Ion Beam Induced Deposition through Solid-Layer Decomposition Analysis"</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The collected experimental data for tested four different metal-organic precursor layers:</p> <p>&nbsp;</p> <p>1 - [Cu<sub>2</sub>(&mu;-O<sub>2</sub>C<sup>t</sup>Bu)<sub>4]n</sub></p> <p>2 - [Cu<sub>2</sub>(NH<sub>2</sub>(NH=)CC<sub>2</sub>F<sub>5</sub>)<sub>2</sub>(&micro;-O<sub>2</sub>CC<sub>2</sub>F<sub>5</sub>)<sub>4</sub>]</p> <p>3 - [Cu<sub>2</sub>(&micro;-O<sub>2</sub>CC<sub>2</sub>F<sub>5</sub>)<sub>4</sub>]</p> <p>4 - [Ag<sub>2</sub>(&mu;-O<sub>2</sub>CC<sub>2</sub>F<sub>5</sub>)<sub>2</sub>]</p> <p>&nbsp;</p> <p>BSE.zip - includes SEM BSE images of the precursor layers</p> <p>EDX.zip - inludes SEM EDX results of the precursor layers together with exemplary python jupyter notebook to analyze EDX hyperspectral data</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Data and statistical analysis scripts for manuscript on X-ray Microscopy of pennycress seeds

<p>Data and R statistical analysis code for manuscript on X-ray Microscopy of pennycress seeds</p> <blockquote> <p><strong>Evaluation of 3D seed structure and cellular traits in-situ using X-ray microscopy</strong></p> </blockquote> <p>The following files contains:</p> <ul> <li><code>Griffiths_et_al_2024_SeedXRM.R</code> - R statistics script for data processing of raw output from X-ray microscopy data and Marvin Seed analyzer data</li> <li><code>Raw_Data.zip</code> - Raw tabular data for use with R script from X-ray microscopy</li> <li><code>Data.zip</code> - Pre-processed tabluar data for use with R script</li> <li><code>Output.zip</code> - Output files that are generated from the R script</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Data for: High-throughput micro-CT analysis identifies sex-dependent biomarkers of erosive arthritis in TNF-Tg mice and differential response to anti-TNF therapy

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Single and few cell analysis for correlative light microscopy, metabolomics, and targeted proteomics (Data)

<p>Combined data for the manuscript `Single and few cell analysis for correlative light microscopy, metabolomics, and targeted proteomics` for all manuscript and supplemental information figures.</p> <p>Every folder contains the raw data and Jupyter notebook (python) for graph creation.</p> <p>Images are not enclosed but are shown in the manuscript.</p> <p>&nbsp;</p>

opengpl-3.0-or-laterJun 2024View details →
dryad32/100

Coded respondent survey data analysis of the socio-economic status of jasmine growers in Huvina Hadagali

<p><strong>Background:</strong></p> <p>This socio-economic analysis studies the influence of jasmine production on the economic well-being of farmers in Huvina Hadagali, a region known for its high-quality jasmine flowers. The Vijaya Nagara district's Havina Hadagali area is well known throughout the country for its jasmine flower farming. In addition to being referred to as Mallige Nadu, this location is also known as Malligeya Tavaru. The cultivation of the jasmine flower is protected by the Geographical Indication (GI) Tag, and this flower has been popular in this region for a substantial amount of time.</p> <p><strong>Methods:</strong></p> <p>Data was collected from a sample of 364 jasmine growers using a structured questionnaire in Huvina Hadagali, Vijayanagar district. The data focused on different socio-economic factors such as income levels, employment, market access, and agricultural techniques. The study is analysed using IBM SPSS through frequency analysis and 2-step clustering.</p> <p><strong>Result:</strong></p> <p>The results demonstrate that the cultivation of jasmine makes a substantial contribution to the local economy, serving as a main or additional source of income for numerous households. Jasmine farming often contributes 40% of the whole household income, and during peak seasons, it provides significant economic advantages. Nevertheless, the highlighted obstacles were volatile market pricing, pest infestations, and limited access to contemporary farming practices. The study emphasizes the crucial significance of cooperative societies and local marketplaces in stabilising income and offering essential resources and training to farmers.</p> <p><strong>Conclusion:</strong></p> <p>The research highlights the necessity of governmental interventions focused on developing market infrastructure, offering financial assistance, and improving access to agricultural innovations to maintain and augment the economic advantages of jasmine cultivation in Huvina Hadagali.</p>

opencc-zeroJun 2024View details →
zenodo32/100

Sample data for analysis of FFPE sequencing data

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Growth Dynamics and System Models for the Restaurant Industry: Data and Analysis from Taiwanese Chains

<p><span>This dataset includes raw and curated data, system dynamics models, and feedback loop diagrams used in the study of growth dynamics in the restaurant industry, focusing on Taiwanese chains. The data supports the findings presented in the paper "The Growth Dynamics of the Restaurant Industry from Single Store to Chain Store in Taiwan: A Systems Thinking Perspective.&rdquo;</span></p>

opencc-by-4.0Jun 2024View details →
dryad32/100

Data from: Pharmacoeconomic study of anti-influenza virus drugs in Japan based on a network meta-analysis

<p><strong>Objectives:</strong> An analysis was conducted in Japan to determine the most cost-effective neuraminidase inhibitor for the treatment of influenza virus infections from the healthcare payer's standpoint.</p> <p><strong>Methods:</strong> This study reanalyzed the findings of a previous study that had some limitations (no probabilistic sensitivity analysis, quality of life scores measured by the EQ-5D-3L instead of the EQ-5D-5L, and the use of a decision tree model with only three health conditions) by using data from a network meta-analysis study. A decision tree model with eight health conditions was constructed, and costs were identified as medical costs and drug prices (the 2020 version of the Japanese medical fee index). The effectiveness outcomes were measured using EQ-5D-5L questionnaires for adult patients who had previously experienced influenza virus infections. The time horizon was 14 days. Both deterministic and probabilistic sensitivity analyses were performed to examine the robustness of the results.</p> <p><strong>Results:</strong> The base-case cost-effectiveness analysis revealed that oseltamivir outperformed laninamivir, zanamivir, and peramivir, making it the most cost-effective neuraminidase inhibitor. The deterministic and probabilistic sensitivity analyses showed robust results that validated oseltamivir as the most cost-effective among the four neuraminidase inhibitors.</p> <p><strong>Conclusions:</strong> This study thus reconfirmed oseltamivir's position as the most cost-effective neuraminidase inhibitor for the treatment of influenza virus infections in Japan from the standpoint of healthcare payment. These findings can help decision-makers and healthcare providers in Japan, including pharmacists, create and manage formularies.</p>

opencc-zeroJul 2024View details →
zenodo32/100

ARC³N: A Collaborative Uncertainty Catalog to Address the Awareness Problem of Model-Based Confidentiality Analysis - Data Set

<p>Data set of the Paper "ARC&sup3;N: A Collaborative Uncertainty Catalog to Address the Awareness Problem of Model-Based Confidentiality Analysis". For more information, please see the README.md. For even more information please visit https://abunai.dev</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Data set for paper: "Numerical analysis of plastic deformation evolution in polycrystalline copper during cyclic loading with different frequencies"

<p>The dataset contains information necessary for performing the numerical analysis presented in the related paper. The input data are suited for the finite element code Z-set (http://www.zset-software.com/). The resulting data from the numerical calculations and source data for paper figures can be used for further analysis. These data are provided in ASCII format in text files and can be processed by any relevant software.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset for Improved differential expression analysis of miRNA-seq data by modeling competition to be counted

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo32/100

Fig. 2 in Phylogenomic analysis and morphological data suggest left-right swimming behavior evolved prior to the origin of the pelagic Phylliroidae (Gastropoda: Nudibranchia)

Fig. 2 Subtree of Phylliroe and its closest relatives from Dendronotida s.s., Dendronotidae, Scyllaeidae, and Tethyidae. Images from life animals and histological slides. a Dendronotus venustus. b, c Dendronotus frondosus. d–f Crosslandia viridis. g–i Melibe leonina. j–l Phylliroe

opennotspecifiedSep 2020View details →
zenodo32/100

Fig. 1 in Phylogenomic analysis and morphological data suggest left-right swimming behavior evolved prior to the origin of the pelagic Phylliroidae (Gastropoda: Nudibranchia)

Fig. 1 Maximum likelihood phylogeny of Cladobranchia from RAxML-NG using a concatenat- ed nucleotide matrix of 292 genes partitioned by codon position. All nodes with no support values in- dicated have 100% bootstrap support and a posterior probability of 1.0 in our analyses. The blue box indicates Dendronotida sensu stricto and the red outline shows the closest relatives to Phylliroe in our analysis (Melibe leonina, Tethyidae; Scyllaea fulva, Scyllaeidae; Dendronotus venustus, Dendronotidae), highlighted further in Fig. 2

opennotspecifiedSep 2020View details →
zenodo32/100

Data analysis scripts and results for Yu et al. (ANNOgesic)

<p>Supplementary data analysis scripts and results for <em>Yu et al.</em> - &quot;ANNOgesic: A Swiss army knife for the RNA-Seq<br> based annotation of bacterial/archaeal genomes&quot;.</p>

opencc-by-4.0Jan 2018View details →
zenodo32/100

Phased genotype data for FILET's analysis of introgression in D. simulans and D. sechellia

<p>This dataset contains all genotype files prepared for Schrider et al.&#39;s analysis of introgression between D. simulans and D. sechellia. For more information see https://www.biorxiv.org/content/early/2017/09/25/170670</p>

opencc-by-4.0Feb 2018View details →
zenodo32/100

Data set for "Analysis of the Ub to Ub-CR transition in ubiquitin"

<p>Data set for&nbsp;&quot;Analysis of the Ub to Ub-CR transition in ubiquitin&quot; containing the data bases for the wild type and the four mutants analysided, for use with PATHSAMPLE (http://www-wales.ch.cam.ac.uk/PATHSAMPLE.2.1.doc/PATHSAMPLE.html)</p>

opencc-by-4.0May 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record