Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

334

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

334 results for “Python”

Learn how ShareScore rates datasets ↗
zenodo40/100

Reference dataset to check the Python package for the calculation of the OLR-based MJO index (OMI)

<p>The Python package (<a href="https://doi.org/10.5281/zenodo.3613752">10.5281/zenodo.3613752</a>) for the calculation of the OLR-based MJO (OMI) index comes with integration tests, with which the users can check the calculation results.</p> <p>However, in order to run some of these tests, input and reference datasets are needed. Specifically, a dataset of Outgoing Longwave Radiation (OLR) described by Liebmann and Smith (1996) is needed as the input and OMI values described by Kiladis et al. (2014) are needed as the reference.</p> <p>These files are provided by this upload and can be downloaded either as .tar.gz or as .zip file.</p> <p>Whereas exactly the files provided here are needed for the tests, updated versions of these datasets can be found on the websites of NOAA: <a href="https://psl.noaa.gov/data/gridded/data.interp_OLR.html">https://psl.noaa.gov/data/gridded/data.interp_OLR.html</a> and <a href="https://www.esrl.noaa.gov/psd/mjo/mjoindex/">https://www.esrl.noaa.gov/psd/mjo/mjoindex/</a>.</p>

opengpl-3.0-or-laterApr 2020View details →
zenodo40/100

Anonymized monthly StackOverflow activity for the Python tag

<p>CSV files containing metadata for each question, answer and comment posted on StackOverflow during 2019. Each file contains one month of activity with the body text, titles and usernames removed. User and post ids have been obscured so that they are not traceable, but are consistent throughout the dataset.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

The dataset for the study of code change patterns in Python

<p>The dataset of Python projects used for the study of code change patterns and their automation. The dataset lists 120 projects, divided into four domains &mdash; Web, Media, Data, and ML+DL.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

An Empirical Study of Flaky Tests in Python

<p><a href="https://arxiv.org/pdf/2101.09077.pdf">Paper</a></p> <p><a href="https://github.com/se2p/FlaPy">Flapy (our test execution tool)</a></p> <p>TestsOverview.csv<br> - Project Columns: Project_Name, Project_URL, Project_Hash<br> - Test Columns: Test_filename, Test_classname, Test_funcname, Test_parametrization<br> - Flaky_sameOrder_withinIteration: the test is non-order-dependent flaky<br> - Order-dependent: the test is order-dependent flaky<br> - Flaky_Infrastructure: the test shows infrastructure flakiness</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Interaction data for standard tournament for the Axelrod Python project.

<p>This is the data from the tournament described here:&nbsp;http://axelrod-tournament.readthedocs.org/en/latest/standard/strategies.html</p> <p>A description of the format is available here:&nbsp;http://axelrod.readthedocs.org/en/latest/tutorials/further_topics/reading_and_writing_interactions.html</p>

opencc-zeroApr 2016View details →
zenodo40/100

Interaction data for noisy tournament for the Axelrod Python project.

<p>This is the data from the tournament described here:&nbsp;http://axelrod-tournament.readthedocs.org/en/latest/noisy/strategies.html</p> <p>A description of the format is available here:&nbsp;http://axelrod.readthedocs.org/en/latest/tutorials/further_topics/reading_and_writing_interactions.html</p> <p>&nbsp;</p>

opencc-zeroApr 2016View details →
zenodo40/100

Interaction data for probabilistic ending tournament for the Axelrod Python project.

<p>This is the data from the tournament described here:&nbsp;http://axelrod-tournament.readthedocs.org/en/latest/prob_end/ordinary_strategies.html</p> <p>A description of the format is available here:&nbsp;http://axelrod.readthedocs.org/en/latest/tutorials/further_topics/reading_and_writing_interactions.html</p>

opencc-zeroApr 2016View details →
zenodo40/100

Interaction data for the Axelrod Python Project

<p>This contains interaction data for the Axelrod&nbsp;Python project tournaments. Results are available here:&nbsp;http://axelrod-tournament.readthedocs.org/</p>

opencc-zeroApr 2016View details →
zenodo40/100

Madagascar data for deforestprob Python module

<p>This is a data-set for Madagascar to be used with the deforestprob Python module (https://ghislainv.github.io/deforestprob) to compute the spatial probability of deforestation for the year 2010. The data-set includes the following variables:</p> <ul> <li>forest/non-forest for the period 2000-2010 (1)</li> <li>distance to forest edge in 2010, distance to previous deforestation (period 1990-2000) (1)</li> <li>altitude, slope, aspect (2)</li> <li>distance to road, town, river (3)</li> <li>protected area network (4)</li> </ul> <p><strong>Sources</strong></p> <ol> <li>http://bioscenemada.cirad.fr, forest maps derived from Harper et al. 2007 and Hansen et al. 2013</li> <li>http://srtm.csi.cgiar.org, SRTM 90m Digital Elevation Database v4.1</li> <li>http://www.geofabrik.de, data extracts from the OpenStreetMap project for Madagascar,</li> <li>http://rebioma.net, SAPM ("Système des Aires Protégées à Madagascar"), 20/12/2010 version</li> </ol> <p>The variables have been computed using the gdal tools and executing a bash script available at https://github.com/ghislainv/deforestprob/blob/master/notebook/scripts/dataMada.sh</p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

SleepEEGpy: a Python-based software integration package to organize preprocessing, analysis, and visualization of sleep EEG data

<p>This dataset includes three high-density sleep EEG recordings of healthy participants, downsampled to 250 Hz and stored in FIF format:</p> <ol> <li>Nap recording of a young adult participant</li> <li>Overnight recording of a young adult participant</li> <li>Overnight recording of an older adult participant</li> </ol> <p>Additionally, the dataset includes three text files for each recording:</p> <ul> <li>bad_channels.txt: Indexes of noisy channels</li> <li>annotations.txt: Onset and duration of noisy temporal intervals</li> <li>staging.txt: Sleep staging vector</li> </ul> <p>The corresponding package can be found&nbsp;on <a href="https://github.com/NirLab-TAU/sleepeegpy">GitHub.</a></p> <p>For citation, please use:<br>Falach, R., G. Belonosov, J. F. Schmidig, M. Aderka, V. Zhelezniakov, R. Shani-Hershkovich, E. Bar, and Y. Nir. "SleepEEGpy: a Python-based software integration package to organize preprocessing, analysis, and visualization of sleep EEG data." Computers in Biology and Medicine 192 (2025): 110232.<br><a href="https://doi.org/10.1016/j.compbiomed.2025.110232" rel="nofollow">https://doi.org/10.1016/j.compbiomed.2025.110232</a></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

TREXIO files used for the validation tests in the paper entitled 'TurboGenius: Python suite for high-throughput calculations of ab initio quantum Monte Carlo methods'.

<p>The TREXIO files used for the validation tests in the paper entitled TurboGenius: Python suite for high-throughput calculations of ab initio quantum Monte Carlo methods. The detail about the TREXIO library is described in the JCP article [J. Chem. Phys. 158, 174801 (2023)] and the GitHub repository [https://github.com/TREX-CoE/trexio]. The TREXIO files were generated using TREXIO version 2.3.2 (and the corresponding Python API version 1.3.2).</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Artifacts of the CGO 2024 Paper: EasyTracker: A Python Library for Controlling and Inspecting Program Execution

<p>This is the archive of the artifacts for the CGO 2024 Paper <em>"EasyTracker: A Python Library for Controlling and Inspecting Program Execution"</em></p> <p>The EasyTracker library is an open source project, refer to the Home page and Gitlab repository for up-to-date versions:&nbsp;</p> <ul> <li>Home Page: <a href="https://corse.gitlabpages.inria.fr/easytracker">https://corse.gitlabpages.inria.fr/easytracker</a></li> <li>Repository: <a href="https://gitlab.inria.fr/CORSE/easytracker">https://gitlab.inria.fr/CORSE/easytracker</a></li> </ul> <p>The details of the artifacts generation are described in the paper appendix or in the&nbsp;<code>README.md</code> file included in the main artifacts archive <code>easytracker-artifacts-cgo-2024-v1.2.0.tar.gz</code>.</p> <p>Summary of artifacts construction steps (execution in a Docker container):</p> <ul> <li>ensure Docker is installed with&nbsp;<code>docker --version</code>,</li> <li>download the artifacts archive <code>easytracker-artifacts-cgo-2024-v1.2.0.tar.gz</code>,</li> <li>extract with <code>tar xvzf easytracker-artifacts-cgo-2024-v1.2.0.tar.gz</code>,</li> <li>change dir with <code>cd eastracker-artifacts-cgo-2024</code>,</li> <li>download the EasyTracker sources archive&nbsp; <code>easytracker-archive-dfe8aa888f.tar.gz</code><a href="../api/records/10428215/draft/files/easytracker-archive-dfe8aa888f.tar.gz/content" target="_blank" rel="noopener noreferrer">,</a></li> <li>extract with&nbsp;<code>tar xvzf easytracker-archive-dfe8aa888f.tar.gz</code>,</li> <li>download the docker image <code>docker-image-easytracker-1.2.0.tar</code>,</li> <li>load the Docker image with <code>docker load -i&nbsp;docker-image-easytracker-1.2.0.tar</code>,&nbsp;</li> <li>generate all artifacts with&nbsp;<code>./in-docker.sh ./run-all.sh</code>,</li> <li>all artifacts are generated in <code>figure-*/</code> directories,</li> <li>refer to the artifacts archive <code>README.md</code> file for more details, or refer to the paper artifacts appendix.</li> </ul> <p>Note that this artifacts archive is extracted from the artifacts repository at tag <code>v1.2.0</code>: &nbsp;<a title="Opens in new tab" href="https://gitlab.inria.fr/CORSE/easytracker-artifacts-cgo-2024/-/tree/v1.2.0" target="_blank" rel="noopener">https://gitlab.inria.fr/CORSE/easytracker-artifacts-cgo-2024/-/tree/v1.2.0&nbsp;</a></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Supplementary files for the dingo Python library

<p>We&nbsp;compare the&nbsp;<a href="https://drops.dagstuhl.de/opus/frontdoor.php?source_opus=13820">Multiphase Monte Carlo&nbsp;flux Sampling (MMCS)</a> feature&nbsp;of the <a href="https://github.com/GeomScale/dingo">dingo</a> library&nbsp;against the combined method of <a href="https://gitlab.com/csb.ethz/PolyRound">PolyRound</a> (for rounding) followed by <a href="https://modsim.github.io/hopsy/">hopsy</a> (for sampling) on a set of 7 models with a ranging&nbsp;dimension&nbsp;(<em>ext_data.zip</em>).</p> <p>The&nbsp;<em>simpl_transf_polytopes.zip</em> contains the polytopes retrieved after the <em>simplify()</em> and <em>transform() </em>functions of the PolyRound library. <br>These polytopes were used as input for the dingo implementation of the MMCS algorithm asking for an ESS of 1000. <br>Under the&nbsp;<em>dingo_samples_on_simpl_transf_polytopes.zip</em>&nbsp;the resulting&nbsp;samples from <em>dingo</em> can be found using the MMCS algorithm.&nbsp;</p> <p>Similarly, <em>polyrounded_polytopes.zip&nbsp;contains&nbsp;</em>the polytopes retrieved after&nbsp;applying <em>simplify()</em>, <em>transform()</em> and&nbsp;<em>round()</em> functions of the PolyRound library. <br>These polytopes were used as input for the <a href="https://modsim.github.io/hopsy/">hopsy</a> library, again, asking for&nbsp;an ESS of 1000. <br>The <em>hopsy_samples.zip</em>&nbsp;folder contains the resulting&nbsp;samples from hopsy<em> </em>library, using&nbsp;a thinning of 100<em>d </em>;<em> </em>only in the case of Recon3D a thinning of 200<em>d </em>was used as suggested by the authors. Under the&nbsp;<em>hopsy_samples_ess_1000.zip&nbsp;</em>folder, we provide the&nbsp;<em>hopsy</em> samples with an ESS of 1000. <br>Last,&nbsp;the samples produced using the efficient Billiard Walk implementation of <em>dingo</em> can be found under the <em>BWRsamples.zip </em>file. <br>In this case, 20000 points were sampled for each model.</p> <p>Further, the <em>sars_samples.zip&nbsp;</em>file contains&nbsp;<em>dingo</em>&nbsp;samples from the solution space of the SARS-CoV-2&nbsp;integrated&nbsp;model of <a href="https://doi.org/10.1093/bioinformatics/btaa813">Renz et <em>al</em> (2020)</a>&nbsp;for the following cases:&nbsp;</p> <ul> <li>unbiased; where the zero vector has been used as the objective function of the model</li> <li>after maximising for the human biomass&nbsp;</li> <li>after maximising for the&nbsp;virus biomass objective function (VBOF)</li> </ul> <p>The following Python scripts to perform these experiments are included:</p> <ul> <li><em>polyround_preproces.py&nbsp;</em>: runs the <em>PolyRound&nbsp;</em>functions and builds the simplified and transformed polytopes that&nbsp;<em>dingo&nbsp;</em>will use as well as the simplified, transformed and rounded polytopes&nbsp;<em>hopsy</em>&nbsp;uses</li> <li><em>hopsy_on_polyrounded_polytopes.py&nbsp;</em>: performs&nbsp;sampling with <em>hopsy&nbsp;</em></li> <li><em>dingo_on_simpl_transf_polytopes.py&nbsp;</em>: performs sampling with&nbsp;<em>dingo&nbsp;</em></li> <li><em>run_bwr_exp.py</em> computes samples using the efficient Billiard Walk of <em>dingo</em></li> <li><em>binary_search.py</em>&nbsp;: a function to return the index in the chain where ESS becomes 1000</li> <li><em>compute_ess.py: </em>based on a model's&nbsp;<em>hopsy</em> samples (under the&nbsp;<em>hopsy_samples.zip </em>folder)&nbsp;it retrieves the samples&nbsp;with an ESS of 1000 and the corresponding&nbsp;required time for <em>hopsy&nbsp;</em>to build them. The script requires the total time of the&nbsp;<em>hopsy&nbsp;</em>experiment&nbsp;recorded in the model's corresponding <em>.txt&nbsp;</em>file (you can find this under the&nbsp;<em>hopsy_samples.zip)</em></li> <li><em>compute_ess_psrf_per_phase.py&nbsp;</em>computes ESS and PSRF in specific indices which correspond to those when MMCS switches from a phase a to next one</li> </ul> <p>A <a href="https://github.com/hariszaf/dingo/blob/vbof/tutorials/vbof.ipynb">notebook</a> is available for how the integrated model was sampled.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Replication package for the paper: "A Study on the Pythonic Functional Constructs' Understandability"

<h1>Replication Package for "<em>A Study on the Pythonic Functional Constructs' Understandability</em>" to appear at ICSE 2024</h1> <ul> <li><strong>Authors</strong>: Cyrine Zid, Fiorella Zampetti, Giuliano Antoniol, Massimiliano Di penta</li> <li><strong>Article Preprint:</strong> <a href="https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf">https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf</a></li> <li><strong>Artifacts:</strong> <a href="https://doi.org/10.5281/zenodo.8191782">https://doi.org/10.5281/zenodo.8191782</a></li> <li><strong>License</strong>: <a>GPL V3.0</a></li> </ul> <p>This package contains folders and files with code and data used in the study described in the paper. In the following, we first provide all fields required for the submission, and then report a detailed description of all repository folders.</p> <h2>Artifact Description</h2> <h3>Purpose</h3> <p>The artifact is about a controlled experiment aimed at investigating the extent to which Pythonic functional constructs have an impact on source code understandability. The artifact archive contains:</p> <ol> <li>The material to allow replicating the study (see Section <a>Experimental-Material</a>)</li> <li>Raw quantitative results, working datasets, and scripts to replicate the statistical analyses reported in the paper. Specifically, the executable part of the replication package reproduces figures and tables of the quantitative analysis (RQ1 and RQ2) of the paper starting from the working datasets.</li> <li>Spreadsheets used for the qualitative analysis (RQ3).</li> </ol> <p>We apply for the following badges:</p> <ul> <li><strong>Available and reusable</strong>: because we provide all the material that can be used to replicate the experiment, but also to perform the statistical analyses and the qualitative analyses (spreadsheets, in this case)</li> </ul> <h3>Provenance</h3> <ul> <li><strong>Paper preprint link:</strong> <a href="https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf">https://mdipenta.github.io/files/ICSE24_funcExperiment.pdf</a></li> <li><strong>Artifacts:</strong> <a href="https://doi.org/10.5281/zenodo.8191782">https://doi.org/10.5281/zenodo.8191782</a></li> </ul> <h3>Data</h3> <p>Results have been obtained by conducting the controlled experiment involving <a href="https://www.prolific.com/">Prolific</a>workers as participants. Data collection and processing followed a protocol approved by the University ethical board. Note that all data enclosed in the artifact is completely anonymized and does not contain sensible information.</p> <p>Further details about the provided dataset can be found in the Section <a>Results' directory and files</a></p> <h3>Setup and Usage (for executable artifacts):</h3> <p>See the Section <a>Scripts to reproduce the results, and instructions for running them</a></p> <h2><a>Experiment-Material/</a></h2> <p>Contains the material used for the experiment, and, specifically, the following subdirectories:</p> <h3><a>Google-Forms/</a></h3> <p>Contains (as PDF documents) the questionnaires submitted to the ten experimental groups.</p> <h3><a>Task-Sources/</a></h3> <p>Contains, for each experimental group (G-1...G-10), the sources used to produce the Google Forms, and, specifically: - The cover letter (Letter.docx). - A directory for each experimental task (Lambda 1, Lambda 2, Comp 1, Comp 2, MRF 1, MRF 2, Lambda Comparison, Comp Comparison, MRF Comparison). Each directory contains: (i) the exercise text (in both Word and .txt format), the source code snippet, and its .png image to be used in the form. <strong>Note:</strong> the "Comparison" tasks do not have any exercise as the purpose is always the same, i.e., to compare the (perceived) understandability of the snippets and return the results of the comparison.</p> <h3><a>Code-Examples-Table1/</a></h3> <p>Contains the source code snippets used as objects of the study (the same you can find under "Task-Sources/"), named as reported in Table 1.</p> <h2>Results' directory and files</h2> <p>&nbsp;</p> <h3><a>raw-responses/</a></h3> <p>Contains, as spreadsheets, the raw responses provided by the study participants through Google forms.</p> <h3><a>raw-results-RQ1/</a></h3> <p>Contains the raw results for RQ1. Specifically, the directory contains a subdirectory for each group (G1-G10). Each subdirectory contains: - For each user (named using their Prolific IDs, a directory containing, for each question (Q1-Q6) the produced python code (Qn.py) its output (QnR.txt) and its StdErr output (QnErr.txt). - "expected-outputs/": A directory containing the expected outputs for each task (Qn.txt).</p> <h4><a>working-results/RQ1-RQ2-files-for-statistical-analysis/</a></h4> <p>Contains three .csv files used as input for conducting the statistical analysis and drawing the graphs for addressing the first two research questions of the study. Specifically:</p> <ul> <li> <p><a>ConstructUsage.csv</a> contains the declared frequency usage of the three functional constructs object of the study. This file is used to draw Figure 4. The file contains an entry for each participant, reporting the (text-coded) frequency of construct usage for Comprehension, Lambda, and MRF.</p> </li> <li> <p><a>RQ1.csv</a> contains the collected data used for the mixed-effect logistic regression relating the use of functional constructs with the correctness of the change task, as well as the logistic regression relating the use of map/reduce/filter functions with the correctness of the change task. The csv file contains an entry for each answer provided by each subject, and features the following columns:</p> <ul> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>User</em>: user ID</li> <li><em>Time</em>: task time in seconds</li> <li><em>Approvals</em>: number of approvals on previous tasks performed on Prolific</li> <li><em>Student</em>: whether the participant declared themselves as a student</li> <li><em>Section</em>: section of the questionnaire (lambda, comp, or mrf)</li> <li><em>Construct</em>: specific construct being presented (same as "Section" for lambda and comp, for mrf it says whether it is a map, reduce, or filter)</li> <li><em>Question</em>: question id, from Q1 to Q6, indicate the ordering of the question</li> <li><em>MainFactor</em>: main factor treatment for the given question - "f" for functional, "p" for procedural counterpart</li> <li><em>Outcome</em>: TRUE if the task was correctly performed, FALSE otherwise</li> <li><em>Complexity</em>: cyclomatic complexity of the construct (empty for mrf)</li> <li><em>UsageFrequency</em>: usage frequency of the given construct</li> </ul> </li> <li> <p><a>RQ1Paired-RQ2.csv</a> contains the collected data used for the ordinal logistic regression of the relationship between the perceived ease of understanding of the functional constructs and (i) participants' usage frequency, and (ii) constructs' complexity (except for map/reduce/filter). The file features a row for each participant, and the columns are the following:</p> <ul> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>User</em>: user ID</li> <li><em>Time</em>: task time in seconds</li> <li><em>Approvals</em>: number of approvals on previous tasks performed on Prolific</li> <li><em>Student</em>: whether the participant declared themselves as a student</li> <li><em>LambdaF</em>: result for the change task related to a lambda construct</li> <li><em>LambdaP</em>: result for the change task related to the procedural counterpart of a lambda construct</li> <li><em>CompF</em>: result for the change task related to a comprehension construct</li> <li><em>CompP</em>: result for the change task related to the procedural counterpart of a comprehension construct</li> <li><em>MrfF</em>: result for the change task related to an MRF construct</li> <li><em>MrfP</em>: result for the change task related to the procedural counterpart of a MRF construct</li> <li><em>LambdaComp</em>: perceived understandability level for the comparison task (RQ2) between a lambda and its procedural counterpart</li> <li><em>CompComp</em>: perceived understandability level for the comparison task (RQ2) between a comprehension and its procedural counterpart</li> <li><em>MrfComp</em>: perceived understandability level for the comparison task (RQ2) between a MRF and its procedural counterpart</li> <li><em>LambdaCompCplx</em>: cyclomatic complexity of the lambda construct involved in the comparison task (RQ2)</li> <li><em>CompCompCplx</em>: cyclomatic complexity of the comprehension construct involved in the comparison task (RQ2)</li> <li><em>MrfCompType</em>: type of MRF construct (map, reduce, or filter) used in the comparison task (RQ2)</li> <li><em>LambdaUsageFrequency</em>: self-declared usage frequency on lambda constructs</li> <li><em>CompUsageFrequency</em>: self-declared usage frequency on comprehension constructs</li> <li><em>MrfUsageFrequency</em>: self-declared usage frequency on MRF constructs</li> <li><em>LambdaComparisonAssessment</em>: outcome of the manual assessment of the answer to the "check question" required for the lambda comparison ("yes" means valid, "no" means wrong, "moderate<em>chatgpt" and "extreme</em>chatgpt" are the results of GPTZero)</li> <li><em>CompComparisonAssessment</em>: as above, but for comprehension</li> <li><em>MrfComparisonAssessment</em>: as above, but for MRF</li> </ul> </li> </ul> <h3><a>working-results/inter-rater-RQ3-files/</a></h3> <p>This directory contains four .csv files used as input for computing the inter-rater agreement for the manual labeling used for addressing RQ3. Specifically, you will find one file for each functional construct, i.e., comprehension.csv, lambda.csv, and mrf.csv, and a different file used for highlighting the reasons why participants prefer to use the procedural paradigm, i.e., procedural.csv.</p> <h3><a>working-results/RQ2ManualValidation.csv</a></h3> <p>This file contains the results of the manual validation being done to sanitize the answers provided by our participants used for addressing RQ2. Specifically, we coded the behaviour description using four different levels: (i) correct ("yes"), (ii) somewhat correct ("partial"), (iii) wrong ("no"), and (iv) automatically generated. The file features a row for each participant, and the columns are the following:</p> <ul> <li><em>ID</em>: ID we used to refer the participant in the paper's qualitative analysis</li> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>ProlificID</em>: user ID</li> <li><em>Comparison for lambda construct description</em>: answer provided by the user for the lambda comparison task</li> <li><em>Final Classification</em>: our assessment of the lambda comparison answer</li> <li><em>Comparison for comprehension description</em>: answer provided by the user for the comprehension comparison task</li> <li><em>Final Classification</em>: our assessment of the comprehension comparison answer</li> <li><em>Comparison for MRF description</em>: answer provided by the user for the MRF comparison task</li> <li><em>Final Classification</em>: our assessment of the MRF comparison answer</li> </ul> <h3><a>working-results/RQ3ManualValidation.xlsx</a></h3> <p>This file contains the results of the open coding applied to address our third research question. Specifically, you will find four sheets, one for each functional construct and one for the procedural paradigm. Each sheet reports the provided answers together with the categories assigned to them. Each sheet contains the following columns:</p> <ul> <li><em>ID</em>: ID we used to refer the participant in the paper's qualitative analysis</li> <li><em>Group</em>: experimental group to which the participant is assigned</li> <li><em>ProlificID</em>: user ID (as in the tables from the quantitative analysis)</li> <li>: question asked to the user</li> <li><em>Final Classification</em>: The outcome of our categorization according to the taxonomy shown in Table 10.</li> </ul> <h2>Scripts to reproduce the results and instructions for running them</h2> <p>&nbsp;</p> <h3><a>FuncConstructs-Statistics.r</a></h3> <p>This file contains an R script that you can reuse to re-run all the analyses conducted and discussed in the paper.</p> <h3><a>FuncConstructs-Statistics.ipynb</a></h3> <p>This file contains the code to re-execute all the analysis conducted in the paper as a Jupyter Notebook (using the R Kernel).</p> <h3><a>run-analysis.sh</a></h3> <p>This script can be used to run the R script <code>FuncConstructs-Statistics.r</code> using a Docker container (see Option 1 below) in Unix operating systems.</p> <h3><a>run-analysis.bat</a></h3> <p>This script can be used to run the R script <code>FuncConstructs-Statistics.r</code> using a Docker container (see Option 1 below) in Windows operating systems (power shell recommended).</p> <h3><a>run-jupyter-container.sh</a></h3> <p>This script can be used to run a local Jupyter server (with R kernel and all required packages) from a Docker container (see Option 3 below) in Unix operating system.</p> <h3><a>run-jupyter-container.bat</a></h3> <p>This script can be used to run a local Jupyter server (with R kernel and all required packages) from a Docker container (see Option 3 below) in Windows operating systems (power shell recommended).</p> <h3>How to Run the scripts</h3> <p>There are four options to run the scripts. In all cases, one has first to open a shell terminal window (e.g., bash or sh in Unixes) in the replication package directory. For Windows, we suggest to use a Power Shell.</p> <ol> <li> <p><strong>Running the R script using Dockerized R installation</strong>: this is the simplest option, and it simply requires a running Docker engine. In <strong>Unix (MacOS, Linux)</strong>, to produce the results, one has to run the shell script "run-analysis.sh" (e.g., by typing <code>sh run-analysis.sh</code> or simply <code>./run-analysis.sh</code> after making it executable). In <strong>Windows</strong>, one has to run the script "run-analysis.bat" instead (by typing <code>.\run-analysis.bat</code>). This script (either .sh or .bat) will:</p> <ul> <li>Pull a docker image named <code>mdipenta/rexp</code> which contains an R installation with all required packages.</li> <li> <p>Run R from the container created from the image and produce the paper's results under a directory named <code>results/</code>.</p> </li> <li> <p><strong>Note:</strong> An alternative would be to run everything from inside the container, after running it in interactive mode. To this aim, please execute the following commands:</p> <ol> <li>In <strong>Unix</strong>: <code>docker run -v${PWD}:/data --rm -ti --name shell mdipenta/rexp:latest bash</code>in <strong>Windows</strong>: <code>docker run -v %cd%:/data --rm -ti --name shell mdipenta/rexp:latest bash</code></li> <li><code>cd data</code></li> <li><code>R --no-save &lt; FuncConstructs-Statistics.r</code> After exiting the container, the "results" directory will be again populated with the study results.</li> </ol> </li> </ul> </li> <li> <p><strong>Running the R script from own R installation</strong>: this option works if one has an R installation already (or wants to use an R installation) without relying on the Docker image. The steps to be followed are:</p> <ul> <li>Uncomment the <code>install.packages(..)</code> instruction in the first lines of the script. This will allow for the installation of the required packages.</li> <li>Just run, from the current directory, the script <code>FuncConstructs-Statistics.r</code>, using the command <code>Rscript FuncConstructs-Statistics.r</code> (making sure the directory containing Rscript is in your PATH, this should work fine in Unixes, it might require to modify the PATH environment variable in Windows). Should you experience problems with the first part of the script (installations), try to execute the <code>install.packages(..)</code> statement from your R GUI, and then run the script again.</li> </ul> </li> <li> <p><strong>Using the Jupyter Notebook using a Dockerized Jupyter lab with R kernel</strong>: this option allows for opening the Jupyter Notebook with all results without having to install Jupyter with the R kernel, nor all the required R packages. The steps required are:</p> <ul> <li>Run the <code>./run-jupyter-container.sh</code> (<strong>Unix</strong>) or <code>.\run-jupyter-container.bat</code> (<strong>Windows</strong>). It will download the <code>mdipenta/myjupyter</code> image and run Jupyter lab from it.</li> <li>Open a browser on <a href="http://localhost:8888/">localhost:8888</a> (or if it does not work, <a href="http://127.0.0.1:8888/">127.0.0.1:8888)</a> and, when being asked for a password, type <code>docker</code>.</li> <li>From the Jupyter lab page, open the "FuncConstruct-Statistics.ipynb" notebook, and (if you wish) re-run it, or simply browse its results. Note: differently from options 1 and 2, results are not saved, but just displayed in the notebook.</li> </ul> </li> <li> <p><strong>Using the Jupyter Notebook from your installation</strong>: this is similar to Option 3, but it can work if you have already Jupyter lab installed, with the R kernel enabled (for details see: <a href="https://github.com/IRkernel/IRkernel">https://github.com/IRkernel/IRkernel</a>). The steps to follow are:</p> <ul> <li>Run jupyter lab (e.g., jupyter lab from the command line) and open it on a webpage.</li> <li>Open the <code>FuncConstruct-Statistics.ipynb</code> notebook.</li> <li>If you want to re-execute it, uncomment the <code>install.packages()</code> line.</li> <li>Re-run it (if you wish).</li> </ul> </li> </ol> <h3>The output</h3> <p>If using Option 1 or 2, the <code>results</code> directory will contain the following files:</p> <ul> <li><strong>Figures 4 and 5</strong> as in the paper.</li> <li><strong>Tables 2-9</strong> as in the paper in various formats (csv, tex, and for Tables 3-5 also .txt). Some notes: The diagnostics (top part, up to "Fixed effects") for Tables 3-5 are shown in the .txt files only. However, these files do not report the "OR" columns that correspond to exp(Estimate). This is because the .txt file contains the statistics dump which does not include the ORs. The .csv and .tex tables report the Fixed effects as shown in the paper, including the ORs.</li> <li><strong>rq1-rq2-correlation</strong> (.tex and .csv) contains the correlation analysis between RQ1 and RQ2 results as discussed in the "Threats to construct validity" (Section 6).</li> <li><strong>rq3-inter-rater</strong> (.tex and .csv) contains the results of the inter-rater agreements analysis discussed in Section 3.6.</li> </ul>

opengpl-3.0-or-laterDec 2023View details →
zenodo40/100

Dataset supporting the tool 'delfies: a Python package for the detection of DNA breakpoints with neo-telomere addition'

<h2>Purpose</h2> <p><br>These data can be used to test my tool&nbsp;<a href="https://github.com/bricoletc/delfies">delfies </a>on real data, to get a concrete sense of its inputs/outputs and test that it is&nbsp;<br>properly installed.</p> <h2>Description</h2> <h3>Genome</h3> <p>I downloaded the genome of&nbsp;<em>Oscheius onirici</em>, accession: <a href="https://www.ebi.ac.uk/ena/browser/view/GCA_932521025.1">GCA_932521025</a>.</p> <p>I subsampled the genome to the last 2kbp of chromosome I, which contains an elimination breakpoint,&nbsp;<br>using `seqkit` v2.8.2, giving the FASTA file in this release.</p> <h3>Sequencing data</h3> <p>I then downloaded the following sequencing data for *O. onirici*, from the European Nucleotide Archive:</p> <ul> <li>ERR5967937: Illumina NovaSeq 6000 paired end short reads. Reads are 2x150bp with average per-base quality of Q27.</li> <li>ERR10796202: Oxford Nanopore PromethION long reads. Reads have average length 11.9kbp and average per-base quality Q11.4.</li> <li>ERR7979900: Pacific Biosciences (PacBio) Sequel II long reads. Reads have average length 11.1kbp and average per-base quality Q28.<br><br></li> </ul> <p>And aligned them to the above genome with `minimap2` version 2.26-r1175, using the following presets:&nbsp;<br>"map-ont" for the Nanopore data, "map-hifi" for the PacBio data, "sr" for the Illumina data.</p> <p>After sorting with `samtools`, this gives the BAM files in this release.</p> <h3>Running delfies</h3> <p>I then ran `delfies` version 0.6.0 on each BAM and genome, as:</p> <p>```sh<br>delfies --threads 16 \<br>&nbsp; &nbsp; --telo_forward_seq TTAGGC \<br>&nbsp; &nbsp; --breakpoint_type all \<br>&nbsp; &nbsp; --min_mapq 20 \<br>&nbsp; &nbsp; --min_supporting_reads 6 \<br>&nbsp; &nbsp; \${genome} \${bam} \${odirname}<br>```</p> <p>The three resulting output directories are in this release, prefixed with `delfies_`.</p> <p><strong>A single, identical breakpoint is found using all three BAMs</strong> (see files '*breakpoint_locations.bed').</p> <h3>Data source</h3> <p>The above raw data were produced and released by the Wellcome Sanger Institute as part of projects&nbsp;<br><a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB51305">PRJEB51305</a> and <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB59023">PRJEB59023.</a></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Dataset - A Python-Based Approach to Sputter Deposition Simulations in Combinatorial Materials Science

<p>This dataset accompanies the publication <em>"A Python-Based Approach to Sputter Deposition Simulations in Combinatorial Materials Science,"</em> which presents and validates pySIMTRA, a Python wrapper for the Monte Carlo-based SIMTRA simulation tool. The dataset includes all measured and simulated data shown in the publication, as well as additional animations visualizing the compositions in the multinary composition space.</p> <p>The dataset contains the compositions for each of the seven materials libraries (in at.%) in the quaternary Ni-Pd-Pt-Ru system. Additionally, it provides the simulated number of particles as outputted by SIMTRA, which serve as the basis for composition estimation. Both the compositional data and the particle counts are supplied in .csv format. To supplement the results, 3D animations of the quaternary compositional spaces are included, showing the comparison between simulated compositions (red dots) and measured compositions (blue dots). These animations offer a more intuitive visualization of the data compared to the static Figures in the publication and are supplied as .gif files.</p> <p>Due to the in-depth analysis of cathode tilt discussed in the paper, the dataset also includes simulation results for the ternary Pd-Pt-Ru library, highlighting the effect of varying the cathode tilt angle. Simulations were conducted for tilt angles of 10&deg;, 9.5&deg;, 9&deg;, and 8.5&deg;.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

CrossDomainTypes4Py: A Python Dataset for Cross-Domain Evaluation of Type Inference Systems

<p>This dataset contains python repositories mined on GitHub on January 20, 2021. It allows a cross-domain evaluation of type inference systems. For this purpose, it consists of two sub-datasets, each containing only projects from the web or scientific calculation domain, respectively. Therefore&nbsp;we searched for projects with dependencies to either <a href="https://numpy.org/">NumPy</a>&nbsp;or <a href="https://flask.palletsprojects.com/en/2.0.x/">Flask</a>. Furthermore, only projects with dependencies to <a href="http://mypy-lang.org/">mypy</a>&nbsp;were considered, because this should ensure that at least parts of the projects have type annotations. These can be used later as ground truth. Further details about the dataset will be described in an upcoming paper, as soon as it is published it will be linked here.<br> The dataset consists of two files for the two sub-datasets. The web domain dataset contains 3129 repositories and the scientific calculation domain dataset contains 4783 repositories. The files have&nbsp;two columns with the URL to the GitHub repository and the used&nbsp;commit hash. Thus, it is possible to download the dataset using shell or python&nbsp;scripts, for example, the pipeline provided by <a href="https://github.com/saltudelft/many-types-4-py-dataset">ManyTypes4Py</a>&nbsp;can be used.<br> If repositories do not exist anymore or are private, you can contact us via the following email address: bernd.gruner@dlr.de. We have a backup of all repositories and will be happy to help you.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

"Python for Data Science" (AY250; UC Berkeley) Data files

<p>Data files for &quot;Python for Data Science&quot; (AY250; UC Berkeley)</p> <ul> <li><a href="https://zenodo.org/api/files/796932a4-1a76-4467-a66e-ac0d47e029c7/homework1_data.tgz">homework1_data.tgz </a>- Data for HW1</li> </ul> <p>Course website:&nbsp;https://github.com/profjsb/python-seminar</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

E3SM simulation results and associated python analysis scripts

<p>This archive contains E3SM Land Model simulation results associated with the <em>Journal of Advances in Modeling Earth Systems&nbsp;(JAMES)</em><em>&nbsp;</em>article&nbsp;titled &quot;More Realistic Intermediate Depth Dry Firn Densification in the Energy&nbsp;Exascale Earth System Model (E3SM),&quot; by Adam M. Schneider, Charles&nbsp;S. Zender, and Stephen F.&nbsp;Price.&nbsp; Also included in the archive are python scripts used to analyze associated data and&nbsp;an offline, statistical firn model.</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

spectrapepper: A Python toolbox for advanced analysis of spectroscopic data for materials and devices.

<p>spectrapepper&nbsp;is a Python package that makes advanced analysis of spectroscopic data easy and accessible through straightforward, simple, and intuitive code. This library contains functions for every stage of spectroscopic methodologies, including data acquisition, pre-processing, processing, and analysis. In particular, advanced and high statistic methods are intended to facilitate, namely combinatorial analysis and machine learning, allowing also fast and automated traditional methods. The following is a short list of some main procedures that&nbsp;spectrapepper&nbsp;package enables: i) Baseline removal functions, ii) Normalization methods, iii) Noise filters, trimming tools, and despiking methods, iv) Chemometric algorithms to find peaks, fit curves, and deconvolution of spectra, v) Combinatorial analysis tools, such as Spearman, Pearson, and n-dimensional correlation coefficients, vi) Tools for Machine Learning applications, such as data merging, randomization, and decision boundaries, and vii) Sample data and examples</p>

opencc-by-4.0Nov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record