Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
334
datasets available to search
ShareScore release 0.7.1
Dataset results
334 results for “python”
Interaction data for the Axelrod Python Project (v1.14.0)
<p>This contains interaction data for the Axelrod Python project tournaments. Results are available here: http://axelrod-tournament.readthedocs.org/</p>
Interaction data for the Axelrod Python Project (v1.15.0)
<p>This contains interaction data for the Axelrod Python project tournaments. Results are available here: http://axelrod-tournament.readthedocs.org/</p>
Interaction data for the Axelrod Python Project (v1.16.0)
<p>This contains interaction data for the Axelrod Python project tournaments. Results are available here: http://axelrod-tournament.readthedocs.org/</p>
Interaction data for the Axelrod Python Project (v1.19.0)
<p>This contains interaction data for the Axelrod Python project tournaments. Results are available here: http://axelrod-tournament.readthedocs.org/</p>
Interaction data for the Axelrod Python Project (v2.3.0)
<p>This contains interaction data for the Axelrod Python project tournaments. Results are available here: http://axelrod-tournament.readthedocs.org/</p>
Interaction data for the Axelrod Python Project (v2.4.0)
<p>This contains interaction data for the Axelrod Python project tournaments. Results are available here: http://axelrod-tournament.readthedocs.org/</p>
Python Systems for Empirical Analysis
<p>Reference</p> <p>Studies who have been using the data (in any form) are required to include the following reference:</p> <p>@inproceedings{Orru2015, abstract = {The aim of this paper is to present a dataset of metrics associated to the first release of a curated collection of Python software systems. We describe the dataset along with the adopted criteria and the issues we faced while building such corpus. This dataset can enhance the reliability of empirical studies, enabling their reproducibility, reducing their cost, and it can foster further research on Python software.}, author = {Orrú, Matteo and Tempero, Ewan and Marchesi, Michele and Tonelli, Roberto and Destefanis, Giuseppe}, booktitle = {Submitted to PROMISE '15}, keywords = {Python, Empirical Studies, Curated Code Collection}, title = {A Curated Benchmark Collection of Python Systems for Empirical Studies on Software Engineering}, year = {2015} }</p> <p>About the Data</p> <p>Overview</p> <p>This paper presents a dataset of metrics taken from a curated collection of 51 popular Python software systems.</p> <p>The dataset reports 41 metrics of different categories: volume/size, complexity and object oriented metrics. These metrics and computed both at file and class level. We provide metrics for any file and class of each system and global metrics (computed on the entire system). Moreover we provide 14 meta-data for each system.</p> <p>Paper Abstract</p> <p>The aim of this paper is to present a dataset of metrics associated to the first release of a curated collection of Python software systems. We describe the dataset along with the adopted criteria and the issues we faced while building such corpus. This dataset can enhance the reliability of empirical studies, enabling their reproducibility, reducing their cost, and it can foster further research on Python software.</p>
Replication package for the paper :The Relationship Between Different Python Argument-Passing Mechanisms and Fixes: An Empirical Study
<p><strong>Abstract:</strong></p> <p>Modern programming languages, such as Python, have introduced a variety of constructs and syntactical elements to make software development more efficient and concise. Examples include lambda functions, comprehension collections, or mechanisms to facilitate the passing of arguments to a function. While many of such constructs may, in principle, be beneficial for developers, recent studies have shown that certain programming constructs may affect program understanding and even induce more fixes than other changes. <br>This paper studies the effect of different Python argument-passing mechanisms to investigate their relationship with code proneness to be fixed. Specifically, we study the fix-proneness for what concerns function definitions and invocations. This is done by analyzing the evolutionary history of 200 Python projects, for a total of about 3M functions and 12M call sites. While there are varying effects for what concerns parameter declaration mechanisms, we found evidence that keyword-based argument passing is less defect-prone than positional argument passing, and this is not affected by size-related confounding factors.</p>
APIBench: A Benchmark Dataset for Evaluating API Recommendation Approaches in Python and Java
<p>APIBench is the benchmark dataset APIBench released in the paper "<a href="https://yunpeng.site/files/apirec.pdf"><em>Revisiting, Benchmarking and Exploring APIRecommendation: How Far Are We?</em></a>". </p> <p>APIBench contains two sub-dataset for evaluating the performance of query-based and code-based API recommendation approaches, namely APIBench-Q and APIBench-C. Each sub-dataset has a Java version and a Python version.</p> <p>APIBench-Q contains 4,309 Python queries and 6,563 Java queries collected from Stack Overflow posts generated from Aug 2008 to Feb 2021 and tutorial websites Geeks4Geeks, Java2s, and Kode Java in April 2021.</p> <p>APIBench-C contains 2,361 Python projects and 1,477 Java projects mined from GitHub in April 2021.</p> <p>Please read the <strong>README.md</strong> file for detailed information about the benchmark.</p> <p>The evaluation results of existing API recommendation approaches can be found in <a href="https://github.com/JohnnyPeng18/APIBench">this GitHub repository</a>.</p>
Automated analysis of HPLC chromatograms obtained during the dehydration of N-Acetylglucosamine into 3-Acetamido-5-acetylfurane using Python
<p><strong>Content: </strong>This dataset contains High Performance Liquid Chromatography (HPLC) chromatograms, obtained while studying the dehydration of N-Acetylglucosamine into 3-Acetamido-5-acetylfurane in DMAc, DMF and NMP employing various catalysts and screening different reaction conditions. Reaction conditions can be retrieved from the metadata attached. Importantly, a python script to automatically extract, plot and analyze the HPLC-data obtained is provided.</p><p><strong>Acknowledgements</strong>: The authors acknowledge support by the German Research Foundation (DFG) within NFDI4Cat (ID 441926934). Parts of this work were funded by the Cluster of Excellence Fuel Science Center (EXC 2186, ID: 390919832) funded by the Excellence Initiative by the German federal and state governments. Furthermore, the authors thank Jens Heller and Frederic Thilmany for performing the HPLC measurements.</p><p> </p>
Enquête par questionnaire humanités "Manipuler des données en Sciences Humaines et Sociales (SHS) : R, Python, ou autre ?"
<p>Ce sondage réalisé avec Framaform https://framaforms.org/manipuler-des-donnees-en-sciences-humaines-et-sociales-shs-r-python-ou-autre-1675889669 était destiné à tous les personnels impliqués dans la recherche et / ou l'enseignement en sciences humaines et sociales mobilisant du traitement de données (humanités numériques, sciences sociales computationnelles, etc.). Il a été diffusé sur la liste de diffusions DH, sur les sites de l'Observatoire des Humanités numériques de l'ENS PSL et de l'INSHS du CNRS.</p><p>L'enquête visait à mieux connaître les usages de la programmation chez les chercheurs, enseignants-chercheurs, étudiants et personnels de soutien à la recherche. Les résultats obtenus permettent de proposer un état des lieux de l'existant afin d'accompagner et d'améliorer les pratiques en proposant des ressources pour s'informer ou se former.</p><p>217 personnes ont répondu à cette enquête ce qui nous a permis de dresser un panorama réaliste des pratiques actuelles relevant de la programmation en SHS.</p><p>Le fichier .json permet de recoder le nom des colonnes.</p><p>Un notebook d'analyse est disponible ici : https://github.com/emilienschultz/digit_hum_2023/blob/main/2023_Digit_Hum_Exploration_sondage_v2.ipynb</p>
EZBugs4Py: A benchmark of simple, easily reproducible Python bugs
<p>This is the appendix for paper entitled "EZBugs4Py: A benchmark of simple, easily reproducible Python bugs" submitted to MSR 2024.</p><p> </p><p>The dataset itself is available on <a href="https://github.com/gaborantal/ezbugs4py">GitHub.</a></p><p>This appendix contains:</p><ul><li>The results of GPT in the following tasks: automated program repair for to so-called "buggy" versions of the programs, the "failing" versions of the program, and the code synthesis based on the descriptions of the tasks.</li><li>The exact prompts we used in the paper.</li><li>The runner scripts to query GPT-4.</li><li>The categorization of the bugs.</li></ul><p> </p>
Python scripts for input and post-processing of fuzz sputtering TRI3DYN simulations
<p>The influence of a fuzzy surface on the physical sputtering of Mo in He plasmas has been studied with hyperspectral imaging (HSI) measurements and simulations that couple the TRI3DYN code with an impurity transport code. The 2D profiles of the Mo I line emission intensity from HSI images reveal that the sputtering yield, Y, is reduced to ~40% of the smooth-surface value due to the presence of a fuzz layer, while the angular distribution of the sputtered Mo atoms might not change significantly. The simulations reproduce the Y reduction successfully, but indicate that fuzz causes an increase in the small-angle distribution of sputtered atoms. However, the increase is too small to produce an observable change in the Mo I emission profiles. A simple analytical model that assumes a single collision mean free path for a fuzz layer and considers only the primary sputtering events qualitatively reproduces the Y reduction and the small-angle distribution enhancement, explaining the geometrical effect of fuzz on physical sputtering.</p>
Example configurations and test cases for the Python HDF5Translator framework.
<p>This is a set of use examples for the <a href="https://github.com/BAMresearch/HDF5Translator">HDF5Translator framework</a>. This framework lets you translate measurement files into a different (e.g. NeXus-compatible) structure, with some optional checks and conversions on the way. For an in-depth look at what it does<a href="https://lookingatnothing.com/?p=4087">, there is a blog post here. </a></p> <p>The use examples provided herein are each accompanied by the measurement data necessary to test and replicate the conversion. The README.md files in each example show the steps necessary to do the conversion for each. </p> <p>We encourage those who have used or adapted one or more of these exampes to create their own conversion, to get in touch with us so we may add your example to the set. </p>
Example data sets and input parameters for running various features in the Python code Dynpy
<p>This is a set of examples intended to be used with the Python code called Dynpy at <a href="https://zenodo.org/records/13241475">https://zenodo.org/records/13241475</a></p>
An Empirical Study on the Usage and Availability of Machine Learning Libraries in Open-Source Python Projects - Dataset
<p>This repository contains the dataset of the manuscript:</p> <p>"An Empirical Study on the Usage and Availability of Machine Learning Libraries in Open-Source Python Projects"</p>
PyVOLCANS: A Python package to flexibly explore similarities and differences between volcanic systems
<p>Python tool to identify analogue volcanoes via <a href="https://doi.org/10.1007/s00445-019-1336-3">VOLCANS</a>.</p> <p>The main goal of PyVOLCANS is to help alleviate data-scarcity issues in volcanology, and contribute to developments in a range of topics, including (but not limited to): quantitative volcanic hazard assessment at local to global scales, investigation of magmatic and volcanic processes, and even teaching and scientific outreach. We hope that future users of PyVOLCANS will include any volcano scientist or enthusiast with an interest in exploring the similarities and differences between volcanic systems worldwide. Please visit our <a href="https://github.com/BritishGeologicalSurvey/pyvolcans/wiki">wiki pages</a> for more information.</p>
Dataset for the research paper "How and Why Developers Migrate Python Tests from unittest to pytest"
<p>This is the dataset for the proposed paper "How and Why Developers Migrate Python Tests from unittest to pytest". <br> <br> <br> The `10_systems` zip file contains the aggregated and intermediate files for the systems used for precision and recall analysis.</p> <p>The `top_100_systems` zip file contains the aggregated and intermediate files for the top 100 python systems analyzes.</p> <p>The `__rq_reason` contains data to assess the advantages and disadvantages found in 100 issues or pull requests. The second column indicates whether issues/PRs were selected to be analyzed and the following columns indicate if the advantages (A) or disadvantages (D) are present or not.</p>
Dataset of the paper "An Empirical Study on the Fault-Inducing Effect of Functional Constructs in Python"
<p>This package contains the dataset of the manuscript "An Empirical Study on the Fault-Inducing Effect of Functional Constructs in Python"</p>
Reference Datasets for: SHAFTS (v2022.3): a deep-learning-based Python package for Simultaneous extraction of building Height And FootprinT from Sentinel Imagery
<p>These are reference building height and footprint datasets which consist of 46 cities worldwide and support the development of SHAFTS (https://github.com/LllC-mmd/3DBuildingInfoMap).</p> <p>The snapshot of original reference datasets from 46 cities and related GitHub repository has been created as a zipped file named <strong><em>SHAFTS_220527_snapshot.zip</em></strong>.</p> <p>On 2022.5.27, we added 8 additional cities from ArcGIS Hub when compared with the previous version (https://doi.org/10.5281/zenodo.6370003).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.