Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
334
datasets available to search
ShareScore release 0.7.1
Dataset results
334 results for “Python”
Empirical Analysis of Security Vulnerabilities in Python Packages
<p>This dataset contains the data files used to analyze our RQs in the manuscript "Empirical Analysis of Security Vulnerabilities in Python Packages".</p> <p>For more information on how to understand the folder structure and dataset, please read the README.md.</p>
Does Coding in Pythonic Zen Peak Performance? Preliminary Experiments of Nine Pythonic Idioms at Scale
<p><strong>Link to pre-print: </strong><a href="https://arxiv.org/abs/2203.14484">https://arxiv.org/abs/2203.14484</a><br> <br> <strong>How to run</strong></p> <ol> <li>Extract <strong>pythonnic_performance.zip </strong>and move into the extracted directory.</li> <li>Execute <strong>install.sh</strong> to install required dependencies.</li> <li>Execute <strong>run.sh </strong>to start the experiment.</li> <li>Execute <strong>python statistic.py </strong>to perform the statistical test and show data statistics.</li> </ol> <p>Note that files in <strong>submitted_output</strong> are our experimental results shown in the paper.</p> <p><br> <strong>Abstract</strong></p> <p>In field of data science, and for academics in general, the Python programming language is a popular choice, mainly because of its libraries for storing, manipulating, and gaining insight from data.<br> Evidence includes the versatile set of machine learning, data visualization, and manipulation packages used for the ever-growing size of available data.<br> The <em>Zen of Python</em> is a set of guiding design principles that developers use to write acceptable and elegant Python code.<br> Most principles revolve around simplicity.<br> However, as the need to compute on large amounts of data, performance has become a necessity for the Python programmer.</p> <p>The new idea in this paper is to empirically confirm whether writing the Pythonic way peaks performance at scale.<br> As a starting point, we conduct a set of preliminary experiments to evaluate nine Pythonic code examples by comparing the performance of both Pythonic and Non-Pythonic code snippets.<br> Our results reveal that writing in Pythonic idioms may save memory and time.<br> We show that incorporating list comprehension, generator expression, zip, and itertools.zip\_longest can save up to 7,000 MB and 32.25 seconds.<br> The results open more questions on how they could be utilized in a real-world setting.</p>
APL synapses distribution on PN and KC meshes, primary data, calcium imaging macros and python scripts from Prisco et al.
<p class="MsoBodyText">To identify and memorize discrete but similar environmental inputs, the brain needs to distinguish between subtle differences of activity patterns in defined neuronal populations. The Kenyon cells of the <i>Drosophila </i>adult mushroom body (MB) respond sparsely to complex olfactory input, a property that is thought to support stimuli discrimination in the MB. To understand how this property emerges, we investigated the role of the inhibitory anterior paired lateral neuron (APL) in the input circuit of the MB, the calyx. Within the calyx, presynaptic boutons of projection neurons (PNs) form large synaptic microglomeruli (MGs) with dendrites of postsynaptic Kenyon cells (KCs). Combining EM data analysis and <i>in vivo</i> calcium imaging, we show that APL, via inhibitory and reciprocal synapses targeting both PN boutons and KC dendrites, normalizes odour-evoked representations in MGs of the calyx. APL response scales with the PN input strength and is regionalized around PN input distribution. Our data indicate that the formation of a sparse code by the Kenyon cells requires APL-driven normalization of their MG postsynaptic responses. This work provides experimental insights on how inhibition shapes sensory information representation in a higher brain centre, thereby supporting stimuli discrimination and allowing for efficient associative memory formation.</p>
Python Type Information
<p>Python type information of 1000 projects. </p>
Additional 1000 notebooks for the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts"
<p>Additional 1000 notebooks for the review of the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts".</p> <p>The notebooks can be processed using Matroskin tool from the supplementary materials. To do that, place:</p> <p>- the folder "1k_notebooks_dataset" into the folder ".../databases/datasets"</p> <p>- the file "mapping_of_1k_notebooks.json" into the folder ".../databases/mappings"</p>
Supplementary material of "VIGIL: A Python tool for automatized probabilistic VolcanIc Gas dIspersion modeLling"
<p>This dataset includes the Supplementary material of the paper: </p> <p>Dioguardi F, Massaro S, Chiodini G, Costa A, Folch A, Macedonio G, Sandri L, Selva J, Tamburello G. VIGIL: A Python tool for automatized probabilistic VolcanIc Gas dIspersion modeLling. Ann. Geophys. [Internet]. 2022Apr.3 [cited 2022Apr.7];65(1). Available from: https://www.annalsofgeophysics.eu/index.php/annals/article/view/8796</p>
Tool and python programs for the paper "The Impact of Altering Emission Data Precision on Compression Efficiency and Accuracy of Simulations of the Community Multiscale Air Quality Model"
<p>Here is the content:</p> <p> * file dir_list which contains information about each file's content</p> <p> * the tool is used to alter a data file by keeping a specific number of significant digits for the paper "The Impact of Altering Emission Data Precision on Compression Efficiency and Accuracy of Simulations of the Community Multiscale Air Quality Model'</p> <p> * pythons program and its associated data to create each figure and table in the paper (data for Table 07 is not included due to size is larger than 50GB)</p>
Python notebooks as a pedagogical tool to teach NOT data reduction
<p>We will present a series of seven Python Jupyter Notebooks designed to teach master-level students the basic steps of data reduction for observations with the Alhambra Faint Object Spectrograph and Camera (ALFOSC). This pedagogical tool, which translates IRAF tasks into the widespread Python language, has been successfully deployed for the course “Observational Astrophysics II” at the Department of Astronomy of Stockholm University. Each notebook introduces the students to one specific task of the data reduction explaining the purpose of that task, how it is implemented and guiding the students through the completion of the task. With a hands-on approach, this allows the students to understand the reason behind each step and to test directly what is the effect of each step, thanks to interactive plots of the data and the intermediate products. In addition to its educational value, this material can be expanded to reach the quality needed for scientific works. Therefore it could offer a starting point for developing a personalised data reduction for a given scientific problem. A complete version of the material is publicly available at https://github.com/astrojuggler/data-reduction-Obs-II-course. This project has been financed by Stockholm University - Department of Astronomy with a grant earned by Professor Matthew Hayes.</p>
Dataset and Jupyter notebook for "pyDARN: A Python Software for Visualizing SuperDARN Radar Data"
<p>SuperDARN radar dataset and Jupyter notebook used to generate figures for "pyDARN: A Python Software for Visualizing SuperDARN Radar Data".</p>
Evolutionary models demonstrate rapid and adaptive diversification of Australo-Papuan pythons
<p>Lineages may diversify when they encounter available ecological niches. Adaptive divergence by ecological opportunity often appears to follow the invasion of a new environment with open ecological space. This evolutionary process is hypothesized to explain the explosive diversification of numerous Australian vertebrate groups following the collision of the Eurasian and Australian plates 25 million years ago. One of these groups is the pythons, which demonstrate their greatest phenotypic and ecological diversity in Australo-Papua (Australia and New Guinea). Here, using an updated and near complete time-calibrated phylogenomic hypothesis of the group, we show that following invasion of this region, pythons experienced a sudden burst of speciation rates coupled with multiple instances of accelerated phenotypic evolution in head and body shape and body size. These results are consistent with adaptive radiation theory with an initial rapid niche filling phase and later slow-down approaching niche saturation. We discuss these findings in the context of other Australo-Papuan adaptive radiations and the importance of incorporating adaptive diversification systems that are not extraordinarily species-rich but ecomorphologically diverse to understand how biodiversity is generated.</p>
Questões em C e Python, abordando padrões de equívoco.
<p>O material contem questões que abordam padrões de equívoco, presentes na linguagem C e Python.</p>
Updated dartmouth.tri for Isochrones python package
<p>This replaces the `dartmouth.tri` file from https://zenodo.org/record/161241.</p>
CoqPyt: Proof Navigation in Python in the Era of LLMs
<p>Replication package with code and CompCert dataset accompanying the paper "CoqPyt: Proof Navigation in Python in the Era of LLMs".<br>The file provided here is a Docker image. To use it, follow the steps:</p><p> </p><p>1. docker load < coqpyt.tar</p><p>2. docker run -it --name coqpyt -d coqpyt /bin/bash</p><p>3. docker attach coqpyt<br> </p><p>After these steps, see the README.</p>
Duplicate Question Dataset - Python
Open the record for dataset details and reuse information.
Data for: Exploring Emerging Social Media: Acquiring, Processing, and Visualizing Data with Python and OSoMe Web Tools
<p>Data collected from Bluesky and Mastodon via streaming covering the period between 2024-06-25 and 2024-07-02. Entries contains any the of following terms: biden, trump, or debate. Data also contains embedding precalculated for each dataset. The data also contains embeddings pre calculcated for the datasets.</p>
PyShEx - Python implementation of Shape Expressions
<p>ShEx interpreter for ShEx 2.0</p>
Fotografias Python Gado
<p>Fotografias utilizadas para teste de algoritmos de visão computacional com OpenCV.</p>
ICSE 2025 Artifact for "An Empirical Study on Package-Level Deprecation in Python Ecosystem"
<div> <h1>Artifact</h1> <div>This artifact includes the source code and data needed to reproduce the results of our paper.</div> <h2>Files</h2> <div> <ul> <li><strong>inactive_task</strong>: This folder contains the dataset we collected, including:</li> </ul> </div> <ul> <li> <ul> <li>A list of all packages in PyPI as of 2023.1.9 in <strong>all_package.json</strong></li> <li>A list of halted packages in <strong>halted_packages.pkl</strong></li> <li>A list of packages that haven't received any commit for a long time in <strong>long_time_no_commit.json</strong></li> <li>A list of inactive packages and their corresponding GitHub repositories in <strong>inactive_pkg_repo_list.json</strong></li> <li>A mapping of packages to their corresponding GitHub links in <strong>pgk2url.json</strong></li> <li>A dataset with details in <strong>deprecation_dataset/</strong>, which includes rationales, alternative solutions, and package characteristics (RQ1). Click [here](inactive_task/deprecation_dataset/README.md) for more details.</li> </ul> </li> </ul> <div> <ul> <li><strong>ghd_dataset</strong>: This folder contains the dependency information for PyPI, which we used to build the dependency network. Please unzip the file before use.</li> </ul> </div> <div> <ul> <li><strong>down_deps</strong>: The folder contains scripts to process our data, including</li> </ul> </div> <ul> <li> <ul> <li><strong>similar_brothers</strong>: A script to find similar brother packages of deprecated packages.</li> <li><strong>delta_of_downdeps.py</strong>: A script that calculates the gain of downstream dependencies.</li> </ul> </li> <li><strong>regression</strong>: This folder contains scripts for the models that estimate the effect of deprecation announcements, which can be used to reproduce the results of RQ2.</li> <li><strong>questionnaire_data</strong>: The folder includes the questionnaire prototype and responses (RQ3, 4). Click [here](./questionnaire_data/README.md) for more details.</li> </ul> </div>
Dataset of the Paper "Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using Copilot"
<p>This dataset contains a list of 102 code smells detected from Copilot-generated Python code, along with Python code files generated by Copilot from the <em>Repositories</em> and <em>Code</em> label, respectively. This dataset also includes Copilot Chat’s responses to fixing the 102 detected code smells. A brief description of each document and folder in the dataset is provided below:</p> <p><strong>1. files folder</strong></p> <p>contains 311 Python code files generated by Copilot. In the 311 Python files, 171 are retrieved under the <em>Repositories</em> label, indicating Python code files entirely generated by Copilot, and 140 are retrieved under the <em>Code</em> label, indicating Python code snippets generated by Copilot.</p> <p><strong>2. results of RQ1.xlsx</strong></p> <p>contains a list of 102 code smells detected from Copilot-generated Python code.</p> <p><strong>3. results of RQ2.xlsx</strong></p> <p>contains Copilot Chat’s responses to fixing the detected 102 code smells instructed by three prompts of varying detail levels.</p>
Python analyses: Dissolved oxygen
<p>Contains the python scripts associated with (i) the development of the hybrid DynQual_Random Forest model; and (ii) the analyses of the dissolved oxygen output from the hybrid model.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.