Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
334
datasets available to search
ShareScore release 0.7.1
Dataset results
334 results for “Python”
Investigando o Uso da Inteligência Artificial em Projetos Python Hospedados no GitHub
<p>This dataset was utilized in the research paper titled "Investigando o Uso da Inteligência Artificial em Projetos Python Hospedados no GitHub," accepted for publication in the 12th Workshop on Software Visualization, Evolution, and Maintenance. This study aims to explore the presence and usage of artificial intelligence libraries in Python projects hosted on GitHub.</p> <p>To achieve this, we developed a Python script that analyzes repositories listed in a CSV file, checking for dependencies related to AI. The script utilizes GitHub’s API to access repository data and inspect `requirements.txt` files for mentions of AI libraries. The dataset includes information on repositories that use specific AI libraries, such as TensorFlow, PyTorch, and scikit-learn. By examining these dependencies, the study provides insights into how frequently and in what contexts these libraries are used in real-world Python projects. This research contributes to understanding the adoption of AI technologies in software development and supports practitioners in identifying relevant AI tools in the Python ecosystem.</p>
Data and Python scripts for "Manifold increase in the spatial extent of heatwaves in the terrestrial Arctic"
<p>Data and Python scripts for reproducing the figures and results included in the manuscript "Manifold increase in the spatial extent of heatwaves in the terrestrial Arctic".</p> <p>Figures 1-4 are produced via respective Python codes. Heatwave magnitude index daily (HWMId) for ERA5-Land is available from. hw_era5land.nc file. HWMId fields for CMIP6 models are included in cmip6_hwmid.zip. The underlying data behind the figures 1-4 are included in data_to_produce_figs.zip.</p> <p>The paper is published in Rantanen, M., Kämäräinen, M., Luoto, M. <em>et al.</em> Manifold increase in the spatial extent of heatwaves in the terrestrial Arctic. <em>Commun Earth Environ</em> <strong>5</strong>, 570 (2024). https://doi.org/10.1038/s43247-024-01750-8</p>
Replication Package for "PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems"
<p>This package contains the traceback data, pre-trained models, and static word embeddings used in the paper, PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems.</p>
raw crater counts for background regions & python code
<p>JMARS .jlf shape files that contain the measurement area polygon and crater measurements (diameters, center lat/lon, and a degradation classification). And .csv files of crater measurements. Names of files indicate the region for the counts. See associated publication for descriptions of those regions.</p>
Datasets for Earth Observation Using Python: A Practical Programming Guide
<p>These are the datasets used in the <a href="https://agupubs.onlinelibrary.wiley.com/doi/book/10.1002/9781119606925">2021 Edition</a> of "Earth Observation Using Python: A Practical Programming Guide." These contain example text, csv, netCDF, and GRIB files that are used in code exercises within the book. Datasets are original files and modified public domain files from large Earth satellite data providers, such as NASA, NOAA, EUMETSAT, and ESA. These files may be used freely, but if not heavilly modified, please attribute credit to the original author for collecting the examples.</p>
Dataset for Medical Image Processing in Python Carpentries lesson
<p>This dataset contains a collection of medical imaging files for use in the <a href="https://github.com/esciencecenter-digital-skills/medical-image-processing">"Medical Image Processing with Python" lesson</a>, originally developed by the <a href="https://www.esciencecenter.nl/">Netherlands eScience Center</a>. </p> <p>The dataset includes:</p> <ol> <li>SimpleITK compatible files: MRI T1 and CT scans (<em>training_001_mr_T1.mha, training_001_ct.mha</em>), digital X-ray (<em>digital_xray.dcm</em> in DICOM format), neuroimaging data (<em>A1_grayT1.nrrd, A1_grayT2.nrrd</em>). Data have been downloaded from <a href="https://insightsoftwareconsortium.github.io/SimpleITK-Notebooks/Python_html/00_Setup.html">here</a>. </li> <li>MRI data: a T2-weighted image (<em>OBJECT_phantom_T2W_TSE_Cor_14_1.nii</em> in NIfTI-1 format). Data have been downloaded from <a href="https://zenodo.org/records/6467772">here</a>. </li> <li>Example images for the machine learning lesson: chest X-rays (<em>rotatechest.png, other_op.png</em>), cardiomegaly example (<em>cardiomegaly_cc0.png</em>).</li> <li>Array data: Array data for the Intro to Medical Imaging lesson. Numpy arrays were created by processing and manipulation of publicly available data i.e. from <a href="https://doi.org/10.1109/TNS.1974.6499235">the Schepp Logan phantom</a> and from the <a href="https://fastmri.med.nyu.edu/">NYU FastMRI dataset</a> <div> </div> </li> <li>Additional data: to be added</li> </ol> <p>These files represent various medical imaging modalities and formats commonly used in clinical research and practice. They are intended for educational purposes, allowing students to practice image processing techniques, machine learning applications, and statistical analysis of medical images using Python libraries such as scikit-image, pydicom, and SimpleITK.</p>
Data set from Fischertechnik Smart Factory Model at University of St.Gallen (Custom Python Configuration)
<p>This is about 60 mins worth of data collected from Fischertechnik Industry 9.0V smart factory model available at the University of St.Gallen.</p> <p>In this data set, we used a custom Python-based software stack to control the smart factory via a business process system (Camunda Platform) that calls the functionality of the smart factory via web services implemented in Python flask. MQTT is used to collect the data.</p> <p>Each entry in the file (low-level_log_20230206-140808.txt) corresponds to one message (as JSON object) received on a specific topic via MQTT. Each line contains all the readings of all the sensors, actuators and additional data from <strong>one </strong>CPS component (i.e., production station) at <strong>one </strong>point in time.</p> <p>The data set contains the following files</p> <ul> <li>low-level_log_20230206-140808.txt: low-level IoT data from all the sensors and actuators <ul> <li>*.bpmn: executable BPMN 2.0 models of three different processes that have been executed several times via the Camunda Platform BPM system to control the smart factory</li> </ul> </li> <li>camunda_process-instance.json: event log generated by the BPM system regarding the process instance execution</li> <li>camunda_activity-instance.json: event log generated by the BPM system regarding the activity instance execution</li> </ul> <p>Check the following publications to learn more about our research using the model factory:</p> <p>Malburg, L., Seiger, R., Bergmann, R., & Weber, B. (2020). Using physical factory simulation models for business process management research. In <em>Business Process Management Workshops: BPM 2020 International Workshops, Seville, Spain, September 13–18, 2020, Revised Selected Papers 18</em> (pp. 95-107). Springer International Publishing.</p> <p>Seiger, R., Zerbato, F., Burattin, A., García-Bañuelos, L., & Weber, B. (2020, October). Towards iot-driven process event log generation for conformance checking in smart factories. In <em>2020 IEEE 24th International Enterprise Distributed Object Computing Workshop (EDOCW)</em> (pp. 20-26). IEEE.</p> <p>Seiger, R., Malburg, L., Weber, B., & Bergmann, R. (2022). Integrating process management and event processing in smart factories: A systems architecture and use cases. <em>Journal of Manufacturing Systems</em>, <em>63</em>, 575-592.</p>
Python codes for deconstructing the effects of stochasticity on transmission of hospital-acquired infections in ICUs
<p>The inherent stochasticity in transmission of hospital-acquired infections (HAIs) has complicated our understanding of transmission pathways. It is particularly difficult to detect the impact of changes in the environment on the acquisition rate due to stochasticity. In this study, we investigated the impact of uncertainty (epistemic and aleatory) on nosocomial transmission of HAIs by evaluating the effects of stochasticity on the detectability of seasonality on admission. For doing so, we developed an agent-based model of an ICU and simulated the acquisition of HAIs considering the uncertainties in the behavior of the healthcare workers (HCWs) and transmission of pathogens between patients, HCWs, and the environment. Our results show that stochasticity in HAI transmission weakens our ability to detect the effects of a change, such as seasonality, on the acquisition rate, particularly when transmission is a low-probability event. In addition, our findings demonstrate that data compilation can address this issue, while the amount of required data depends on the size of the said change and the amount of stochasticity. Our methodology can be used as a framework to assess the impact of interventions and provide decision-makers with insight about the minimum required size and target of interventions in a healthcare facility.</p>
Technoableism & Social Media: TikTok Python Scrape
<p>This Notebook uses Deen Freelon's module "Pyktok" to scrape metadata from videos that mention terms related to disability and technology. Search terms include "wearable tech," "disability," "technology," and "biohack." </p>
Example sonifications from the Astronify open-source Python package.
<p>The file named "10_galexFlare.wav" is a sonification of a stellar flare observed by the GALEX space telescope. This sonification uses a linear stretch on the pitch range 100-10,000 Hz, with a note duration of 0.8 seconds and 0.04 seconds between notes. The file named "1_kepler12b.wav" is a sonification of a transiting exoplanet observed by the Kepler space telescope. This sonification uses a linear stretch on the pitch range 100-10,000 Hz, with a note duration of 0.5 seconds and 0.01 seconds between notes.</p>
The Relationship Between Different Python Argument-Passing Mechanisms and Fixes: An Empirical Study
<p>Replication package for the paper: The Relationship Between Different Python Argument-Passing Mechanisms and Fixes -- An Empirical Study</p>
Opportunities and Limitations of Running Python Code in the Web Browser
<p>Python is a popular programming language that is widely used for a variety of applications such as web development, data analysis, scientific computing, and education. This versatility makes it a popular choice for developers and educators who need a language that can be used for a wide variety of tasks. While Python is typically run on the server side or on the desktop, there is a growing interest in running Python code directly on the client side, i.e., in the browser.</p> <p>There are several possibilities available for running Python code in the browser, among others,</p> <ul> <li>transpiling it into JavaScript (cf. e.g., Transcrypt),</li> <li>running it by making use of an interpreter implemented in JavaScript (cf. e.g., Brython or Skulpt),</li> <li>or executed it by leveraging a Python interpreter compiled to WebAssembly (cf. e.g., Pyodide or CoWasm), an open standard defined by the World Wide Web Consortium specifying a bytecode for running programs in browsers.</li> </ul> <p>Since the choice of the right tool for a given application depends on the specific requirements and constraints, the aim of this bachelor thesis is to provide an overview of the state of the art in running Python code in the browser. This is done by presenting a variety of different tools and environments available, their capabilities and limitations, as well as some application examples and sample code. Furthermore, similar to Kiyokawa’s & Jin’s (2022) work on “A Front-End Framework Selection Assistance System [...]”, the objective of this bachelor’s thesis is to develop criteria to help developers and educators in the selection process.</p> <p> </p> <p>Link to Respository: <a href="https://git.uibk.ac.at/csav4362/running-python-in-web-browser">https://git.uibk.ac.at/csav4362/running-python-in-web-browser</a></p>
Additional ASAS-SN 100 Million Variable Star Database Python Filter CSV files
<p>Additional ASAS-SN 100 Million Variable Star Database Python Filter CSV files</p>
Datasets and python code for "Distribution of telecom Time-Bin Entangled Photons through a 7.7 km Hollow-Core Fiber"
<p>Datasets and python code for "Distribution of telecom Time-Bin Entangled Photons through a 7.7 km Hollow-Core Fiber"</p>
3D CAD models exemples to run "ArtificialReef_Complexity" Python script (STL files)
<p>Here you will find 3D CAD models in STL format.</p> <p>These are 3D CAD models of fractal pyramid.</p> <p>These STL files can be used as an example to run the Python script "ArtificialReef_Complexity: v.1.3" available on GitHub (<a href="https://github.com/ELI-RIERA/ArtificialReef_Complexity/tree/V1.3">https://github.com/ELI-RIERA/ArtificialReef_Complexity/tree/V1.3</a>)</p>
Python code for Titan's spin state
Open the record for dataset details and reuse information.
Custom made python script using network assignment and scoring to estimate the impact of biological processes.
Open the record for dataset details and reuse information.
Python codes for deconstructing the effects of stochasticity on transmission of hospital-acquired infections in ICUs
Open the record for dataset details and reuse information.
Cardio PyMEA: A user-friendly, open-source Python application for cardiomyocyte microelectrode array analysis
Open the record for dataset details and reuse information.
Type Ia supernovae from non-accreting progenitors: data, python scripts and mesa inlists
<p>This release contains the inlists and final profiles described in: Antoniadis et al., "Type Ia supernovae from non-accreting progenitors" Mesa v. 10398</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.