Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
334
datasets available to search
ShareScore release 0.7.1
Dataset results
334 results for “Python”
Technical Leverage Analysis in the Python Ecosystem
<p>Technical Leverage Analysis in the Python Ecosystem</p> <p>This dataset is the original dataset used in the publication [1]. It includes 21205 distinct package versions from the top 600 Python packages. An online demo for computing the proposed metrics for real-world software libraries is also available under the following URL: https://techleverage.eu/.</p> <p>This work has been partially funded by the EU under the H2020 Program AssureMOSS (Grant n. 952647). </p> <p>[1] DOI: 10.1007/s10664-023-10355-2</p>
GCMMA-MMA-Python: Python implementation of the Method of Moving Asymptotes
This record contains the Python implementation of the Method of Moving Asymptotes (MMA), originally developed and written in MATLAB by Krister Svanberg. The MMA algorithm is used for solving non-linear programming problems. Users of this code are encouraged to inform Krister Svanberg of their application and intentions via email, as provided on his website. When publishing work that uses this code, please cite Krister Svanberg's original academic work.
Introduction to Ancient Metagenomics Textbook (Edition 2025): Introduction to Python and Pandas
<p>Data and conda software environment file for the chapter 'Introduction to Python and Pandas' of the SPAAM Community's textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>
ManyTypes4Py: A Benchmark Python Dataset for Machine Learning-Based Type Inference
<ul> <li>The dataset is gathered on Sep. 17th 2020 from GitHub.</li> <li>It has <em>clean</em> and <em>complete</em> versions (from v0.7): <ul> <li>The clean version has 5.1K <strong>type-checked </strong>Python repositories and 1.2M type annotations.</li> <li>The complete version has 5.2K Python repositories and 3.3M type annotations.</li> </ul> </li> <li>The dataset's source files are type-checked using <a href="https://mypy.readthedocs.io/">mypy</a> (clean version).</li> <li>The dataset is also de-duplicated using the <a href="https://github.com/saltudelft/CD4Py">CD4Py</a> tool.</li> <li>Check out the <strong>README.MD</strong> file for the description of the dataset.</li> <li>Notable changes to each version of the dataset are documented in <strong>CHANGELOG.md</strong>.</li> <li>The dataset's scripts and utilities are available on <a href="https://github.com/saltudelft/many-types-4-py-dataset">its GitHub repository</a>.</li> </ul>
Spatial transcriptome analysis defines heme as a hemopexin-targetable inflammatoxin in the brain - Datasets and Python notebooks
<p>This dataset and the associated Python notebooks and R-code are related to the publication "Spatial transcriptome analysis defines heme as a hemopexin-targetable inflammatoxin in the brain".</p>
Python code for "Evolutionary epidemiology consequences of trait-dependent control of heterogeneous parasites"
<p>The file contains the Python code used to run the agent-based simulation of the selection-mutation model presented in "Evolutionary epidemiology consequences of trait-dependent control of heterogeneous parasites"</p>
HURRECON Model for Estimating Hurricane Wind Speed, Direction, and Damage (R and Python)
The HURRECON model estimates wind speed, wind direction, enhanced Fujita scale wind damage, and duration of EF0 to EF5 winds as a function of hurricane location and maximum sustained wind speed. Results may be generated for a single site or an entire region. Hurricane track and intensity data may be imported directly from the US National Hurricane Center's HURDAT2 database. HURRECON is available in R and Python. The R version is available on CRAN as HurreconR. The model is an updated version of the original HURRECON model written in Borland Pascal for use with Idrisi (see HF025). New features include support for: (1) estimating wind damage on the enhanced Fujita scale, (2) importing hurricane track and intensity data directly from HURDAT2, (3) creating a land-water file with user-selected geographic coordinates and spatial resolution, and (4) creating plots of site and regional results. The model equations for estimating wind speed and direction, including parameter values for inflow angle, friction factor, and wind gust factor (over land and water), are unchanged from the original HURRECON model. For more details and sample datasets, see the project website on GitHub (https://github.com/hurrecon-model).
EXPOS Model for Estimating Topographic Exposure to Wind (R and Python)
The EXPOS model uses a digital elevation model (DEM) to estimate exposed and protected areas for a given hurricane wind direction and inflection angle. The resulting topograhic exposure maps can be combined with output from the HURRECON model to estimate hurricane wind damage across a region. EXPOS is available in R and Python. The R version is available on CRAN as ExposR. The model is an updated version of the original EXPOS model written in Borland Pascal for use with Idrisi (see HF024). For more details and sample datasets, see the project website on GitHub (https://github.com/expos-model).
Evaluating the microscopic effect of brushing stone tools as a cleaning procedure [Python analysis]
<p>This upload includes the following files related to the Python analysis:</p> <ol> <li>Raw data as a XLSX table (brushing_v2.xlsx), i.e. results from R Script #1 (see <a href="https://doi.org/10.5281/zenodo.3632517">https://doi.org/10.5281/zenodo.3632517</a>)</li> <li>Python script of the whole analysis (RunEveryParameter.py)</li> <li>Convenience script for running RunEveryParameter.py in background and logging all output (RunSingleParametesBash.sh)</li> <li>Log file for output of sampling from the model for each parameter in a loop (logAll.txt)</li> <li>Jupyter notebooks of the analysis run on <em>epLsar</em> as an example (Notebook_SingleParameter.inpyb) and of a summary of the whole analysis (Notebook_Overview.ipynb), plus associated HTML output files (*.html)</li> <li>For each parameter:</li> </ol> <ul> <li>Full samples of parameter values (*.pkl)</li> <li>Energy plots of Hamiltonian Monte Carlo (*_Energy.pdf)</li> <li>Contrast plots between each treatment (BrushDirt = Is_Is, BrushNoDirt = Is_No, RubDirt = No_Is) and the control (No_No) (*_Contrasts.pdf)</li> <li>Trace plots for each parameter (*_Trace.pdf)</li> <li>Distribution of posteriors for each parameter (*_Posterior.pdf)</li> <li>Prior and posterior predictive distributions for each parameter (*_PriorPosterior.pdf)</li> </ul> <p>Instructions to download all files at once are given here: <a href="https://doi.org/10.5281/zenodo.4011952">https://doi.org/10.5281/zenodo.4011952</a></p>
OntoUML Vocabulary Python Library
A Python library designed to simplify the development of software applications using the OntoUML vocabulary.
Data and code for "Tweezepy: A Python package for calibrating forces in single-molecule video-tracking instruments"
<p>Data and code for "Tweezepy: A Python package for calibrating forces in single-molecule video-tracking instruments."</p> <p>Data includes representative real and simulated bead trajectories used in the manuscript.</p> <p>Code includes all simulations, analysis, and plot details for the Figures in the manuscript. </p> <p>See included README.txt for more details.</p>
Python Time Normalized Superposed Epoch Analysis (SEAnorm) Example Data Set
<p>Solar Wind Omni and SAMPEX ( Solar Anomalous and Magnetospheric Particle Explorer) datasets used in examples for <a href="https://github.com/samwalton7645/SEA_Code">SEAnorm</a>, a time normalized superposed epoch analysis package in python.</p> <p>Both data sets are stored as either a HDF5 or a compressed csv file (csv.bz2) which contain a Pandas DataFrame of either the Solar Wind Omni and SAMPEX data sets. The data sets where written with pandas.DataFrame.to_hdf() and pandas.DataFrame.to_csv() using a compression level of 9. The DataFrames can be read using pandas.DataFrame.read_hdf( ) or pandas.DataFrame.read_csv( ) depending on the file format. </p> <p>The Solar Wind Omni data sets contains solar wind velocity (V) and dynamic pressure (P), the southward interplanetary magnetic field in Geocentric Solar Ecliptic System (GSE) coordinates (B_Z_GSE), the auroral electrojet index (AE), and the Sym-H index all at 1 minute cadence. </p> <p>The SAMPEX data set contains electron flux from the Proton/Electron Telescope (PET) at two energy channels 1.5-6.0 MeV (ELO) and 2.5-14 MeV (EHI) at an approximate 6 second cadence.</p> <p> </p>
LLM generated Python Compiler Test Dataset
<p>This dataset is generated by integrating Large Language Models (LLMs) with AFL++ fuzzing to enhance compiler testing for CPython. It includes original Python test scripts created by LLMs such as Mistral 7B, Codellama 7B, and Gemma 7B, targeted at various compiler functionalities. These scripts were subjected to fuzzing, resulting in a rich collection of test cases that tests potential vulnerabilities. An optional minimization process with AFL-cmin refined the dataset, ensuring it focuses on test cases that significantly contribute to code coverage and bug discovery. This dataset serves as a valuable resource for improving compiler design and testing efficiency, supporting further research and development in AI-driven software testing methods.</p> <p>please see references for citations of software used in this development</p>
Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline
<p>This upload contains the HZV029 Plasma and HZV029 Two-Phase dataset for reviewers of the "Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline" submission. </p> <p>Both datasets will be uploaded to metabolomics workbench and the upload completed before final publication of the manuscript. For the he HZV029 Plasma datasets only the final run is included for any sample (i.e., failed injections or other samples with data quality issues that were reran during acquisition were omitted).</p> <p>Also included in the upload is the source code for the MetDataModel and the pcpfm at the time of manuscript re-submission and the pcpfm itself. If you find this upload in the future, please check out the github repos for more updated versions:</p> <p>https://github.com/shuzhao-li-lab/PythonCentricPipelineForMetabolomics</p> <p>https://github.com/shuzhao-li-lab/metDataModel</p> <p>The github repo does not store the input the data for space reasons, they only have the notebooks. However, the .zip here has both the notebooks by themselves in the notebook subdirectory and a separate directory with the notebooks and the data used to generate all the figures and results in the manuscript.</p> <p><strong>Some information that is needed to rerun this analysis:</strong></p> <p>Sequence files are critical to the functioning of the pipeline. The sequence files for all analyses are provided under sequence_files.zip. These can be used to recapitulate the analysis by eitehr changing the filepath to each acquisition to where you put it on your sytem or by placing the sequence file in the same directory as the mzml or raw. In the latter case, the pipeline will search for filenames matching the sample names. The sequence files also store some sample metadata such as the type of sample a given acquisition is (unknown, pooled, qc, etc...)</p> <p>.raw to .mzML conversion works well on MacOS but may not work well on other systems. You will need to use the ability to specify your own conversion command or convert files outside of the pipeline. </p> <p>To replicate the results, you do need to have the annotation sources downloaded which can be done using the pipeline. MS2 annotation requires the files in the AcquireX directory which is MS2 acquisitions on pooled HZV029 plasma samples.</p> <p>For the comparison between MetaboAnalystR and the pcpfm, subsets of the datasets were used. These subsets and the sequence files are in Subsets_for_performance_testing.zip. The sequences are also in the sequence_files directory as well</p> <p>The notebooks reference data in the analysis folders. Copies of these files are located with the notebooks to ease reproduction of the exact results in the paper; however, to do so, you will need to change paths to this data in the notebook. This lets the notebooks be ran during a rerun without copying intermediates back and forth and it keeps the github repo clean.</p> <p><strong>Version History:</strong></p> <p>This version is after reviewer comments and is for resubmission.</p> <p> </p> <p><strong>Contributions:</strong></p> <p>Joshua M Mitchell implemented the pipeline and was first author on the manuscript. Shuzhao Li is the corresponding author on the manuscript. </p> <p>Maheshwor Thapa performed the experiments to collect the HZV029 data. Yuanye Chi helped with testing and documenting the pipeline. </p> <p>Jiangou (Jeff) Xia and Zhiqiang Pang provided the R portion of the analysis. </p>
Supporting Jupyter Python notebook for "A new class of efficient randomized benchmarking protocols"
<p>Python notebook containing the code used to generate the data for figure 2 in the appendix of "A new class of efficient randomized benchmarking protocols" (arXiv:1806.02048).</p>
Dataset supplementing the journal article describing the pyfastspm Python package
<p>This dataset represents a fast scanning tunneling microscopy image sequence, acquired with the FAST SPM module, showing a reduced Fe3O4(001) surface observed at 657 K during oxygen exposure. This is used as example dataset in the journal article describing the pyfastspm Python package (SoftwareX 21 (2023) 101269, Figure 3 and 4, Supporting Information).</p>
Brown Dwarfs are Violet: Python Tools for the Estimation of Human-eye Colors of Stars and Substellar Objects
<p>The accompanying files include a Python Jupyter notebook (and associated data files read in by the Python code) that carry out the calculations described by Cranmer (2023), talk 246.05 presented at the 241st Meeting of the American Astronomical Society (AAS) in Seattle, Washington. The abstract of the talk is provided here:</p> <p>There has always been interest in the perceived colors of the stars. They were key to the development of the H-R diagram, and they are also used widely in educational and public-outreach imagery. Thus, it is useful to develop software tools to compute these colors, as accurately as possible, from spectral energy distributions. This presentation follows up on an RNAAS paper (<a href="https://ui.adsabs.harvard.edu/abs/2021RNAAS...5..201C/abstract">Cranmer 2021</a>) that presented a collection of objective (CIE coordinate) and subjective (RGB triple) colors for main-sequence stars and brown dwarfs. A new empirical method of converting from CIE to RGB values is described, and results for various stellar spectra are presented. Although brown dwarfs over a wide range of effective temperatures (400 to 2000 K) emit most of their flux in the infrared, their visible spectra often exhibit a local maximum around a strong dip in the Na I cross section at 0.4-0.5 microns. Thus, they may appear purple to human eyes. Also, the hottest (O-type) main-sequence stars may appear even "bluer than the blue sky" because of Paschen continuum absorption. This presentation will update earlier stellar and brown-dwarf color estimates using more recently published synthetic spectra, and it will also investigate the effects of atmospheric absorption, over a range of air-mass values, on these perceived colors. Python Jupyter notebooks that carry out these calculations will be uploaded to the Zenodo repository for open-access distribution.</p> <p><strong>NOTE 1: </strong>The algorithms described here, for computing RGB triples, ought to be considered as preliminary results in ongoing research; i.e., they need additional testing and validation by comparing to the results of other more established ways of converting astronomical spectra to perceived colors.</p> <p><strong>NOTE 2:</strong> These files follow on from those provided in another Zenodo upload associated with the 2021 RNAAS paper: <a href="https://doi.org/10.5281/zenodo.5293307">https://doi.org/10.5281/zenodo.5293307</a></p>
Demonstration data for the python package applefy.
<p>Demonstration data for the python package applefy. </p> <p>30_data: Contains the NACO L' dataset of Beta Pic (planet removed) as used in the user documentation</p> <p>70_results: Contains the results of the user documentation tutorials</p> <p>laplace_lookup_tables.csv: Contains the lookup table for the LaplaceBootstrapTest</p>
Default datasets for NAU-PIXEL/roughness Python package
<p>Default datasets for running the roughness python package (see docs at <a href="https://nau-pixel.github.io/roughness/">https://nau-pixel.github.io/roughness/</a>).</p>
Preprocessed Python Code Corpus
<p>A preprocessed code corpus for the Python programming language.<br> The corpus was used for the experiments in the paper Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code.<br> It contains preprocessed-tokenized files for training, validation, testing, and BPE encoding learning.<br> The BPE segmented versions of the above files are also included for three different encoding sizes i,e., 2000, 5000, and 10000 BPE merge operations as well as the learned BPE encodings.<br> Similar versions are also contained for splitting compound identifiers on camelCase and snake_case as in (Allamanis et al., 2015) as well as the corresponding subtoken maps.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.