Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
281
datasets available to search
ShareScore release 0.9.0
Dataset results
281 results for “source code”
Source code for dynamic models and simulations of mate sampling behavior
Open the record for dataset details and reuse information.
Code from: Source–sink dynamics explains the coexistence of the invasive pest <em>Dryocosmus kuriphilus</em> and its biological control agent <em>Torymus sinensis</em> across French Eastern Pyrenees
Open the record for dataset details and reuse information.
Codes and source data files for: Proximity labeling identifies LOTUS domain proteins that promote the formation of perinuclear germ granules in C. elegans
Open the record for dataset details and reuse information.
Data and source code from: Contingency and selection in mitochondrial genome dynamics
Open the record for dataset details and reuse information.
Dataset and Code for "Mining and Predicting Micro-Process Patterns of Issue Resolution for Open Source Software Projects"
<p>Dataset and Code for "Mining and Predicting Micro-Process Patterns of Issue Resolution for Open Source Software Projects" with README included</p>
Open source physiological data and physiological-based kinetic model code for the chicken (Gallus gallus domesticus)
<p>This excel file and mode code (DOI:10.5281/zenodo.3603114) provides:</p> <p>1. Physiological parameters and associated inter-individual variability (sample size, mean, coefficient of variation,) for chicken (<em>Gallus gallus domesticus</em>). These physiological parameters were estimated based on the results of extensive literature searches and specific experimental data described in Lautz et al., (2020).</p> <p>2. An R code for the generic chicken physiologically based model as well as the “soboljansen” code to carry out sensitivity analysis using sobol plots. The code for the generic model allows to run:</p> <p>a. A deterministic PBK model which represents only a single animal.</p> <p>b. A probabilistic PBK model to simulate individual differences in physiological parameters within a population. Sensitivity analyses can be performed to identify which parameters have the most impact on the model’s outputs. Predictions can be compared with experimental data. The model can be used to assess the influence of physiological parameters on the kinetics of chemicals. For PBK modelling purposes, species and chemical specific kinetics (e.g clearance, absorption rate, etc…) should be provided by the user.</p> <p>The full data collection and implementation of the models using case studies are described in (Lautz et al., 2020).</p> <p><strong>The dataset providing the physiological parameters is available in Excel.<br> The R code is presented as meta data to be implemented in R.</strong></p>
Datasets for: Semantic Robustness of Models of Source Code
<p>Datasets for Semantic Robustness of Models of Source Code.</p> <p>Includes the c2s/java-small, csn/java, csn/python, and sri/py150 in the following representations:</p> <ol> <li>Raw [in raw.tar.gz]</li> <li>Normalized [in normalized.tar.gz]</li> <li>Pre-processed (<em>ast-paths and tokens</em>) [in preprocessed.tar.gz]</li> <li>Transformed [in transformed.tar.gz] <ol> <li>Normalized <ol> <li>transforms.All</li> <li>transforms.ShuffleLocalVariables</li> <li>transforms.ShuffleParameters</li> <li>transforms.RenameLocalVariables</li> <li>transforms.RenameFields</li> <li>transforms.RenameParameters</li> <li>transforms.ReplaceTrueFalse</li> <li>transforms.InsertPrintStatements</li> <li>transforms.Identity</li> </ol> </li> <li>Pre-processed (<em>ast-paths and tokens</em>) <ol> <li>transforms.Identity</li> <li>transforms.InsertPrintStatements</li> <li>transforms.ReplaceTrueFalse</li> <li>transforms.RenameParameters</li> <li>transforms.RenameFields</li> <li>transforms.RenameLocalVariables</li> <li>transforms.ShuffleParameters</li> <li>transforms.ShuffleLocalVariables</li> <li>transforms.All</li> </ol> </li> </ol> </li> </ol>
Replication Package: A Study on the Accuracy of OCR Engines for Source Code Transcription from Programming Screencasts
<p>The replication package of the paper "A Study on the Accuracy of OCR Engines for Source Code Transcription from Programming Screencasts" including the dataset, results and tools</p>
Model input and output, performance measures and modifications in the source code for PALM simulations on Mäkelänkatu in Helsinki, Finland
<p>This dataset is for air quality simulations conducted on Mäkelänkatu in Helsinki, Finland, using the PALM model system 6.0.</p> <p>By default, simulations use modelled data as boundary conditions. This includes modelled meteorological data from MEPS (MetCoOp Ensemble Prediction System) and air pollutant background concentrations from the ADCHEM model. Alternatively, measured meteorology from the Kivenlahti mast in Espoo, Finland, and aerosol size distribution from the SMEAR III station in Kumpula, Helsinki, is applied.</p> <p>Simulations:</p> <ul> <li>9 June morning: <ul> <li>0609_morning: use modelled boundary conditions for meteorology and air pollutants</li> <li>0609_morning_allmet_smear: use measured boundary conditions for meteorology and aerosol size distribution</li> <li>0609_morning_smear: use modelled boundary conditions for meteorology and measured for aerosol size distribution</li> <li>0609_morning_wd_smear: use modelled boundary conditions for meteorology, but modify the wind direction by using the measured wind direction at Kivenlahti. For aerosol size distribution, use the measured boundary conditions.</li> <li>0609_morning_wdk_smear: use modelled boundary conditions for meteorology, but modify the wind direction by using the measured wind direction at SMEAR III. For aerosol size distribution, use the measured boundary conditions.</li> <li>precursor_0609_morning: precursor with modelled boundary conditions for meteorology</li> <li>precursor_0609_morning_allmet: precursor with measured boundary conditions for meteorology</li> <li>precursor_0609_morning_wd: precursor with modelled boundary conditions, but the wind direction is modified by using the measured wind direction at Kivenlahti.</li> <li>precursor_0609_morning_wdk: precursor with modelled boundary conditions, but the wind direction is modified by using the measured wind direction at SMEAR III.</li> </ul> </li> <li>9 June evening: <ul> <li>0609_evening: use modelled boundary conditions for meteorology and air pollutants</li> <li>0609_evening_allmet_smear: use measured boundary conditions for meteorology and aerosol size distribution</li> <li>precursor_0609_evening: precursor with modelled boundary conditions for meteorology</li> <li>precursor_0609_evening_allmet: precursor with measured boundary conditions for meteorology</li> </ul> </li> <li>12 December morning: <ul> <li>1207_morning: use modelled boundary conditions for meteorology and air pollutants</li> <li>1207_morning_allmet_smear: use measured boundary conditions for meteorology and aerosol size distribution </li> <li>precursor_1207_morning: precursor with modelled boundary conditions for meteorology</li> <li>precursor_1207_morning_allmet: precursor with measured boundary conditions for meteorology</li> </ul> </li> </ul> <p> </p> <p>Datasets are given separately for the root (no suffix), parent (suffix _N02) and child (_N03) domain. The content is following:</p> <ul> <li>input_monitoring_output_usercode <ul> <li>Input data <ul> <li><run_identifier>_chemistry: emission data for gases</li> <li><run_identifier>_dynamic: initialisation and forcing data for meteorological variables and air pollutants</li> <li><run_identifier>_p3d: parameter file for model steering</li> <li><run_identifier>_salsa: emission data for aerosol particles</li> <li><run_identifier>_static: topography information</li> </ul> </li> <li>Simulation performance information <ul> <li><run_identifier>_cpu: information on the CPU time consumed</li> <li><run_identifier>_header: information about the selected model parameters</li> <li><run_identifier>_rc: time step control output</li> </ul> </li> <li>Output data <ul> <li><run_identifier>_av_masked_N03_M01.nc: temporally averaged wind speed data close to the ground</li> <li><run_identifier>_av_masked_N03_M04.nc: temporally averaged aerosol particle concentration data close to the ground</li> <li><run_identifier>_av_masked_N03_M06.nc: temporally averaged aerosol particle concentration data in a vertical column next to the air quality monitoring station on Mäkelänkatu</li> <li><run_identifier>_av_masked_N03_M07.nc: temporally averaged aerosol particle concentration data in a vertical column on the other side of the street from the air quality monitoring station on Mäkelänkatu</li> <li><run_identifier>_pr.nc: temporally vertical profile data on meteorological variables</li> <li><run_identifier>_ts.nc: flow statistics data</li> </ul> </li> <li>Modifications made to the source code (PALM revision, https://palm.muk.uni-hannover.de/trac/browser?rev=4416, last access: 10 Sept 2019) <ul> <li>chem_gasphase_mod.f90: chemical mechanism salsa+simple</li> <li>chem_emissions_mod.f90 (modifications indicated with "MONA")</li> <li>user_module.f90 (modification listed under "Current revisions")</li> <li>Makefile (modification listed under "Current revisions")</li> </ul> </li> </ul> </li> </ul> <p>See the PALM model webpage (https://palm.muk.uni-hannover.de) for details.</p> <p> </p>
Dynamic Load Balancing for Predictions of Storm Surge and Coastal Flooding-Model setup and source code
<p>Source code and model setup/inputs for the paper titled "Dynamic Load Balancing for Predictions of Storm Surge and Coastal Flooding" article. Simulations were conducted using a modified version of ADCIRC+DLB (ADCIRC + Dynamic Load Balancing) on unstructured triangular meshes.</p> <p>Contains:</p> <ol> <li>Model input files. <ol> <li>ADCIRC model input files for the ideal channel setup and Hurricane Irene simulation (*.13, *.14, *.15)</li> </ol> </li> <li>Zipped archive of the ADCIRC code (adcirc-cg-DLB.zip) used to produce the simulations for the paper.</li> <li>Step-by-step compilation and usage instructions for ADCIRC+DLB. <ol> <li>Installation.html </li> <li>Usage.html</li> </ol> </li> </ol>
Amory et al. (2021), Geoscientific Model Development : data, model outputs and source code
<p><strong>Data and model outputs for the replication of the analysis made in:</strong><br> (see the published version of this article in Geoscientific Model Development, 2021 - please cite this version if you use these data)<br> C. Amory, C. Kittel, L. Le Toumelin, C. Agosta, A. Delhasse, V. Favier, and X. Fettweis: Performance of MAR (v3.11) in simulating the drifting-snow climate and surface mass balance of Adelie Land, East Antarctica, Geoscientific Model Development, accepted, 2021. </p> <p>See README.txt for a full description of the dataset content</p> <p>Please contact me at amory.charles@live.fr if you need other half-hourly outputs or for more details on the dataset</p>
Test cases (input), test scripts (output), and source code
<p>Test cases (input), test scripts (output), and source code used in the paper "NLP-assisted Test Generation Toward Script-free Web Testing" submitted to ICSME2021 NIER track</p>
Supporting data and source code for Hnilica et al. (submitted to HESS)
<p>Data and code to reproduce the results and plots presented in Technical note: Changes of cross- and auto-dependence structures in climate projections of daily precipitation and their sensitivity to outliers (submitted to Hydrology and Earth System Sciences)</p>
Refactoring Code Smells in Open Source Projects: A Hands-on Approach to Teaching Software Maintenance
<p>Code smells are suboptimal code structures that can undermine software quality and maintainability. On the one hand, software engineers commonly apply refactoring techniques to address these deficiencies and improve internal quality attributes. On the other hand, when performed manually and without discipline, refactoring can lead to code degradation. Despite its importance, refactoring and code smells are rarely explored in depth in undergraduate computing courses, which can be reflected in industry practices. To address this gap, this paper presents a hands-on approach to teaching code smell refactoring through contributions to Open Source Software (OSS) projects, an environment where developers with diverse skill levels collaborate, and maintaining code quality is particularly challenging. Code smells accumulate over time in such scenarios, hindering software evolution and collaboration. Our study in two undergraduate Software Quality and Software Maintenance courses expands on previous findings by incorporating an in-depth analysis of students’ learning experiences. The results indicate that: (i) students rec- ognized improvements in code quality after refactoring; (ii) they identified strong connections between refactoring, testing, and debugging; (iii) their confidence decreased when refactoring required changes across multiple files; (iv) code complexity posed a significant challenge to refactoring; (v) students’ choices of refactoring techniques were influenced by project structure and personal preferences, often combining multiple techniques to address a single smell; (vi) in some cases, refactoring introduced new code smells; (vii) the longest refactoring efforts were also the most likely to reintroduce code smells; (viii) contributing to OSS projects improved students’ programming skills and fostered a sense of professional growth; (ix) students faced challenges in understanding OSS contribution processes, particularly regarding issue resolution, adherence to contribution guidelines, and responding to maintainer feedback; (x) automated checks and review workflows varied across projects, affecting students’ ability to submit successful contributions; and (xi) despite these challenges, engagement with OSS enabled students to gain practical experience in collaborative software development. Our findings offer valuable insights for software engineering educators seeking to integrate refactoring practices into coursework while leveraging OSS contributions as an educational tool.</p>
Inlist and Source Code Files for "Fossil Signatures of Main-sequence Convective Core Overshoot Estimated through Asteroseismic Analyses"
<p>MESA (r12778) and GYRE (version 6.0) inlist files used in the work described in "Fossil Signatures of Main-sequence Convective Core Overshoot Estimated through Asteroseismic Analyses".</p><p>Two subdirectories are provided in the archive:</p><p>1) The directory called "inlists" contains different MESA inlist files for different evolutionary period (pms=pre main sequence, ms=main sequence, and rgb=red giant branch) of the stellar model. The file named "inlist_0all" is applied to all evolutionary periods. The different inlist files for the different evolutionary states are called by putting their names in the "inlist" file, for example the include file named "inlist" evolves a MESA model from the pre main sequence until ZAMS (using the stop_near_zams = .true. option in the "inlist_1pms" file). An example gyre (version 6.0) inlist is also included in the "inlists" subdirectory. The Python scripts used to evaluate the matrix elements (which are used to determine the dipolar mixed-mode frequencies) discussed in this work are available at https://gitlab.com/darthoctopus/mesatricks. </p><p>2) Custom stopping conditions (used for stopping a model before the red giant branch) as well as a custom diffusion cutoff (see Viani et al. 2018, ApJ, 858, 28) are included in the run_star_extras.f file in the "src" subdirectory. </p>
Source code and model outputs for 'Increasing Aerosol Direct Effect Despite Declining Global Emissions'
<p>Source code, model outputs and python scripts used for the publication of Hermant, A., Huusko, L., & Mauritsen, T: Increasing Aerosol Direct Effect Despite Declining Global Emissions</p>
The evaluation data and source codes of a new conceptual coupled Earth system model and the MOC box model.
<p>The dataset contains the results of a conceptual Atmosphere-Ocean-Ice-Land coupled Earth system model and a MOC box model and the evaluation data of their.</p>
X-ray Fluorescence Core Scanning Dataset and Calibration Source Code
<p>This repository includes the raw output from X-ray fluorescence (XRF) core scanning, reference element concentration data, and the source code used for calibrating XRF counts. It is associated with the following publication:</p> <p>Kabiri S, Holden NM, Flood RP, Turner JN, O’Rourke SM. X-ray Fluorescence Core Scanning for High-Resolution Geochemical Characterisation of Soils. Soil Systems. 2024; 8(2):56. https://doi.org/10.3390/soilsystems8020056.</p> <p>This research was funded by the Irish Research Council, grant no. IRCLA/2017/137.</p>
Build Prediction in Continuous Integration Using Textual Analysis of Source Code and Traditional Software Metrics
<p>Continuous Integration (CI) systems integrate code changes committed by software developers, tests the results of the integration, and feed developers with information about the outcome of the integration and testing. Predicting the outcome of the integration is important since it reduces the feedback time between the CI system and the developers. This data-set comprises of historical code changes extracted from the TravisTorrent data-set (found in the train-lines folder) and their corresponding feature vectors (found in the train-bag-of-words folder) for Java projects. It also includes a set of files that contains historical build records and a set of traditional software metrics.</p>
Dataset and model codes for "Large daytime molecular chlorine missing source at a suburban site in East China"
<p>This dataset is for measurements of trace gases (Cl<sub>2</sub>, ClNO<sub>2</sub>, N<sub>2</sub>O<sub>5</sub>, HONO, NO<sub>2</sub>, NO, O<sub>3</sub>, NH<sub>3</sub>, SO<sub>2</sub>, CO, VOCs), aerosols (Cl<sup>−</sup>, NO<sub>3</sub><sup>−</sup>, SO<sub>4</sub><sup>2−</sup>, NH<sub>4</sub><sup>+</sup>, organics), and meteorological parameters (<em>T</em>, RH, <em>j</em><sub>NO2</sub>) at a suburban site (32.12°N, 118.95°E) in Nanjing, China during April 13-20, 2018.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.