Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
281
datasets available to search
ShareScore release 0.9.0
Dataset results
281 results for “source code”
Source Code Snippets and Quality Analytics Dataset
<p>This dataset contains the Java code snippets of <a href="https://github.com/github/CodeSearchNet">CodeSearchNet</a>, processed along with their abstract syntax trees and clustered according to their similarity. It also includes static analysis metrics, PMD violations and readability metrics for each snippet.</p> <p>You can use the dataset simply with the following steps:</p> <p> 1. Download the data.</p> <p> 2. Navigate to the download folder and use the mongorestore (<a href="https://docs.mongodb.com/manual/reference/program/mongorestore/">https://docs.mongodb.com/manual/reference/program/mongorestore/</a>) command. (Have in mind to use the --gzip flag)</p>
Supporting dataset for tertiary study on source code metrics
<p>The dataset includes search results (SearchResultsFromScopusIEEEACM June23.xlsx) and compiled results from the included secondary studies (Extracted Data For Tertiary Study June23 15 SS.xlsx). </p> <p>Search results description: Please use Figure 1 from paper to trace the data made available. The excel file contains the following data</p> <p>1) validation set of studies used to quasi-gold standard validation of search string. </p> <p>2) known set of papers used to formulate the search string</p> <p>3) search results from databased Scopus, IEEE, and ACM</p> <p>4) secondary studies from tertiary study on code smells by Lacerda et al.</p> <p>5) Combined results from all sources</p> <p>6) papers removed after duplicates removal</p> <p>7) Preliminary search results</p> <p>8) Search Strings used for Scopus, IEEE, ACM</p> <p>Extracted Data Description: Extracted data contains the following:</p> <p>1) meta data and characteristics of secondary studies (title, author, year of publication)</p> <p>2) Raw data on evidence on maintainability, reliability and security from included set of studies</p> <p>3) compiled data on evidence on maintainability, reliability and security from included set of studies </p> <p>4) Description of prediction models used</p> <p>5) Description of source code metrics</p> <p>6) Secondary studies that used metrics but reported no evidence and secondary studies removed due to low DARE score</p> <p>7) Quality assessment score of included secondary studies</p> <p>8) Quality assessment questions for (additional quality assessment score)AQAS</p> <p> </p> <p> </p>
Supplementary data for a systematic literature review on source code similarity measurement and clone detection: techniques, applications, and challenges
<p>The Microsoft Excel files containing the supplementary data and diagrams for the paper:</p> <p><strong>A systematic literature review on source code similarity measurement and clone detection: techniques, applications, and challenges</strong></p> <p>The article is under review in the Journal of Systems and Software.</p> <p>In this version, the literature search has been performed between April and May 2023, and studies published until that date has been mentioned in the attached Excel files.</p> <p>This version (v3.4.0) corresponds to the third revision (R3) of the manuscript.</p> <p> </p>
An Empirical Comparison of Pre-Trained Models of Source Code
<p>The replication package of the paper "An Empirical Comparison of Pre-Trained Models of Source Code". For the source code, please refer to <a href="https://github.com/NougatCA/FineTuner">https://github.com/NougatCA/FineTuner</a>.</p>
CPT-1 pre-computed whole-proteome variant effect predictions and model source code
<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p><p>This repository contains the variant effect predictions of CPT-1 for 18,602 human proteins, initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects". The proteins are split into three files.</p><p><i>CPT1_score_EVE_set.zip</i>: Proteins in the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>)</p><p><i>CPT1_score_no_EVE_set_1.zip</i> & <i>CPT1_score_no_EVE_set_2.zip</i>: Proteins not in the EVE set. Predictions for these proteins use imputed values for features depending on the EVE MSA.</p><p>The protein names are UniProt gene names.</p><p>We also provide source code to train CPT-1 model and reproduce results in the manuscript :</p><p><i>source_code.zip </i>(corresponds to GitHub repository songlab-cal/CPT version as of Jul 12, 2023)</p><p> </p><p><strong>Citation</strong></p><p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br>"Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p><p>*These authors contributed equally to this work.<br>†To whom correspondence should be addressed: <a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p><p>DOI: <a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p><p> </p>
Source code and data for AIMMD-TPS and PE estimate
<p>This repository contains the source code, simulation data, and analysis from the study by Lazzeri, Jung, Bolhuis, and Covino (paper submitted in 2023). The code, adaptable for any two-state molecular dynamics system, is paired with annotations to promote understanding. The dataset allows full replication of the paper's results and figures.</p>
Source code: Bilateral human laryngeal motor cortex in perceptual decision of lexical tone and voicing of consonant
<p>Source code (and data) for the paper <em>Bilateral Human Laryngeal Motor Cortex in Perceptual Decision of Lexical Tone and Voicing of Consonant</em>.</p> <p><a href="https://zenodo.org/api/files/d8e78f58-333f-45fb-954c-4e3a343fa5a1/Exp1_datacollection.zip">Exp1_datacollection.zip</a>:code for data collection in Experiment 1.</p> <p><a href="https://zenodo.org/api/files/d8e78f58-333f-45fb-954c-4e3a343fa5a1/Exp2_datacollection.zip">Exp2_datacollection.zip</a>: code for data collection in Experiment 2.</p> <p><a href="https://zenodo.org/api/files/d8e78f58-333f-45fb-954c-4e3a343fa5a1/codes_for_dataprocess.zip">codes_for_dataprocess.zip</a>:data processing code and source data.</p> <p>Exp1 and Exp2 collection codes are provided for reference, but for practice, due to copyright and software environment issues, please contact Baishen (liangbs@psych.ac.cn, liangbs95@gmail.com) for technical assistant. </p>
Dense vegetation hinders sediment transport towards saltmarsh interiors - Supporting data and source code (Part I: Pre-processing)
<p>This is Part I of the supporting data and source code for the paper entitled "Dense vegetation hinders sediment transport towards saltmarsh interiors", submitted to <em>Limnology and Oceanography Letters.</em> It contains all input and output files to generate the simulation grids.</p>
2023_ Datasets and R source code of "Allometry Bird Mitochondrial Bioenergetics"
<p>METHODOLOGICAL INFORMATION</p> <p>This dataset contains mitochondrial bioenergetic data and enzymatic activities of our studies from two tissues: skeletal and cardiac muscles, of 13 bird species ranging from 15 g to 160 kg.<br> Methodology: mitochondrial isolation, respiration, enzyme assays, measurement of body mass<br> All analyses were performed in R version 4.2.1 (R Core Team 2022), using phylogenetic comparative analyses.</p> <p><strong>## Description of the Data and file structure "2023_Data_Allometry Bird mitochondrial bioenergetics"</strong></p> <p>The file contains two sheets: one for the skeletal muscle data and the second for the cardiac muscle. <br> For each section you will find: the name of the species studied, the number of individuals, their body mass (in grams), mitochondrial flux measurements (oxygen consumption, ATP synthesis, ROS generation), ratios (RCR, Slope, Mitochondrial efficiency ATP/O...) and enzymatic activity measurements. </p> <p>Missing data correspond to individuals for whom we were unable to collect data (e.g. no heart samples, not enough tissue for analysis...)</p> <p><strong>## Phylogenetic tree " BirdTree_MCMCglmm "</strong></p> <p>The phylogenetic tree combining the 13 species studied was obtained from the BirdTree.org website (Rubolini et al., 2015).The tree source used was Hackett Sequenced Species: a set of 10 000 trees with 6670 OTUs each (Hackett et al., 2008). We performed 1000 simulations to create the most parsimonious tree. The avian tree was summarized using BEAST (v1.10.4, 2002-2018) to create a target tree usable in nexus format in R version 4.2.1 (R Core Team 2022). The parameters used were: burnin as a number of trees (100), maximum clade credibility tree as target tree type, and common ancestor heights.</p>
Data and codes: Who is calling? Optimising source identification from marmoset vocalisations with hierarchical machine learning classifiers
<p>Data and codes that accompany the article titled "Who is calling? Optimising source identification from marmoset vocalisations with hierarchical machine learning classifiers".</p>
Models and post-processing codes for paper "Quantitative stratigraphic analysis in a source-to-sink numerical framework"
<p>This package contains all the files required to reproduce the experiments in the manuscript: <strong>Quantitative stratigraphic analysis in a source-to-sink numerical framework</strong>.</p>
Relationships of climate, human activity, and fire history to spatiotemporal variation in annual fire probability across California: Source Code and Core Data
Open the record for dataset details and reuse information.
Source code for R tutorials and dataset for empirical case study on Malurus elegans (red-winged fairy wren)
Open the record for dataset details and reuse information.
Data and source code for: Recent adaptation in a threatened salmonid revealed by museum genomics
Open the record for dataset details and reuse information.
Source code and data from: Foraging personalities modify effects of habitat fragmentation on biodiversity
Open the record for dataset details and reuse information.
Data and source code for: ClinVar and HGMD genomic variant classification accuracy has improved over time, as measured by implied disease burden
Open the record for dataset details and reuse information.
Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants
Open the record for dataset details and reuse information.
FEHM source code modifications and executables for use with ocean-world gravity
Open the record for dataset details and reuse information.
CROP: linking code reviews to source code changes
<p>he Code Review Open Platform, a.k.a. CROP, is an open-source dataset of code review data. CROP collects code review information from open-source software systems and links this data to complete versions of the code base for each of these systems. CROP was first designed by <a href="https://mhepaixao.github.io/homepage/">Matheus Paixao</a> as part of his PhD thesis in the <a href="http://crest.cs.ucl.ac.uk/about/">CREST Centre</a> at University College London. <a href="http://www0.cs.ucl.ac.uk/staff/j.krinke/">Dr. Jens Krinke</a>, <a href="https://donggyun.com/">Donggyun Han</a> and <a href="http://www0.cs.ucl.ac.uk/staff/M.Harman/">Prof. Mark Harman</a> have also contributed for the first incarnation of CROP.</p> <p> </p> <p>CROP collects code review information from open-source software systems and links this data to complete versions of the code base for each of these systems.</p> <p> </p> <p>Given a certain software system, CROP contains code review data and versions of the code base for each revision ever submitted for review, including intermediary revisions before merging and revisions that were even abandoned by the system's developers. Each version of the system represents a complete snapshot of the system's code base, in a way that each revision of the system is fully buildable, compilable and testable.</p> <p> </p> <p>By leveraging the data contained in CROP, software engineering researchers and practitioners can perform empirical studies to assess how effective the code review process is for different aspects of software development. Since CROP provides complete snapshots of the software system, these experiments can be enhanced by using a wide range of approaches for static and dynamic analysis.</p> <p> </p> <p>Moreover, during code review, developers are constantly providing reasoning and rationale for the changes they make in the system, both when they submit code for review and when they inspect code from their peers. Thus, the data contained in CROP is a valuable source of knowledge regarding motivation for and explanation of software changes.</p> <p> </p> <p>For more information on the CROP dataset, including its structure, technical details, publication history and so on, please visit its official website in <a href="https://crop-repo.github.io/">crop-repo.github.io</a>.</p>
supplementary materials (source code and raw experimental results) for the paper Precision, Recall, and Sensitivity of Monitoring Partially Synchronous Distributed Programs
<p>This is the supplementary materials (source code and raw experimental results) for the paper Precision, Recall, and Sensitivity of Monitoring Partially Synchronous Distributed Systems.</p> <p>Please contact Duong Nguyen (<a href="mailto:nguye476@msu.edu">nguye476@msu.edu</a>) for any questions.</p> <p>Content:</p> <pre><code>TPDS-Precision-Recall-Upload | +- program: source code | +- aws-experiment: source code for running programs and monitors on Amazon Web Service (AWS) | +- interval-based | +- lib: relevant libraries for compiling code | +- wcp-process: source code of programs | +- wcp_process_compile_and_export_jar.sh: script for compiling program | +- simulations: source code for simulations | +- data: experiment and simulation results</code></pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.