Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,363

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

3,363 results for “Replication”

Learn how ShareScore rates datasets ↗
zenodo44/100

replicAnt - Plum2023 - Pose-Estimation Datasets and Trained Models

<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks&nbsp;used for semantic and instance segmentation&nbsp;experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data&nbsp;have been generated with the open-source photgrammetry platform scAnt&nbsp;<a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>.&nbsp;&nbsp;All synthetic data has been generated with the associated replicAnt project available from&nbsp;<a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created&nbsp;replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt&nbsp;places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Two pose-estimation datasets were procured. Both datasets used first instar <em>Sungaya&nbsp;nexpectata</em>&nbsp;(Zompro 1996) stick insects as a model species. Recordings from an evenly lit platform served as representative for controlled laboratory conditions; recordings from a hand-held phone camera served as approximate example for serendipitous recordings in the field.&nbsp;</p> <p>For the platform experiments, walking <em>S. inexpectata</em>&nbsp;were recorded using a calibrated array of five FLIR blackfly colour cameras (Blackfly S USB3, Teledyne FLIR LLC, Wilsonville, Oregon, U.S.), each equipped with 8 mm c-mount lenses (M0828-MPW3 8MM 6MP F2.8-16 C-MOUNT, CBC Co., Ltd., Tokyo, Japan). All videos were recorded with 55 fps, and at the sensors&rsquo; native resolution of 2048 px by 1536 px. The cameras were synchronised for simultaneous capture from five perspectives (top, front right and left, back right and left), allowing for time-resolved, 3D reconstruction of animal pose.<br> <br> The handheld footage was recorded in landscape orientation with a Huawei P20 (Huawei Technologies Co., Ltd., Shenzhen, China) in stabilised video mode: <em>S. inexpectata </em>were recorded walking across cluttered environments (hands, lab benches, PhD desks etc), resulting in frequent partial occlusions, magnification changes, and uneven lighting, so creating a more varied pose-estimation dataset.<br> <br> Representative frames were extracted from videos using DeepLabCut (DLC)-internal k-means clustering. 46 key points in 805 and 200 frames for the platform and handheld case, respectively, were subsequently hand-annotated using the DLC annotation GUI.</p> <p><strong>Synthetic data</strong></p> <p>We generated a synthetic dataset of 10,000 images at a resolution of 1500 by 1500 px, based on a 3D model of a first instar <em>S. inexpectata </em>specimen, generated with the <a href="https://peerj.com/articles/11155/"><em>scAnt</em>&nbsp;photogrammetry workflow</a>. Generating 10,000 samples took about three hours on a consumer-grade laptop (6 Core 4 GHz CPU, 16 GB RAM, RTX 2070 Super). We applied 70\% scale variation, and enforced hue, brightness, contrast, and saturation shifts, to generate 10 separate sub-datasets containing 1000 samples each, which were combined to form the full dataset.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College&rsquo;s President&rsquo;s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

replicAnt - Plum2023 - Segmentation Datasets and Trained Models

<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks&nbsp;used for semantic and instance segmentation&nbsp;experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data&nbsp;have been generated with the open-source photgrammetry platform scAnt&nbsp;<a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>.&nbsp;&nbsp;All synthetic data has been generated with the associated replicAnt project available from&nbsp;<a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created&nbsp;replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt&nbsp;places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Semantic and instance segmentation is used only rarely in non-human animals, partially due to the laborious process of curating sufficiently large annotated datasets. <em>replicAnt&nbsp;</em>can produce pixel-perfect segmentation maps with&nbsp;minimal manual effort. In order to assess the quality of the segmentations inferred by networks trained with these maps, semi-quantitative verification was conducted using a set of macro-photographs of <em>Leptoglossus zonatus</em>&nbsp;(Dallas, 1852) and <em>Leptoglossus phyllopus</em> (Linnaeus, 1767), provided by Prof. Christine Miller (University of Florida), and Royal Tyler (Bugwood.org. For further qualitative assessment of instance segmentation, we used laboratory footage, and field photographs of <em>Atta vollenweideri</em>&nbsp;provided by Prof. Flavio Roces. More extensive quantitative validation was infeasible, due to the considerable effort involved in hand-annotating larger datasets on a per-pixel basis.</p> <p><strong>Synthetic data</strong></p> <p>We generated two synthetic datasets from a single 3D scanned <em>Leptoglossus zonatus</em>&nbsp;(Dallas, 1852) specimen: one using the default pipeline, and one with additional plant assets, spawned by three dedicated scatterers. The plant assets were taken from the Quixel library and include 20 grass and 11 fern and shrub assets. Two dedicated grass scatterers were configured to spawn between 10,000 and 100,000 instances; the fern and shrub scatterer spawned between 500 to 10,000 instances. A total of 10,000 samples were generated for each sub dataset, leading to a combined dataset comprising 20,000 image render and ID passes. The addition of plant assets was necessary, as many of the macro-photographs also contained truncated plant stems or similar fragments, which networks trained on the default data struggled to distinguish from insect body segments. The ability to simply supplement the asset library underlines one of the main strengths of <em>replicAnt</em>: training data can be tailored to specific use cases with minimal effort.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College&rsquo;s President&rsquo;s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Replication package for Decomposition of Monolithic Applications into Microservices Architectures: A Systematic Review

<p><strong>Replication Package</strong></p> <p><strong>Title:</strong></p> <p>Replication package for Decomposition of Monolithic Applications into Microservices Architectures: A Systematic Review.</p> <p><strong>Authors:</strong></p> <p>Yalemisew Abgaz, Andrew McCarren, Peter Elger, David Solan, Neil Lapuz, Marin Bivol, Glenn Jackson, Murat Yilmaz, Jim Buckley, and Paul Clarke</p> <p><strong>Year</strong></p> <p>This replication package was initially generated in 2022 and following feedback from reviewers, it is revised in 2023.</p> <p>This package contains two files and three folders that provide additional insight for researchers who wish to replicate our work or who would like to expand the review in the future.</p> <p><strong>Files</strong></p> <ul> <li>The file contains a detailed description of the literature search outlining the steps and the results obtained.</li> <li>Readme.md: A readme file (this file).</li> </ul> <p><strong>Folders</strong></p> <ul> <li>A folder containing the search results and the refinement steps. It contains four files listing studies included in the refinement process (Refinement_Step_1 to Refinement_Step_4) and a master file combining all the steps in one file. The master file contains detailed information about how the refinement steps are executed and all the intermediate results following each refinement step. Users may explore by expanding the filters in Refinement Step 2 (J), Refinement Step 3 (M) and Refinement Step 4 &sect; columns in the master sheet. Readers can also directly go to the sheets that contain the selected studies in any of the refinement stages. A description of each file is also included in the Literature_Search_Strategy.pdf file.</li> <li>This folder contains the list of studies included in the snowballing process, including the last two refinement steps (Refinement_Step_5 and Refinement_Step_6). The snowballing master sheet contains studies extracted using the snowballing process and the data cleaning and filtering criteria used. The different sheets also contain the selected studies at each stage of the snowballing process.</li> </ul> <ul> <li>This folder contains the data extracted from the selected literature by employing the Systematic Review and Ground Theory. The two files included in this folder contain the data extraction template and the data extracted from the 35 selected studies including some intermediate notes.</li> </ul> <p>If you have further questions regarding the survey, feel free to contact us via e-mail. <a href="mailto:Yalemisewm.abgaz@dcu.ie">Yalemisewm.abgaz@dcu.ie</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Replication Package: Exploring the Relationship Between Personality Traits and User Feedback

<p>This is the replication package for the paper titled &#39;Automated User Feedback Analysis: Processing the Feedback Quantity and Quality&#39; accepted for the AffectRE23 workshop track at RE 2023.</p> <p>Full abstract:</p> <p>Previous research has studied the impact of developer personality in different software engineering scenarios, such as team dynamics and programming education. However, little is known about how user personality affect software engineering, particularly user-developer collaboration. Along this line, we present a preliminary study about the effect of personality traits on user feedback. 56&nbsp; university students provided feedback on different software features of an e-learning tool used in the course. They also filled out a questionnaire for the Five Factor Model (FFM) personality test. We observed some isolated effects of neuroticism on user feedback: most notably a significant correlation between neuroticism and feedback elaborateness; and between neuroticism and the rating of certain features. The results suggest that sensitivity to frustration and lower stress tolerance may negatively impact the feedback of users. This and possibly other personality characteristics should be considered when leveraging feedback analytics for software&nbsp; requirements engineering.</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Replication package for "Deliberate Surrender? The Impact of Interwar Indian Protection"

<p>This is a replication package for&nbsp;Deliberate Surrender? The Impact of Interwar Indian Protection, by&nbsp;Vellore Arthi, Markus Lampe, Ashwin Nair, Kevin Hjortsh&oslash;j O&rsquo;Rourke.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Experiment replication - Plattes 1639 (Johns Hopkins University, Baltimore, USA, 02/28/2023 - 03/01/2023)

<p>Replication in laboratory of an experiment reported by&nbsp;Gabriel Plattes,&nbsp;<em>A Discovery of Subterraneall Treasure</em>&nbsp;(1639):</p> <p>&ldquo;Now for a plaine demonstration [of how mineral ores are generated], let this experiment following be tryed, and I make no question, but that it will satisfie every one that hath an inquisitive disposition. Let there bee had a great retort of glasse, and let&nbsp; the same be halfe filled&nbsp; with brimstone, sea-coale, and as many bituminous and sulphurious subterraneall substances as can bee gotten: then fill the necke thereof halfe full with the most free earth from stones that can be found, but thrust it not in too hard, then let it bee luted, and set in an open furnace to distill with a temperate fire, which may onely kindle the said substances, and if you worke exquisitely, you shall finde the said earth petrified, and turned into a stone: you shall also finde cracks and chinkes in it, filled with the most tenacious, clammy, and viscous parts of the said vapours, which ascended from the subterraneall combustible substances. Wereby it appeareth that the same thing is done by Nature, and that the rocks and craggy mountaines are caused by the vapours of bituminous and sulphurious substances kindled in the bowells of the earth, of which there bee divers so well knowne, that they neede not bee heere mentioned [...]&rdquo; (pp. 5-7). This experiment is also mentioned in Emerton 1984, p. 249. See also Debus 1961.</p> <p>&ldquo;Take a peece of the blacke fat earth, which is usually digged up in the west countrey, where there are such a multitude of firre trees covered therewith, and which the people use to cut in the forme of bricks, and to drye them, &amp; so to burne them in stead of coales; use this substance as you did the other earth in the beginning of the booke, to find out the natural&nbsp; cause of rocks,&nbsp; stones, and mettalls, and let it receive the vapours of the cumbustible substances, and you shall find this fat earth hardned into a plaine coale; even as you found the other leane earth hardned into a stone&rdquo; (p. 50).</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Experiment replication - Plattes 1639 (Johns Hopkins University, Baltimore, USA, 6/15/2022)

<p>Replication in laboratory of an experiment reported by&nbsp;Gabriel Plattes, <em>A Discovery of Subterraneall Treasure</em> (1639):</p> <p>&ldquo;Now for a plaine demonstration [of how mineral ores are generated], let this experiment following be tryed, and I make no question, but that it will satisfie every one that hath an inquisitive disposition. Let there bee had a great retort of glasse, and let&nbsp; the same be halfe filled&nbsp; with brimstone, sea-coale, and as many bituminous and sulphurious subterraneall substances as can bee gotten: then fill the necke thereof halfe full with the most free earth from stones that can be found, but thrust it not in too hard, then let it bee luted, and set in an open furnace to distill with a temperate fire, which may onely kindle the said substances, and if you worke exquisitely, you shall finde the said earth petrified, and turned into a stone: you shall also finde cracks and chinkes in it, filled with the most tenacious, clammy, and viscous parts of the said vapours, which ascended from the subterraneall combustible substances. Wereby it appeareth that the same thing is done by Nature, and that the rocks and craggy mountaines are caused by the vapours of bituminous and sulphurious substances kindled in the bowells of the earth, of which there bee divers so well knowne, that they neede not bee heere mentioned [...]&rdquo; (pp. 5-7). This experiment is also mentioned in Emerton 1984, p. 249. See also Debus 1961.</p> <p>&ldquo;Take a peece of the blacke fat earth, which is usually digged up in the west countrey, where there are such a multitude of firre trees covered therewith, and which the people use to cut in the forme of bricks, and to drye them, &amp; so to burne them in stead of coales; use this substance as you did the other earth in the beginning of the booke, to find out the natural&nbsp; cause of rocks,&nbsp; stones, and mettalls, and let it receive the vapours of the cumbustible substances, and you shall find this fat earth hardned into a plaine coale; even as you found the other leane earth hardned into a stone&rdquo; (p. 50).</p>

opencc-by-4.0Oct 2023View details →
edi44/100

Bird Abundances at the Hubbard Brook Experimental Forest (1969-present) and on three replicate plots (1986-2000) in the White Mountain National Forest (Reformatted to a Darwin Core Archive)

This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/355/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-hbr/81/7. The abstract below was extracted from the Level 0 data package and is included for context: Bird abundances have been determined from timed censuses, territory maps and nest locations at the Hubbard Brook Experimental Forest from 1969 to the present. This data set includes counts of the number of adult birds (males and females) per 10 ha at HBEF (1969 - present) and on three additional plots within the White Mountain National Forest (1986 - 2000). These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.

openCC (other)Aug 2021View details →
edi44/100

Soil physical and chemical properties based on genetic horizon from 4 replicate pits placed around the replicate LTER control plots sampled in 1988 and 1989.

Dataset contains the following soil properties for each genetic horizon - site, Soil pit, upper and lower boundary (cm), Mg meq/100gm, Ca meq/100gm, K meq/100gm, CEC meq/100gm, pH, %C, %sand, %silt, %clay, Total %N, Total %P, % organic matter, Mn meq/100gm, Available-P ppm, %CO3, bulk density gm/cm3, Volume wt gm/m2.

openOpenFeb 1998View details →
OpenNeuro40/100

Investigating the visual number form area: A replication study

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo40/100

spanichella/RP_EMSE_MCR_2019 v.1.0.1 Second release of of the replication Package for the paper "An Empirical Investigation of Relevant Changes and Automation Needs in Modern Code Review".

<p>Replication Package for the paper &quot;An Empirical Investigation of Relevant Changes and Automation Needs in Modern Code Review&quot;</p> <p>Structure</p> <pre><code>project_raw_data/ gerrit_review_comments.csv gerrit_review_changes.csv survey_raw_data/ google_forms_survey.pdf google_forms_survey.csv RQ1_taxonomy_mcr/ RQ1_inception_phase/ initial_taxonomy.pdf intermediate_taxonomy.pdf RQ1_definition_phase/ Q1.2_evaluation_survey.csv cram_classified.csv cram.pdf RQ2_automation_needs/ Q2.1-Q2.5_evaluation_survey.xlsx Q2.6-Q2.7_evaluation_survey.xlsx Q2.1-Q2.7_question_index.csv Now "RQ3_automated_support/" contain the results concerning RQ2.1 in the paper. content explained in the README.md file located in "RQ3_automated_support/README.md" </code></pre> <p>Contents of the Replication Package</p> <p><strong>project-raw-data/</strong>&nbsp;contains the data used for the creation of our taxonomies, it includes information about the ten open-source projects.</p> <ul> <li><code>gerrit_review_comments.csv</code>&nbsp;- information about all in-line review comments used for this paper</li> <li><code>gerrit_review_changes.csv</code>&nbsp;- information about all patches analyzed that contain the in-line comments</li> </ul> <p><strong>survey_raw_data/</strong>&nbsp;contains information about the survey conducted for the paper.</p> <ul> <li><code>google_forms_survey.pdf</code>&nbsp;- the distributed&nbsp;<em>Google Forms</em>&nbsp;of our survey</li> <li><code>google_forms_survey.csv</code>&nbsp;- all survey answers obtained from 52 survey participants</li> </ul> <p><strong>RQ1_taxonomy_mcr/</strong>&nbsp;contains information and data about the elicited taxonomies in our paper (RQ1).</p> <ul> <li><strong>RQ1_inception_phase/</strong> <ul> <li><code>initial_taxonomy.pdf</code>&nbsp;- initial taxonomy obtained in the inception phase of our paper</li> <li><code>intermediate_taxonomy.pdf</code>&nbsp;- intermediate taxonomy after integrating and merging the initial taxonomy with the one by Beller&nbsp;<em>et al</em>&nbsp;[1]</li> </ul> </li> <li><strong>RQ1_definition_phase/</strong> <ul> <li><code>Q1.2_evaluation_survey.csv</code>&nbsp;- relevant survey feedback with additional taxonomy categories integrated into&nbsp;<em>CRAM</em></li> <li><code>cram_classified.csv</code>&nbsp;- classified review comments (<code>gerrit_review_comments.csv</code>) into&nbsp;<em>CRAM</em></li> <li><code>cram.pdf</code>&nbsp;-&nbsp;<em>CRAM</em>&nbsp;taxonomy</li> <li><code>cram_classified_with_frequency_information2020.xls</code>&nbsp;- it contain the information used to compute the frequency of CRAM changes, derived by the analysis of the 211 commits</li> </ul> </li> </ul> <p><strong>RQ2_automation_needs/</strong>&nbsp;contains the encoded evaluation of the survey question Q2.1-2.7 for RQ2</p> <ul> <li><code>Q2.1-Q2.5_evaluation_survey.xlsx</code>&nbsp;- the encoded evaluation of the survey questions Q2.1-Q2.5 (used for the&nbsp;<em>Automation Needs</em>&nbsp;Section in the paper) and contains the following sheets: <ul> <li><strong>all Findings</strong>: Very detailed findings matrix distilled from all answers concerning possible automated solutions or general possibilities to achieve automation in MCR. Every feedback was analyzed and decomposed into single findings. These findings are grouped, into categories of our Taxonomy of Code Changes in MCR (CRAM). Red represent in the feedback-text where the corresponding category was distilled from.</li> <li><strong>unique Findings</strong>: As one participant could mention the same categories/solutions in multiple feedbacks, the following matrix is cleaned of any duplication of participant answers. Multiple feedbacks containing the same information by one participant were removed, leaving only distinct occurances.</li> <li><strong>aggregated by Solution</strong>: Aggregated feeback clustered into abstracted solutions and the number of times participants mentioned the solution.</li> <li><strong>aggregated by Taxonomy</strong>: Aggregated feeback grouped by low-level categoried in CRAM.</li> </ul> </li> <li><code>Q2.6-Q2.7_evaluation_survey.xlsx</code>&nbsp;- the encoded evaluation of the survey questions Q2.6-Q2.7 (used for the&nbsp;<em>Automation Needs</em>&nbsp;Section in the paper) and contains the following sheets: <ul> <li><strong>all Findings</strong>: Very detailed findings matrix distilled from all answers concerning possible techniques, approaches and data to achieve automation in MCR. Every feedback was analyzed and decomposed into single findings. These findings are grouped, into categories of our Taxonomy of Code Changes in MCR (CRAM). Red represent in the feedback-text where the corresponding category was distilled from.</li> <li><strong>unique Findings</strong>: As one participant could mention the same categories/solutions in multiple feedbacks, the following matrix is cleaned of any duplication of participant answers. Multiple feedbacks containing the same information by one participant were removed, leaving only distinct occurances.</li> <li><strong>aggregated by low-level taxonomy</strong>: Aggregated mentionings of approaches/data by developers in the survey grouped by low-level taxonomy category.</li> <li><strong>aggregated by high-level taxonomy</strong>: Aggregated mentionings of approaches/data by developers in the survey grouped by high-level taxonomy category.</li> </ul> </li> <li><code>Q2.1-Q2.7_question_index.csv</code>&nbsp;- table of IDs given to each participant-question pair for Q2.1-Q2.7 in order to trace back the feeback.</li> <li><code>cram_survey-with_criticality_and_feasibility2020.xls</code>&nbsp;and&nbsp;<code>cram_survey-with_relevance_and_completeness_information2020.xls</code>: they contain we results of the survey, involving 14 additional participants (12 developers and 2 researchers), not involved in the aforementioned survey, and performed to qualitatively assess the relevance and completeness of the identified MCR change types as well as assess how critical and feasible to implement are some of the identified techniques to support MCR activities.</li> </ul> <p><strong>RQ2_1_automated_support/</strong>&nbsp; (or&nbsp;<strong>RQ3_automated_support/</strong>&nbsp;)- content explained in the README.md file located in &quot;RP_EMSE_MCR_2019/tree/master/EMSE_MCR_2019/RQ3_automated_support/README.md&quot;</p> <p>References</p> <p>[1] Moritz Beller, Alberto Bacchelli, Andy Zaidman, and Elmar Juergens. 2014. Modern code reviews in open-source projects: which problems do they fix?. In Proceedings of the 11th Working Conference on Mining Software Repositories (MSR 2014). ACM, New York, NY, USA, 202-211. DOI:&nbsp;<a href="http://dx.doi.org/10.1145/2597073.2597082">http://dx.doi.org/10.1145/2597073.2597082</a></p>

openother-openFeb 2020View details →
zenodo40/100

Darwin's naturalization conundrum can be explained by spatial scale: R replication code

<p>1) Code for extracting climatic data:<br> 1_extract_env.R<br> 2_extract_env_county.R<br> 3_relatePlotCounty.R</p> <p>2) Code for generating composite phylogenies:<br> 1_making_sunplin_trees.R</p> <p>3) Code for relatedness analyses:<br> DNH_calculations_PHY_obs.cluster.R<br> DNH_calculations_PHY_random.cluster.R<br> DNH_calculations_TAX.cluster_obs.R<br> DNH_calculations_TAX.cluster_random.R</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Replication package of A benchmark-based evaluation of search-based crash reproduction

<p>Release of the reproduction package of Soltani, M., Derakhshanfar, P., Devroey, X. and van Deursen, A. (2020). A benchmark-based evaluation of search-based crash reproduction. In Empirical Software Engineering. 25, 1 (Jan. 2020), pp. 96&ndash;138.</p>

openother-openApr 2020View details →
zenodo40/100

Mass spectrometry output SILAC labelled (F/Y) biological replicate 1 - anti-HLA-A29 antibody DK1G8

<p>Mass spectrometry output from SILAC labelled (F/Y) ERAP2 wildetype versus (CRISPR Cas9-mediated) ERAP2-KO lymphoblastoid cell line from a Birdshot Uveitis patient (ERAP1 hap10/10 ERAP2 hapA/A). Peptides were eluted from immuno-purifications with anti-HLA-A29 antibody DK1G8. The dataset was used for subsequent filtering and differential expression analysis by <em>limma</em>.&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo40/100

A Replication Package of Learning Features that Predict Developer Responses for iOS App Store Reviews

<p>This replication package contains the dataset and script used in our&nbsp;paper &quot;<em>Learning Features that Predict Developer Responses for iOS App Store Reviews.</em>&quot; The paper has been accepted at the ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), 2020. For further modification and versioning of the dataset (as well as the preprint) please go to&nbsp;<a href="https://github.com/Kamonphop/ESEM20-Replication">https://github.com/Kamonphop/ESEM20-Replication</a></p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Replication files for "Building social cohesion between Christians and Muslims through soccer in post-ISIS Iraq"

<p>Replication files for main analyses, and supplementary analyses (comparison group, Muslim player attitudes, fan attitudes, match-level data) for:</p> <p><strong>Mousa, Salma. </strong>&quot;<a href="https://science.sciencemag.org/content/369/6505/866">Building social cohesion between Christians and Muslims through soccer in Post-ISIS Iraq</a>.<strong>&quot; <em>Science</em>. </strong>Vol. 369, Issue 6505, pp. 866-870. DOI: 10.1126/science.abb3153</p> <p>Each R file describes the needed datasets at the top of the script.&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Replication package of "Revisiting Test Smells in Automatically Generated Tests: Limitations, Pitfalls, and Opportunities"

<p><strong>Abstract:</strong><br> Test smells attempt to capture design issues in test code that reduce their maintainability. Previous work found such smells to be highly common in automatically generated test-cases, but based this result on specific static detection rules; although these are based on the original definition of &ldquo;test smells&rdquo;, a recent empirical study showed that developers perceive these as overly strict and non-representative of the maintainability and quality of test suites. This leads us to investigate how&nbsp;effective&nbsp;such test smell detection tools are on automatically generated test suites. In this paper, we build a dataset of 2,340 test cases automatically generated by EVOSUITE for 100 Java classes. We performed a multi-stage, cross-validated&nbsp;manual analysis to identify six types of test smells and label their instances. We benchmark the performance of two test smell detection tools: one widely used in prior work, and one recently introduced with the express goal to match developer perceptions of test smells. Our results show that these test smell detection strategies poorly characterized the issues in automatically generated test suites; the older tool&rsquo;s detection strategies, especially, misclassified over 70% of test smells, both&nbsp;missing&nbsp;real instances (false negatives) and marking many smell-free&nbsp;tests as smelly (false positives). We identify common patterns in these tests that can be used to&nbsp;improve&nbsp;the tools, refine and update the definition of&nbsp;certain&nbsp;test smells, and&nbsp;highlight&nbsp;as of yet uncharacterized issues. Our findings suggest the need for (i) more appropriate metrics to match development practice; and (ii) more accurate detection strategies, to be evaluated primarily in industrial contexts.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Replication Data for: "Parabolic Diamond Scanning Probes for Single-Spin Magnetic Field Imaging

<p>Data repository for: <strong>Parabolic Diamond Scanning Probes for Single-Spin Magnetic Field Imaging</strong></p> <ul> <li><em>DataDescription.pdf</em><strong><em>:&nbsp;</em></strong>describes the uploaded data</li> <li><em>Data (folder):&nbsp;</em>folder containing&nbsp;<em>data.xlsx</em>, which summarizes all the data plotted in the paper as well as additional imaging and simulation data sets</li> <li><em>Code (folder):&nbsp;</em>&nbsp;contains Matlab code for converting and plotting certain data sets</li> </ul>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Replication R code and data for "Geometric morphometric investigation of craniofacial morphological change in domesticated silver foxes"

<p>This repository holds various files and R code used in the publication of the manuscript entitled &quot;Geometric morphometric investigation of craniofacial morphological change in domesticated silver foxes&quot;.</p> <p>Data files include: The 3D landmark coordinates of each individual specimen (Fox_data_Morphologika.txt), the linear measurement data associated with those foxes (fox_linear_volume_data.csv), and replication data measurements.&nbsp;</p> <p>The following files include the R code used to perform the analyses contained within the paper:</p> <p>1_Procrustes_analysis - details the Geometric morphometrics analyses performed</p> <p>2_linear_models - details the model specification for the GLS models employed in the paper</p> <p>3_graph_code - contains R script for the creation of the graphs displayed&nbsp;in the paper</p> <p>4_repeatability_script - contains R code that details the statistical calculations made with the repeatability measurements indicated above.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Replication Package: Assessing time-based and range-based strategies for commit assignment to releases

<p><strong>Abstract:</strong></p> <p>Release is a ubiquitous concept in software development, referring to grouping multiple independent changes into a deliverable piece of software. Mining releases can help developers understand the software evolution at coarse grain, identify which features were delivered or bugs were fixed, and pinpoint who contributed on a given release. A typical initial step of release mining consists of identifying which commits compose a given release. We could find two main strategies used in the literature to perform this task: time-based and range-based. Some release mining works recognize that those strategies are subject to misclassifications but do not quantify the impact of such a threat. This paper analyzed 13,419 releases and 1,414,997 commits from 100 relevant open source projects hosted at GitHub to assess both strategies in terms of precision and recall. We observed that, in general, the range-based strategy has superior results than the time-based strategy. Nevertheless, even when the range-based strategy is in place, some releases still show misclassifications. Thus, our paper also discusses some situations in which each strategy degrades, potentially leading to bias on the mining results if not adequately known and avoided.</p> <p><strong>Instructions:</strong></p> <p>Visit <a href="https://github.com/gems-uff/release-mining">https://github.com/gems-uff/release-mining</a> for instructions about how to use this dataset.</p> <p><strong>Files:</strong></p> <ul> <li>The <em>repos.tgz</em> contains our project <em>corpus </em>comprising 1,414,997&nbsp;releases&nbsp;from&nbsp;100&nbsp;relevant&nbsp;open&nbsp;source&nbsp;projects.</li> <li>The <em>repos.sha1</em> contains the sha1 checksum of <em>repos.tgz</em></li> </ul> <p><strong>Disclaimer:</strong></p> <p>This replication package contains the source code of 100 relevant open source projects. Its purpose is to enable the replication of the study conducted in the paper &quot;Assessing time-based and range-based strategies for commit assignment to releases.&quot;&nbsp;&nbsp;</p> <p>It is essential to check each project license before using the source code or any attached file for any other purposes besides replicating the study.</p> <p>&nbsp;</p>

openmit-licenseJan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record