Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.9.0
Dataset results
19 results for “Breaking changes”
Research compendium for 'Refitting the Context: A Reconsideration of Cultural Change among Early Homo sapiens at Fumane Cave through Blade Break Connections, Spatial Taphonomy, and Lithic Technology'
<div> <h3>Compendium DOI:</h3> <p><a href="../doi/10.5281/zenodo.10965413">https://zenodo.org/doi/10.5281/zenodo.10965413</a> </p> </div> <p>The content available at the above provided URL will reproduce the results as documented in the publication. Instead, the files hosted at <a href="https://github.com/ArmandoFalcucci/Refitting-The-Context">https://github.com/ArmandoFalcucci/Refitting-The-Context</a> represent the developmental versions and might have undergone modifications since the paper's publication.</p> <div> <h3>Maintainer of this repository:</h3> </div> <p>Armando Falcucci (<a href="mailto:armando.falcucci@uni-tuebingen.de">armando.falcucci@uni-tuebingen.de</a>)</p> <div> <h3>Published paper:</h3> </div> <p>Armando Falcucci, Domenico Giusti, Filippo Zangrossi, Matteo De Lorenzi, Letizia Ceregatti, Marco Peresani. Refitting the Context: Revisiting the Aurignacian sequence at Fumane Cave through blade fragment connections, spatial taphonomy, and lithic technology. <em>Journal of Paleolithic Archaeology</em> (2024). DOI: <a href="https://doi.org/10.1007/s41982-024-00203-0" rel="nofollow">10.1007/s41982-024-00203-0</a></p> <div> <h3>Abstract:</h3> </div> <p>High-resolution stratigraphic frameworks are crucial for unraveling the biocultural processes behind the dispersals of Homo sapiens across Europe. Detailed technological studies of lithic assemblages retrieved from multi-stratified sequences allow archaeologists to precisely model the chrono-cultural dynamics of the early Upper Paleolithic. However, it is of paramount importance to verify the integrity of these assemblages before building explanatory models of cultural change. In this study, multiple lines of evidence suggest that the stratigraphic sequence of Fumane Cave in northeastern Italy experienced minor post-depositional reworking, establishing it as a pivotal site for exploring the earliest stages of the Aurignacian. By conducting a systematic search for break connections between blade fragments and applying spatial analysis techniques, we identified three well-preserved areas of the excavation containing assemblages suitable for renewed archaeological investigations. Subsequent technological analyses, incorporating attribute analysis, reduction intensity, and multivariate statistics, have allowed us to discern the spatial organization of the site during the formation of the Protoaurignacian palimpsest A2–A1. Moreover, diachronic comparisons between three successive stratigraphic units prompted us to reject the hypothesis of techno-cultural continuity of the Protoaurignacian in northeastern Italy after the onset of the Heinrich Event 4. Based on the variability of the lithic and osseous artifacts, the most recent assemblage analyzed, D3b alpha, is now ascribed to the Early Aurignacian, aligning the evidence from Fumane with the current understanding of the development of the Aurignacian across Europe. Overall, this study demonstrates the high effectiveness of the break connection method when combined with detailed spatial analysis and lithic technology, providing a methodological tool particularly amenable to be applied to sites excavated in the past with varying degrees of recording accuracy.</p> <div> <h3>Keywords:</h3> </div> <p>Protoaurignacian; Early Aurignacian; Lithics; Refittings; Assemblage integrity; Spatial analysis; Italy</p> <div> <h3>Overview of contents and how to reproduce:</h3> </div> <p>Within this repository, various folders house data (<code>data</code>), code (<code>script</code>), and output files (<code>output</code>) pertinent to the paper. The data folder encompasses the blank and core datasets from the Aurignacian of Fumane Cave and the dataset of the blade fragment connection study. To replicate the results, download the entire repository and employ <code>Refitting-The-Context.Rproj</code> and open the folder <code>script</code>. For ensuring reproducibility, the <code>renv</code> package (v. 1.0.3) was utilized, following the procedures detailed in its vignette. All analyses and visualizations in the paper were conducted using R 4.3.1 on Microsoft Windows 10.0.19045 (64-bit). As the necessary packages are available in the <code>renv</code> folder, they are not explicitly listed here.</p> <div> <h3>Licenses:</h3> </div> <p>Code: <strong>MIT</strong> <a href="http://opensource.org/licenses/MIT" rel="nofollow">http://opensource.org/licenses/MIT</a>, copyright holder: Armando Falcucci (2024).</p> <p>Data and intellectual work: <strong>Creative Commons Attribution 4.0 International License</strong> (<a href="http://creativecommons.org/licenses/by/4.0/" rel="nofollow">http://creativecommons.org/licenses/by/4.0/</a>), copyright holder: the authors (2024).</p>
Breaking Bad? Semantic Versioning and Impact of Breaking Changes in Maven Central (Dataset)
<p>The content presented in this repository accompanies the paper "Breaking Bad? Semantic Versioning and Impact of Breaking Changes in Maven Central" authored by Lina Ochoa, Thomas Degueule, Jean-Rémy Falleri, and Jurgen Vinju. The paper was submitted and accepted in the Journal of Empirical Software Engineering (EMSE'21). This study is an external and differentiated replication study of the paper <a href="https://jstvssr.github.io/assets/pdf/semantic-versioning-maven.pdf">"Semantic Versioning and Impact of Breaking Changes in the Maven Repository"</a> presented by Steven Raemaekers, Arie van Deursen, and Joost Visser.</p> <p><strong>Content</strong></p> <ul> <li><strong>README.md: </strong>document with the main description to start exploring the bundle.</li> <li><strong>data.zip: </strong>contains the datasets used within the study. These datasets must be used to get the same results like the ones presented in the article.</li> <li><strong>maven-api-dataset.zip:</strong> contains the code used to generate the datasets and to analyse the obtained results. Check the README.md file within this bundle for more information.</li> </ul> <p><strong>Relevant Links</strong></p> <ul> <li><strong>maven-api-dataset repository:</strong> <a href="https://github.com/tdegueul/maven-api-dataset">https://github.com/tdegueul/maven-api-dataset</a></li> <li><strong>maracas repository:</strong> <a href="https://github.com/crossminer/maracas">https://github.com/crossminer/maracas</a></li> <li><strong>Companion webpage:</strong> <a href="https://crossminer.github.io/maracas/2021/08/16/emse21/">https://crossminer.github.io/maracas/2021/08/16/emse21/</a></li> </ul>
Break the Code? Breaking Changes and Their Impact on Software Evolution (Artefacts)
<p>The artefacts included in this repository accompany the thesis "Break the Code? Breaking Changes and Their Impact on Software Evolution" authored by Lina María Ochoa Venegas and supervised by prof.dr. Jurgen Vinju, prof.dr. Mark van den Brand, and dr.Thomas Degueule. The thesis was developed at Eindhoven University of Technology (TU/e) in Eindhoven, The Netherlands and Centrum Wiskunde & Informatica (CWI) in Amsterdam, The Netherlands. It was submitted to revision in 2022 and defended in 2023.</p> <p> </p> <p><strong>Relevant Links</strong></p> <ul> <li><strong>Maracas:</strong> https://github.com/alien-tools/maracas</li> <li><strong>BreakBot: </strong>https://github.com/alien-tools/breakbot</li> </ul>
North China Record-Breaking Rainfall of July 2021 tied to the Arctic Sea-ice change
<p><strong>1、General Introduction:</strong></p> <p>This dataset includes the model simulation data used in the article going to be submitted to Geophysical Research Letters, named “North China Record-Breaking Rainfall of July 2021 tied to the Arctic Sea-ice change”. The Community Atmosphere Model version 5 (CAM5; Neale et al., 2012) in the Community Earth System Model, version 1.2.1 (CESM1.2.1) with a 1.9<sup>◦</sup> ×2.5<sup>◦ </sup>finite volume grid and 30 hybrid sigma pressure levels is utilized to conduct atmospheric model experiments. To validate the impact of the Arctic sea-ice, the Community Atmosphere Model version 5 (CAM5; Neale et al., 2012) in the Community Earth System Model, version 1.2.1 (CESM1.2.1) with a 1.9◦ ×2.5◦ finite volume grid and 30 hybrid sigma pressure levels is adopted. Two experiments are designed. One is called the control experiment (CTL). In this experiment, the model is forced by the seasonal-varying climatology of SST and SIC as the lower boundary forcing. Another is called the sensitivity experiment (SEN). This experiment is the same as the CTL except that SIC and SST over the Barents-East Siberian Sea region (50°–165°E, 70°–80°N) in June and July are set to be the observational values in 2021. Both the CTL and SEN experiments are integrated for 35 years including a 10-year spin-up, whilst the last 25-year ensemble mean is analyzed in the article. The differences between the SEN and CTL experiments (SEN-CTL) reflect the atmospheric responses to the reduced SIC and increased SST over the Barents-East Siberian Sea region in June-July 2021.</p> <p><strong>2、Description of this dataset:</strong></p> <p>CTL represents the control experiment while SEN represents the sensitive experiment. OMEGA, PRECC, PRECL, T, U, V, Z3, and TS denote the vertical velocity, convective precipitation rate, large-scale (stable) precipitation rate, temperature, zonal wind, meridional wind, geopotential height, and surface temperature respectively. All files contain a time coverage from 1 to 35 model-year with a monthly time resolution.</p> <p>The post-processing from 30 hybrid sigma pressure levels to 17 pressure levels has been conducted on the raw data of the model simulations for the variables of OMEGA, T, U, V, and Z3.</p>
Automatically fixing dependency breaking changes
<h1>Replication Data for Automated Dependency Update Experiments</h1> <h2>Dataset Description</h2> <p>This repository contains a comprehensive collection of data from experiments on automated dependency updates using Large Language Models (LLMs). The dataset consists of two main components:</p> <ol> <li><strong>Execution Traces</strong>: Detailed logs of LLM interactions and repair attempts, captured using OpenTelemetry and OpenInference (<a href="https://github.com/Arize-ai/openinference">https://github.com/Arize-ai/openinference</a>). These traces provide insights into the behavior and performance of various LLMs in addressing breaking changes caused by dependency updates in Java projects.</li> <li><strong>Docker Images</strong>: Pre-built Docker images containing the state of Java projects after successful repair attempts. These images allow for direct inspection and verification of the changes made by our automated repair system.</li> </ol> <p>This dataset enables analysis and replication of our study on automated dependency updates, offering both low-level interaction data and high-level repair outcomes.</p> <h2>Data Components</h2> <h3>1. Execution Traces</h3> <ul> <li><strong>Format</strong>: JSON Lines, persisted via a custom OpenTelemetry exporter</li> <li><strong>Content</strong>: Spans capturing various aspects of LLM interactions, including: <ul> <li>Input processing</li> <li>LLM query generation</li> <li>LLM response parsing</li> <li>Code modification attempts</li> <li>Compilation and test execution results</li> </ul> </li> <li><strong>Purpose</strong>: Enables detailed analysis of LLM decision-making processes and performance metrics</li> </ul> <h3>2. Docker Images</h3> <ul> <li><strong>Format</strong>: Compressed Docker image files (.tar.gz) in a .zip</li> <li><strong>Content</strong>: Maven project states after successful repair attempts with patch(es) applied and seperated in the second layer of each image.</li> <li><strong>Purpose</strong>: Allows direct inspection and verification of code changes made during repairs and the successful maven outcome</li> </ul> <h2>Research Context and Data Usage</h2> <p>This dataset corresponds to experiments described in our study investigating the effectiveness of zero-shot prompting and agentic approaches in automating dependency updates. It provides valuable insights into:</p> <ul> <li>Performance variations across different LLMs</li> <li>Influence of various factors on repair success rates</li> <li>Specific code changes made during successful repairs</li> </ul> <p>This dataset can be used for:</p> <ol> <li>Replicating the experimental results presented in the associated study</li> <li>Conducting further analysis on LLM behavior in software engineering tasks</li> <li>Developing and benchmarking new approaches for automated dependency updates</li> <li>Investigating the decision-making processes of LLMs in code modification tasks</li> <li>Verifying and inspecting successful repair attempts through Docker images</li> </ol> <p> </p> <div> <h2>Repair Replication with Docker</h2> To replicate specific repairs using our pre-built Docker images:<br> <div>1. Ensure you have Docker installed on your system.</div> <br>2. Use the included docker-images.zip, which first has to be unzipped <div> <pre><code>unzip docker-images.zip</code></pre> </div> <div>after which the images can be loaded as follows:</div> <div> <pre><code>docker load < image_name.tar.gz</code></pre> </div> <br> <div>3. Run the Docker container:</div> <pre><code>docker run -it [image-name]</code></pre> <br> <div>These Docker images contain the state of projects after repair attempts, allowing for easy inspection and verification of the changes made by our automated repair system.</div> </div>
Replication Package For An Extended Study of Syntactic Breaking Changes in the Wild
<p>This is the replication package associated with the paper titled 'An Extended Study of Syntactic Breaking Changes in the Wild' published under the Empirical Software Engineering journal.<br>Modern software applications rely heavily on the usage of libraries, which provide reusable functionality, to accelerate the development process. As libraries evolve and release new versions, the software systems that depend on those libraries (the clients) should update their dependencies to use these new versions as the new release could, for example, include critical fixes for security vulnerabilities. However, updating is not always a smooth process, as it can result in software failures in the clients if the new version<br>includes breaking changes. Yet, there is little research on how these breaking changes impact the client projects in the wild. <br>To identify if changes between two library versions cause breaking changes at the client end, we perform an empirical study on Java projects built using Maven. For the analysis, we used 18,415 Maven artifacts, which declared 142,355 direct dependencies, of which 71.60% were not up-to-date. We updated these dependencies and found<br>that 11.58% of the dependency updates contain breaking changes that impact the client. We further analyzed these changes in the library which impact the client projects and examine if libraries have adhered to the semantic versioning scheme when introducing breaking changes in their releases. Our results show that changes in transitive dependencies were a major factor in introducing breaking changes during dependency updates and almost half of the detected client impacting breaking changes violate the semantic versioning scheme by introducing breaking changes in non-Major upda</p>
FOCI model output used in the study by Ivanciu et al. - On the ridging of the South Atlantic Anticyclone over South Africa: the impact of Rossby wave breaking and of climate change
<p>This dataset contains the model output used in the analysis presented in the study by Ivanciu et al., 2022 - On the ridging of the South Atlantic Anticyclone over South Africa: the impact of Rossby wave breaking and of climate change. Four ensembles of three simulations each were performed with the global coupled climate model FOCI (Flexible Ocean and Climate Infrastructure, Matthes et al., 2020). Details about the ensembles can be found in the above-mentioned publication. The files containing "past" in their name belong to the ensemble "PAST", the files containing "future" in their name belong to the ensemble "FUTURE", the files containing "future_GHG" in their name belong to the ensemble "GHG" and the files containing "future_Ozone" in their name belong to the ensemble "OZONE" from the publication.</p>
Replication Package for Understanding the Impact of APIs Behavioral Breaking Changes on Client Applications
<p>This repository contains the replication package for the paper Understanding the Impact of APIs Behavioral Breaking Changes on Client Applications.<br>This paper will be published in the Proceedings of the ACM on Software Engineering journal. The replication package includes the scripts and data we extracted, leading us to our findings.</p>
Replication Package of the Paper "Towards Better Understanding of Breaking Changes in the NPM Ecosystem"
<p><span><strong>Replication Package of the paper "Towards Better Comprehension of Breaking Changes in the NPM Ecosystem"</strong></span></p> <p> </p> <p>We describe the files in this replication package as follows.</p> <p> </p> <p><strong>collected_data folder:</strong></p> <p>This folder contains the our collected breaking changes from 381 sampled NPM projects. The files include:</p> <p>1) sampled_projects.txt, the 381 projects we sampled.</p> <p>2) original_breaking_commits.csv, the breaking commits obtained from the 381 projects.</p> <p>3) breaking_commits_after_removal.csv, it contains the breaking commits after removing (1) non-JavaScript source code change, (2) contain very long commit messages over 10 lines.</p> <p> </p> <p><strong>RQ1 folder:</strong></p> <p>This folder contains the breaking changes used in RQ1.</p> <p>1) documented_bc_can_be_detected.csv, it contains the breaking changes after removal, indicating whether a documented breaking change can be detected by test cases.</p> <p>2) detected_bc_are_documented.csv, it contains the breaking changes sampled from all commits from 381 projects, indicating whether a detected breaking commit is documented by developers.</p> <p> </p> <p><strong>RQ2-3-4 folder:</strong></p> <p>This folder contains the breaking changes used in analysis process in RQ2, RQ3 and RQ4.</p> <p>1) used_projects.csv, the projects that contain breaking changes.</p> <p>2) analyzed_breaking_changes.csv, the analyzed breaking changes in RQ2 to 4. The breaking changes are annotated. For example, column “category” indicates the type of the breaking change, column “change_signature_type” is related to RQ2, column “change_behavior_type” is related to RQ3 and column “reason” is related to RQ4.</p>
Replication Package: Unboxing Default Argument Breaking Changes in 1 + 2 Data Science Libraries in Python
<p><strong>Replication Package</strong></p> <p>This repository contains data and source files needed to replicate our work described in the paper "Unboxing Default Argument Breaking Changes in Scikit Learn".</p> <p><strong>Requirements</strong></p> <p>We recommend the following requirements to replicate our study:</p> <ol> <li>Internet access</li> <li>At least 100GB of space</li> <li>Docker installed</li> <li>Git installed</li> </ol> <p><strong>Package Structure</strong></p> <p>We relied on Docker containers to provide a working environment that is easier to replicate. Specifically, we configure the following containers:</p> <ul> <li><code>data-analysis</code>, an R-based Container we used to run our data analysis.</li> <li><code>data-collection</code>, a Python Container we used to collect Scikit's default arguments and detect them in client applications.</li> <li><code>database</code>, a Postgres Container we used to store clients' data, obtainer from Grotov et al.</li> <li><code>storage</code>, a directory used to store the data processed in <code>data-analysis</code> and <code>data-collection</code>. This directory is shared in both containers.</li> <li><code>docker-compose.yml</code>, the Docker file that configures all containers used in the package.</li> </ul> <p>In the remainder of this document, we describe how to set up each container properly.</p> <p><strong>Using VSCode to Setup the Package</strong></p> <p>We selected VSCode as the IDE of choice because its extensions allow us to implement our scripts directly inside the containers. In this package, we provide configuration parameters for both <code>data-analysis</code> and <code>data-collection</code> containers. This way you can directly access and run each container inside it without any specific configuration.</p> <p>You first need to set up the containers</p> <pre><code>$ cd /replication/package/folder $ docker-compose build $ docker-compose up # Wait docker creating and running all containers </code></pre> <p>Then, you can open them in Visual Studio Code:</p> <ol> <li>Open VSCode in project root folder</li> <li>Access the command palette and select "Dev Container: Reopen in Container" <ol> <li>Select either <em>Data Collection</em> or <em>Data Analysis</em>.</li> </ol> </li> <li>Start working</li> </ol> <p>If you want/need a more customized organization, the remainder of this file describes it in detail.</p> <p><strong>Longest Road: Manual Package Setup</strong></p> <p><strong>Database Setup</strong></p> <p>The database container will automatically restore the dump in <code>dump_matroskin.tar</code> in its first launch. To set up and run the container, you should:</p> <p>Build an image:</p> <pre><code>$ cd ./database $ docker build --tag 'dabc-database' . $ docker image ls REPOSITORY TAG IMAGE ID CREATED SIZE dabc-database latest b6f8af99c90d 50 minutes ago 18.5GB </code></pre> <p>Create and enter inside the container:</p> <pre><code>$ docker run -it --name dabc-database-1 dabc-database $ docker exec -it dabc-database-1 /bin/bash root# psql -U postgres -h localhost -d jupyter-notebooks jupyter-notebooks=# \dt List of relations Schema | Name | Type | Owner --------+-------------------+-------+------- public | Cell | table | root public | Code_cell | table | root public | Md_cell | table | root public | Notebook | table | root public | Notebook_features | table | root public | Notebook_metadata | table | root public | repository | table | root </code></pre> <p>If you got the tables list as above, your database is properly setup.</p> <p>It is important to mention that this database is extended from the one provided by <a href="https://markdowntohtml.com/">Grotov et al.</a>. Basically, we added three columns in the table <code>Notebook_features</code> (<code>API_functions_calls</code>, <code>defined_functions_calls</code>, and<code>other_functions_calls</code>) containing the function calls performed by each client in the database.</p> <p><strong>Data Collection Setup</strong></p> <p>This container is responsible for collecting the data to answer our research questions. It has the following structure:</p> <ul> <li><code>dabcs.py</code>, extract DABCs from Scikit Learn source code, and export them to a CSV file.</li> <li><code>dabcs-clients.py</code>, extract function calls from clients and export them to a CSV file. We rely on a modified version of <a href="https://markdowntohtml.com/">Matroskin</a> to leverage the function calls. You can find the tool's source code in the `matroskin`` directory.</li> <li><code>Makefile</code>, commands to set up and run both <code>dabcs.py</code> and <code>dabcs-clients.py</code></li> <li><code>matroskin</code>, the directory containing the modified version of matroskin tool. We extended the library to collect the function calls performed on the client notebooks of Grotov's dataset.</li> <li><code>storage</code>, a docker volume where the data-collection should save the exported data. This data will be used later in <a href="https://markdowntohtml.com/#data-analysis-setup">Data Analysis</a>.</li> <li><code>requirements.txt</code>, Python dependencies adopted in this module.</li> </ul> <p>Note that the container will automatically configure this module for you, e.g., install dependencies, configure matroskin, download scikit learn source code, etc. For this, you must run the following commands:</p> <pre><code>$ cd ./data-collection $ docker build --tag "data-collection" . $ docker run -it -d --name data-collection-1 -v $(pwd)/:/data-collection -v $(pwd)/../storage/:/data-collection/storage/ data-collection $ docker exec -it data-collection-1 /bin/bash $ ls Dockerfile Makefile config.yml dabcs-clients.py dabcs.py matroskin storage requirements.txt utils.py </code></pre> <p>If you see project files, it means the container is configured accordingly.</p> <p><strong>Data Analysis Setup</strong></p> <p>We use this container to conduct the analysis over the data produced by the <a href="https://markdowntohtml.com/#data-collection-setup">Data Collection</a> container. It has the following structure:</p> <ul> <li><code>dependencies.R</code>, an R script containing the dependencies used in our data analysis.</li> <li><code>data-analysis.Rmd</code>, the R notebook we used to perform our data analysis</li> <li><code>datasets</code>, a docker volume pointing to the <code>storage</code> directory.</li> </ul> <p>Execute the following commands to run this container:</p> <pre><code>$ cd ./data-analysis $ docker build --tag "data-analysis" . $ docker run -it -d --name data-analysis-1 -v $(pwd)/:/data-analysis -v $(pwd)/../storage/:/data-collection/datasets/ data-analysis $ docker exec -it data-analysis-1 /bin/bash $ ls data-analysis.Rmd datasets dependencies.R Dockerfile figures Makefile </code></pre> <p>If you see project files, it means the container is configured accordingly.</p> <p>A note on <code>storage</code> shared folder</p> <p>As mentioned, the <code>storage</code> folder is mounted as a volume and shared between <code>data-collection</code> and <code>data-analysis</code> containers. We compressed the content of this folder due to space constraints. Therefore, before starting working on <a href="https://markdowntohtml.com/#data-collection-setup">Data Collection</a> or <a href="https://markdowntohtml.com/#data-analysis-setup">Data Analysis</a>, make sure you extracted the compressed files. You can do this by running the <code>Makefile</code> inside <code>storage</code> folder.</p> <pre><code>$ make unzip # extract files $ ls clients-dabcs.csv clients-validation.csv dabcs.csv Makefile scikit-learn-versions.csv versions.csv $ make zip # compress files $ ls csv-files.tar.gz Makefile</code></pre>
ATM-Dependent Phosphorylation of Nemo SQ Motifs is Dispensable for Nemo-Mediated Gene Expression Changes in Response to DNA Double Strand Breaks
GEO Series GSE264315. Mus musculus. 18 samples. Type: Expression profiling by high throughput sequencing.
DNA breaks and chromatin structural changes enhance the transcription of Autoimmune Regulator target genes [RNA-Seq]
GEO Series GSE76240. Homo sapiens. 8 samples. Type: Expression profiling by high throughput sequencing.
Profiling cell-type-specific transcriptional changes and DNA break sites in response to contextual fear learning
GEO Series GSE155095. Mus musculus. 88 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
CGH of stage 13 amplifying follicle cells to measure changes in replication fork progression in DNA damage checkpoint and double-strand break repair mutants
GEO Series GSE66690. Drosophila melanogaster. 25 samples. Type: Genome variation profiling by genome tiling array.
Meiotic DNA breaks activate a streamlined response that largely avoids protein level changes
GEO Series GSE197022. Saccharomyces cerevisiae. 6 samples. Type: Expression profiling by high throughput sequencing.
DNA breaks and chromatin structural changes enhance the transcription of Autoimmune Regulator target genes [FAIRE-Seq]
GEO Series GSE89892. Homo sapiens. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
DNA breaks and chromatin structural changes enhance the transcription of Autoimmune Regulator target genes
GEO Series GSE89893. Homo sapiens. 16 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
Replication Dataset of the paper "Towards Better Understanding of Breaking Changes in NPM Ecosystem"
<p>Replication Dataset of the paper "Towards Better Understanding of Breaking Changes in NPM Ecosystem"</p>
Dataset of the paper "Towards Better Comprehension of Breaking Changes in the NPM Ecosystem"
<p>This is the dataset of the paper "Towards Better Comprehension of Breaking Changes in the NPM Ecosystem".</p> <div> <div>We describe the files in this replication package as follows. For each RQ, we present the dataset of the results of the RQ. For each result table, we provide both Excel and CSV file formats.</div> <br> <div><strong>Scripts folder:</strong></div> <div>This folder contains the scripts for cloning repositories, obtaining breaking change commits and running test cases (in RQ1).</div> <br> <div><strong>Collected_data folder:</strong></div> <div>This folder contains our collected breaking changes from 381 sampled NPM projects. The files include:</div> <div>1) sampled_projects.txt, the 381 projects we sampled.</div> <div>2) original_breaking_commits.{csv, xlsx}, the breaking commits obtained from the 381 projects (5,242 in total)</div> <div>3) breaking_commits_after_removal.{csv, xlsx}, it contains the breaking commits after removing (1) non-JavaScript source code change, (2) contains very long commit messages over 10 lines (The detailed process is in Section 3.1 of the paper). There are 2,724 breaking changes in total.</div> <div> </div> <div><strong>RQ1 folder:</strong></div> <div>This folder contains the breaking changes used in RQ1 (Section 4.1).</div> <div>1) documented_bc_can_be_detected.{csv, xlsx}, it contains the breaking changes after removal, and the “can_be_detected” column indicates whether a documented breaking change can be detected by test cases.</div> <div>2) detected_bc_are_documented.{csv, xlsx}, it contains the breaking changes sampled from all commits from 381 projects. The “can_be_detected” column indicates whether a commit can be detected by test cases, and the “documented” column indicates whether a commit is a documented breaking change.</div> <div>The script of running test cases “run_tests.py” is in “Scripts” folder.</div> <div> </div> <div><strong>RQ2-3-4 folder:</strong></div> <div>This folder contains the breaking changes used in the analysis process of RQ2, RQ3 and RQ4. Compared to the breaking commits in the Collected_data folder, we remove the breaking changes that cannot be linked to reason information. This is detailed in Section 3.3 of the paper:</div> <div>1) used_projects.{csv, xlsx}, the projects that contain breaking changes (131 in total)</div> <div>2) analyzed_breaking_changes.{csv, xlsx}, the analyzed breaking changes in RQ2 to 4. All breaking changes are annotated: column “category” indicates the type of the breaking change (the classification process is detailed in Section 3.2), column “change_signature_type” is for RQ2 (Section 4.2), column “change_behavior_type” is for RQ3 (Section 4.3) and column “reason” is for reasons behind breaking changes in RQ4 (Section 4.4).</div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.