Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
677
datasets available to search
ShareScore release 0.9.0
Dataset results
677 results for “Replication package”
Adjustment costs and factor demand: New evidence from firms' real estate - Replication packages
<p>This is the replication package associated with the article "Adjustment costs and factor demand. New evidence from firms’ real estate" published in Economic Journal.</p>
poojaruhal/RP-commenting-practices-social-media: RP-commenting-practices-social-media: RP_TOSEM_2020 v.1.0.1 Second release of of the replication Package for the paper "What do Developers Discuss about Code Comment Conventions on Social Media"
<p>RP-commenting-practices-social-media</p> <p>Replication Package for the paper "What do Developers Discuss about Code Comment Conventions on Social Media?"</p> <p>Structure</p> <pre><code>Paper-presenation.pdf Makar_tool/ Data/ stackoverfow_questions_with_answers_by_tags.csv stackoverfow_tags_metrics.csv apache_mailing_list.csv mailing_lists_ASF_@dev_@users_1.csv mailing_lists_ASF_@dev_@users_2.csv quora.csv sample_stackoverfow_questions_with_answers_by_tags.csv Schemas/ apache_mailing_lists.json quora.json stackoverfow_questions_answers_by_tag.json stackoverfow_tag_count.json stackoverfow_tag_metrics.json RQ1/ LDA_input/ stackoverfow_raw_dataset.csv LDA_output/ Mallet/ output_csv/ docs-in-topics.csv topic-words.csv topics-in-docs.csv topics-metadata.csv output_html/ all_topics.html Docs/ Topics/ RQ2/ datasource_rawdata/ mailing_lists_selection_criteria.csv quora.csv stackoverflow.csv manual_analysis_output/ stackoverflow_quora_taxonomy.xlsx </code></pre> <p>Contents of the Replication Package</p> <p><strong>Paper-presenation.pdf</strong> presents the highlights of the work in a presenation.</p> <p><strong>Makar_tool/</strong> contains the data processed using the tool for the study</p> <ul> <li> <p><strong>Data/</strong></p> <ul> <li><code>stackoverfow_questions_with_answers_by_tags.csv</code> - all stackoverflow questions used in the study as stored in Makar <ul> <li><code>stackoverfow_tags_metrics.csv</code> - all data containing the calculations done for stackoverflow tag selection</li> <li><code>apache_mailing_list.csv</code> - statistically significant sample of <code>mailing_lists_ASF_@dev_@users_1.csv</code> and <code>mailing_lists_ASF_@dev_@users_2.csv</code> used in the study</li> <li><code>mailing_lists_ASF_@dev_@users_1.csv</code> - mailing list data used in the study as stored in Makar (part 1)</li> <li><code>mailing_lists_ASF_@dev_@users_2.csv</code> - mailing list data used in the study as stored in Makar (part 2)</li> <li><code>quora.csv</code> - all quora questions used in the study as stored in Makar</li> <li><code>sample_stackoverfow_questions_with_answers_by_tags</code> - statistically significant sample of <code>stackoverfow_questions_with_answers_by_tags.csv</code> used in the study</li> </ul> </li> </ul> </li> <li> <p><strong>Schemas/</strong></p> <ul> <li><code>apache_mailing_lists.json</code> - data schema used in Makar to store mailing list data</li> <li><code>quora.json</code> - data schema used in Makar to store quora data</li> <li><code>stackoverfow_questions_answers_by_tag.json</code> - data schema used in Makar to store stackoverflow questions data</li> <li><code>stackoverfow_tag_count.json</code> - data schema used in Makar to lookup number of questions per tag available in stackoverflow</li> <li><code>stackoverfow_tag_metrics.json</code> - data schema used in Makar to stackoverflow tag metrics data</li> </ul> </li> <li> <p><strong>RQ1/</strong> - contains the data used to answer RQ1</p> <ul> <li><strong>LDA_input/</strong> - input data used for LDA analysis <ul> <li><code>stackoverfow_raw_dataset.csv</code> - stackoverflow questions used to perform LDA analysis</li> </ul> </li> <li><strong>LDA_output/</strong> <ul> <li><strong>Mallet/</strong> - contains the LDA output generated by MALLET tool <ul> <li><strong>output_csv/</strong> <ul> <li><code>docs-in-topics.csv</code> - documents per topic</li> <li><code>topic-words.csv</code> - most relevant topic words</li> <li><code>topics-in-docs.csv</code> - topic probability per document</li> <li><code>topics-metadata.csv</code> - metadata per document and topic probability <ul> <li><strong>output_html/</strong> - Browsable results of mallet output</li> </ul> </li> <li><code>all_topics.html</code></li> <li><code>Docs/</code></li> <li><code>Topics/</code></li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> <li> <p><strong>RQ2/</strong> - contains the data used to answer RQ2</p> <ul> <li><strong>datasource_rawdata/</strong> - contains the raw data for each source <ul> <li><code>mailing_lists_selection_criteria.csv</code> - criteria used to select mailing_lists.</li> <li><code>quora.csv</code> - contains the processed dataset (like removing HTML tags). To know more about the preprocessing steps, please refer to the reproducibility section in the paper. The data is preprocessed using Makar tool.</li> <li><code>stackoverflow.csv</code> - contains the processed stackoverflow dataset. To know more about the preprocessing steps, please refer to the reproducibility section in the paper. The data is preprocessed using Makar tool.</li> </ul> </li> <li><strong>manual_analysis_output/</strong> <ul> <li><code>stackoverflow_quora_taxonomy.xlsx</code> - contains the classified dataset of stackoverflow and quora and description of taxonomy. <ul> <li><code>Taxonomy</code> - contains the description of the first dimension and second dimension categories. Second dimension categories are further divided into levels, separated by <code>|</code> symbol.</li> <li><code>stackoverflow-posts</code> - the questions are labelled relevant or irrelevant and categorized into the first dimension and second dimension categories. <ul> <li><code>quota-posts</code> - the questions are labelled relevant or irrelevant and categorized into the first dimension and second dimension categories.</li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> </ul>
Replication package for «Business disruptions from social distancing»
<p>Replication package for</p> <blockquote> <p>Koren, Miklós and Rita Pető. Business disruptions from social distancing; PLoS ONE 2020. <a href="https://doi.org/10.1371/journal.pone.0239113">https://doi.org/10.1371/journal.pone.0239113</a> </p> </blockquote>
The Probabilistic Model Checker Storm: Evaluation Results and Replication Package
<p>This package contains logdata and replication scripts for the evaluation of the model checker Storm (www.stormchecker.org) as part of the paper:<br> "The Probabilistic Model Checker Storm" by Christian Hensel, Sebastian Junges, Joost-Pieter Katoen, Tim Quatmann, and Matthias Volk</p>
Replication package for: The Aggregate Implications of Mergers and Acquisitions
<p>The package contains the codes to replicate the figures and tables in David (forthcoming). "The Aggregate Implications of Mergers and Acquisitions." Review of Economic Studies. Detailed instructions are also given about accessing the raw data.</p>
Replication package for Nonparametric Analysis of Time-Inconsistent Preferences
<p>Stata datasets, R and Julia files.</p>
Replication package for "An Industrial Study on the Challenges and Effects of Diversity-based Testing in Continuous Integration"
<p>This is the replication package for the analysis done in the paper "An Industrial Study on the Challenges and Effects of Diversity-based Testing in Continuous Integration".</p> <p>The package includes: (i) CSV files with data on test case and corresponding feature coverage, as well as test execution data; (ii) CSV files including failure coverage for the executed techniques; (iii) R scripts to re-run our visual and statistical analysis when comparing techniques results; and (iv) An R Markdown file (rendered into HTML) detailing the steps of our analysis with the coresponding code.</p>
Replication package of the paper titled "Mining the ROS ecosystem for Green Architectural Tactics in Robotics and an Empirical Evaluation".
<p>This is the replication package of the paper titled "Mining the ROS ecosystem for Green Architectural Tactics in Robotics and an Empirical Evaluation".</p> <p>The replication package is structured according to the research questions of the study (RQ1 and RQ2) and it is composed of the following elements:</p> <ul> <li><strong>Supplementary Material (RQ1).pdf</strong>: a 93-pages technical report providing the details about our mining activities and application of thematic analysis for identifying green architectural tactics for robotics software.</li> <li><strong>Supplementary Material (RQ2).pdf</strong>: a 75-pages technical report detailing the design, conduction, and results of the empirical assessment of the identified green tactics.</li> <li><strong>RQ1_data_software.zip</strong>: the raw data and mining source code related to all the activities we carried out for answering RQ1.</li> <li><strong>RQ2_data.zip</strong>: the raw data related to all the activities we carried out for answering RQ2.</li> <li><strong>RQ2_ros_implementation.zip</strong>: the source code of the ROS-based system we implemented for performing the experiment in RQ2.</li> </ul>
Replication Package for Bloom, Draca and Van Reenen (2021) "A Reply to Campbell and Mau"
<p>Replication File for Review of Economic Studies manuscript 28949-2 "A Reply to Campbell and Mau"</p>
Replication package for: Recoverability and Expectations-Driven Fluctuations
<p>This package contains all the data and code necessary to reproduce the results in Chahrour and Jurado, "Recoverability and Expectations-Driven Fluctuations." Review of Economic Studies (forthcoming). </p>
Replication package - The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding
<p>Replication package for:</p> <p>Marvin Wyrich, Andreas Preikschat, Daniel Graziotin and Stefan Wagner. The Mind Is a Powerful Place: How Showing Code Comprehensibility Metrics Influences Code Understanding. To appear in: Proceedings of the 43rd International Conference on Software Engineering (ICSE ’21), Madrid, Spain.</p> <p>- The `data` folder contains dataset, R analysis script, and figures for the paper<br> - the `material` folder contains the snippets used for the experimental tasks (`code snippets` subfolder), the environment in which participants had to understand the code (`environment` subfolder) and the task sheet form to evaluate code comprehension.</p> <p>The R analysis script allows to reproduce all statistical analysis steps of the paper as well as its figures. We recommend calling `setwd()` before running the script contents, so that the dataset can be properly loaded.</p> <p>For ethics and privacy reasons, we had to drop the following potentially identifying fields from the openly released dataset, because the fields might enable participants to identify their peers:</p> <p>- Age<br> - Gender<br> - Subject (study plan)<br> - Program<br> - Semester no.</p> <p>These missed details will still allow full reproducibility of the study.</p>
Replication Package: Towards Using Package Centrality Trend to Identify Packages in Decline
<p>Due to its increasing complexity, today’s software systems are frequently built by leveraging reusable code in the form of libraries and packages. Software ecosystems (e.g., npm) are the primary enablers of this code reuse, providing developers with a platform to share their own and use others’ code. These ecosystems evolve rapidly: developers add new packages every day to solve new problems or provide alternative solutions, causing obsolete packages to decline in their importance to the community. Developers should avoid depending on packages in decline, as these packages are reused less over time and may become less frequently maintained. However, current popularity metrics are not fit to provide this information to developers.</p> <p>In this paper, we propose a scalable approach that uses the package’s centrality in the ecosystem to identify packages in decline. We evaluate our approach with the npm ecosystem and show that the trends of centrality over time can correctly distinguish packages in decline with an ROC-AUC of 0.9. The approach can capture 87% of the packages in decline, on average 18 months before the trend is shown in currently used package popularity metrics. We implement this approach in a tool that can be used to augment npms metrics and help developers avoid packages in decline when reusing packages from npm.</p>
Replication Package for the Paper "OSS License Identification at Scale: A Comprehensive Dataset Using World of Code"
<div> <div># Replication Package for the Paper "OSS License Identification at Scale: A Comprehensive Dataset Using World of Code"<br><br>containing the dataset and scripts used to create it.</div> </div>
Replication Package for the Paper: "How Do Users Revise Architectural Related Questions on Stack Overflow: An Empirical Study"
<p>This is the replication package for the paper: "How Do Users Revise Architectural Related Questions on Stack Overflow: An Empirical Study". In the following, we provide a brief description of the folders and files:</p> <p><strong>(1) raw data</strong></p> <p>The raw data folder contains the retrieved 36,417 posts and the SQL query used for retrieving ARPs from Stack Overflow through the query interface provided by Stack Exchange.</p> <p><strong>(2) filtered ARPs</strong></p> <p>The filtered ARPs folder contains 13,205 filtered candidates ARPs from the retrieved 36,417 posts and the results of data analysis for the first RQ (i.e., RQ1).</p> <p><strong>(3) randomly selected posts and labeling results</strong></p> <p>The randomly selected posts and labeling results folder contains 1,068 randomly selected posts and their labeling results (i.e., 21 ARPs, wherein 14.3%, 3 out of 21 ARPs, do not contain “architect*” terms and 85.7%, 18 out of 21 APRs, contain “architect*” terms).</p> <p><strong>(4) relevant ARPs for answering RQs</strong></p> <p>The relevant ARPs for answering RQs folder contains 4,114 ARPs with revision information for answering the last three RQs (i.e., RQ2, RQ3, and RQ4).</p> <p><strong>(5) interview responses</strong></p> <p>The interview responses folder contains 11 collected interview responses from software practitioners. These responses were gathered to evaluate the identified categories related to ARQ revisions.</p> <p><strong>(6) data extraction and analysis</strong></p> <p>The data extraction and analysis folder contains the MAXQDA file. Data Labeling & Encoding for RQs.mx20 is the results of data labeling and encoding for RQ2, RQ3, and RQ4, which were analyzed by the MAXQDA tool. This file can be opened by MAXQDA 2020 or higher versions, which are available at https://www.maxqda.com/ for download. You may also use the free 14 days trial version of MAXQDA 2020, which is available at https://www.maxqda.com/trial for download.</p>
Does Co-Development with AI Assistants Lead to More Maintainable Code? Replication Package
<p>This is the replication package for the study "Echoes of AI: Investigating the Downstream Effects of AI Assistants on Software Maintainability" preregistered aat ICSME 2024 as “Does Co-Development with AI Assistants Lead to More Maintainable Code?”</p> <p>Abstract from the registered report:</p> <p>[Background/Context] AI assistants like GitHub Copilot are transforming software engineering, with several studies highlighting productivity improvements. However, their impact on code quality, particularly in terms of maintainability, requires further investigation.<br>[Objective/Aim] This study aims to examine the influence of AI assistants on software maintainability, specifically assessing how these tools affect the ability of developers to evolve code.<br>[Method] We will conduct a two-phased controlled experiment involving professional developers. In Phase 1, developers will add a new feature to a Java project, with or without the aid of an AI assistant. Phase 2, a randomized controlled trial, will involve a different set of developers evolving random Phase 1 projects - working without AI assistants. We will employ Bayesian analysis to evaluate differences in completion time, perceived productivity, code quality, and test coverage.</p> <p>Note: To maintain the integrity of the study, i.e., preventing any leakage to AI assistants' training data, we choose not to host the code in a public git repository. Instead, all relevant documents and code are shared through a replication package on Zenodo, available as PDF documents generated by repo2pdf (https://github.com/BankkRoll/repo2pdf). We have deliberately used settings to obfuscate the code (e.g., line numbers) to ensure it will not be scraped by any large language models before the study has been completed.</p> <p>Contents:</p> <ul> <li>Task 1 instructions.</li> <li>Task 2 instructions.</li> <li>The source code that the participants received.</li> <li>A causal graph with analysis details.</li> <li>Archives containing anonymized experimental data and analysis scripts (in .zip and .tar.gz for convenience). </li> </ul>
Replication package for: Disability and risk preferences: Experimental and survey evidence from Vietnam
<p>The replication package contains all data files and programs (Stata) in order to reproduce the tables and figures of the manuscript.</p>
Replication Package for 'How do Machine Learning Models Change?'
<h1>Replication Package: How Do Machine Learning Models Change?</h1> <p> </p> <h2>Overview</h2> <div>This replication package accompanies the paper "<strong>How Do Machine Learning Models Change?</strong>". In this study, we conducted a large-scale analysis of <strong>over 680,000 commits from 100,000 models</strong> and <strong>2,251 releases from 202 of these models</strong> on the Hugging Face (HF) platform. Our goal was to understand how machine learning (ML) models evolve by classifying commit types using a detailed ML change taxonomy and analyzing temporal patterns in their activities using Bayesian networks.</div> <p> </p> <div>Our research addresses three main aspects:</div> <div><strong>1. Categorization of Commit Changes:</strong> We classified over 960,000 commits on HF, providing a detailed breakdown of change types and their distribution.</div> <div><strong>2. Analysis of Commit Sequences</strong>: We examined the sequence and dependencies of commit types using Bayesian networks to identify temporal patterns.</div> <div><strong>3. Release Analysis**</strong>: We investigated the distribution and evolution of release types, analyzing how model attributes and metadata change across successive releases.</div> <p> </p> <div>This package provides all the necessary code, data, and documentation to reproduce the results presented in our paper.</div> <p> </p> <h2>Data Collection and Preprocessing</h2> <h3>Data Collection</h3> <div>We collected data from the Hugging Face platform using the Hugging Face Hub API. The data extraction was performed up to <strong>May 2025,</strong> capturing details from over 1 million models available at that time.</div> <p> </p> <div><strong>- Model & Release Information</strong>: We collected model metadata, commit histories, and release information (Git tags) for our sampled models.</div> <div><strong>- Detailed Commit Changes</strong>: To get a detailed list of files modified in each commit, we implemented a direct Git processing approach. For each model, its repository was temporarily cloned to programmatically extract the list of changed files for every commit SHA.</div> <p> </p> <h3>Data Preprocessing</h3> <div><strong>Commit Diffs</strong></div> <div>We computed the differences for key JSON configuration files (e.g., `config.json`) between commits to identify added, deleted, and updated keys, which served as input for classification.</div> <div> </div> <div><strong>Commit Classification</strong></div> <div>We classified each commit according to Bhatia et al.'s ML change taxonomy using the <strong>Gemini 2.5 Flash LLM</strong>. To ensure the reliability of this process, we implemented a rigorous <strong>two-phase validation</strong>:</div> <div><strong>1. Prompt Refinement (Training): </strong>The prompt was iteratively refined over 6 cycles using a curated training set of 143 commits. The process was guided by comparing LLM classifications against a gold standard created by two human annotators (Human-Human IRR on a subset: 𝜅 = 0.7798). The final refined prompt achieved a Kappa of 0.9068 against the training gold standard.</div> <div><strong>2. Final Validation (Testing): </strong>The validated prompt was tested on an independent, statistically significant sample of 384 commits. The LLM's classifications achieved a Cohen's Kappa of 0.8568 when compared against the test set's gold standard, which itself was validated with a human-human IRR of 𝜅 = 0.8150.</div> <p> </p> <div>We also classified commits into Swanson's categories using a fine-tuned DistilBERT model, as detailed in the paper.</div> <p> </p> <h2>Folder Structure</h2> <div>The replication package is organized as follows. The structure has been designed to separate code, data, and metadata for clarity.</div> <p> </p> <ul> <li>`code/`: Contains all Jupyter notebooks for the study. <ul> <li>`Collection/`: Scripts for data extraction from Hugging Face.</li> <li>`HFExtraction.ipynb`: Collects primary model and commit information.</li> <li>`HFReleasesExtraction.ipynb`: Collects release (tag) specific information.</li> </ul> </li> <li>`Preprocessing/`: Scripts for data cleaning, processing, and classification. <ul> <li>`HFCommitsPreprocessing.ipynb`: Processes commits, computes diffs, and prepares data for classification and analysis.</li> <li>`HFReleasesPreprocessing.ipynb`: Processes and classifies release data.</li> </ul> </li> <li>`Analysis/`: Notebooks for reproducing the analysis for each research question. <ul> <li>`HFFileChanges.ipynb`: Contains the preliminary analysis of file change patterns.</li> <li>`RQ1_Analysis.ipynb`: Analysis for Research Question 1.</li> <li>`RQ2_Analysis.ipynb`: Analysis for Research Question 2.</li> <li>`RQ3_Analysis.ipynb`: Analysis for Research Question 3.</li> </ul> </li> <li>`datasets/`: Contains the key final datasets used in the analysis notebooks. <ul> <li>`commits_datasets/`: Contains the main classified commit dataset. <ul> <li>`HFCommitsClassification_final.csv`: The final dataset with over 960,000 classified commits for RQ1 and RQ2.</li> </ul> </li> <li>`releases_datasets/`: Contains the datasets related to releases. <ul> <li>`HFReleasesClassification.csv`: The final dataset of 2,251 classified releases for RQ3.</li> </ul> </li> <li>`model_metadata.csv`: The extracted internal metadata from model files for RQ3.4.</li> </ul> </li> <li>`metadata/`: Contains configuration files and the data used for the validation process. <ul> <li>`validation_data/`: A sub-folder containing the gold standard data. <ul> <li>`Agreement TOSEM Commit Changes.xlsx`: Excel containing details of the classication and validation processes.</li> <li>`prompt_refinement.txt`: The final, validated prompt used for the LLM classification along its previous iterations.</li> <li>`training_set_ground_truth.json`: Gold standard for the 143-commit training set.</li> <li>`training_set_first_classification.json`: First annotator's labels for the training IRR subset.</li> <li>`training_set_second_classification.json`: Second annotator's labels for the training IRR subset.</li> <li>`test_set_ground_truth.json`: Gold standard for the 384-commit test set.</li> <li>`test_first_classification.json`: First annotator's labels for the training IRR subset.</li> </ul> </li> <li>`test_set_second_classification.json`: Second annotator's labels for the test IRR subset.</li> <li>`tags_metadata.yaml`: Auxiliary metadata file used during preprocessing.</li> </ul> </li> <li>`README.md`: This file.</li> <li>`requirements.txt`: Lists the required Python packages.</li> </ul> <p>*Note: Other intermediate CSV files are provided to facilitate re-running specific parts of the analysis without starting from scratch.*</p> <p> </p> <h2>How to Use This Package</h2> <p> </p> <h3>Setup</h3> <div><strong>1. Create and activate a virtual environment</strong> (recommended).</div> <div>```bash</div> <div>python -m venv venv</div> <div>source venv/bin/activate # On Windows: venv\Scripts\activate</div> <div>```</div> <div><strong>2. Install required packages.</strong></div> <div>```bash</div> <div>pip install -r requirements.txt</div> <div>```</div> <p> </p> <h3>Running the Analysis</h3> <div>The Jupyter notebooks in the `code/` directory are numbered and named to be run in a logical sequence: <strong>Collection -> Preprocessing -> Analysis.</strong> We recommend following this order.</div> <div> </div> <div>- <strong>To reproduce our findings directly,</strong> you can start with the notebooks in `code/Analysis/`. They are configured to load the final, processed datasets provided in the `datasets/` folder.</div> <div>- <strong>To re-run the entire pipeline</strong>, start with the notebooks in `code/Collection/`, followed by `code/Preprocessing/`. Please note that running the full data collection and classification pipeline is time-consuming and may require significant computational resources and appropriate API keys for the LLM.</div> <p> </p> <h3>Key Datasets Provided</h3> <div><strong>- For RQ1 & RQ2:</strong> `datasets/commits_datasets/HFCommitsClassification_final.csv` (100,000 models for RQ1; filtered to 14,343 models for RQ2).</div> <div><strong>- For RQ3.1-3.3:</strong>`datasets/releases_datasets/HFReleasesClassification.csv` (2,251 releases from 202 models).</div> <div><strong>- For RQ3.4: </strong>`datasets/releases_datasets/model_metadata.csv` (from 28 models).</div> <p> </p> <h2>Contact</h2> <div>If you have any questions or encounter issues with this package, please contact the corresponding author. If you find our work useful, please consider citing our paper.</div>
"On the Prevalence, Co-occurrence, and Impact of Infrastructure-as-Code Smells" Replication Package
<p>In this package, we provide the dataset for the paper: " On the Prevalence, Co-occurrence, and Impact of Infrastructure-as-Code Smells ''</p><p> </p><p>1 – we provide the generated data for each of the research questions.</p><p> </p><p>2 – we provide the scripts for each of the research questions.</p>
Replication package for: "Corruption and firm growth: evidence from around the world"
<p>Guriev, S., Fisman, R., Ioramashvili, C., & Plekhanov, A. (2023). Corruption and firm growth: evidence from around the world. The Economic Journal.</p>
Replication package for: "Heat and observed economic activity in the rich urban tropics"
<p>The data and code shared here pertain to an article that is forthcoming in the Economic Journal. Please consult the README file herein for details.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.