Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
296
datasets available to search
ShareScore release 0.9.0
Dataset results
296 results for “language models”
Human review for post-training improvement of low-resource language performance in large language models
<p>Large language models (LLMs) have significantly improved natural language processing, holding the potential to support health workers and their clients directly. Unfortunately, there is a substantial and variable drop in performance for low-resource languages. Here we present results from an exploratory case study in Malawi, aiming to enhance the performance of LLMs in Chichewa through innovative prompt engineering techniques. By focusing on practical evaluations over traditional metrics, we assess the subjective utility of LLM outputs, prioritizing end-user satisfaction. Our findings suggest that tailored prompt engineering may improve LLM utility in underserved linguistic contexts, offering a promising avenue to bridge the language inclusivity gap in digital health interventions.</p>
Do Large Language Models Have a Personality? A Psychometric Evaluation with Implications for Clinical Medicine and Mental Health AI Dataset
Open the record for dataset details and reuse information.
Artifacts for paper "Code Refinement with the Assistance of Conversation-Aided Large Language Models"
<p>All the conversation data collected and the source code for our method are inside the uploaded zip file.</p>
Large Language Models for Human-Machine Collaborative Particle Accelerator Tuning through Natural Language
<p>Autonomous tuning of particle accelerators is an active and challenging field of research with the goal of enabling novel accelerator technologies cutting-edge high-impact applications, such as physics discovery, cancer research and material sciences. A key challenge with autonomous accelerator tuning remains that the most capable algorithms require an expert in optimisation, machine learning or a similar field to implement the algorithm for every new tuning task. In this work, we propose the use of large language models (LLMs) to tune particle accelerators. We demonstrate on a proof-of-principle example the ability of LLMs to successfully and autonomously tune a particle accelerator subsystem based on nothing more than a natural language prompt from the operator, and compare the performance of our LLM-based solution to state-of-the-art optimisation algorithms, such as Bayesian optimisation (BO) and reinforcement learning-trained optimisation (RLO). In doing so, we also show how LLMs can perform numerical optimisation of a highly non-linear real-world objective function. Ultimately, this work represents yet another complex task that LLMs are capable of solving and promises to help accelerate the deployment of autonomous tuning algorithms to the day-to-day operations of particle accelerators.</p>
ETimeline: An Extensive Timeline Generation Dataset based on Large Language Model
<div> <div>Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when viewing news in isolation, making it easier to track the development of stories and understand the interrelationships between key events. Timelines have appeared in many commercial products, but there is a noticeable lack of research in this field in academia, and existing datasets need improvement in terms of effectiveness and scale. We propose the ETimeline dataset, which contains over 13,000 news articles, covering 600 bilingual timelines across 23 news domains. We collected more than 120,000 news articles as a candidate news pool and used the large language model (LLM) Pipeline to enhance performance, ultimately obtaining the ETimeline, and the news pool data will also be provided. This work contributes to the advancement of timeline generation research and supports a wide range of tasks, including topic generation and event relationships. We believe that this dataset will serve as a catalyst for innovative research and bridge the gap between academia and industry in understanding the practical application of technology services.</div> </div>
SonicParanoid2: fast, accurate, and comprehensive orthology inference with machine learning and language models
<p>This repository contains the documentation, test datasets and scripts used in the following study:</p> <p>"SonicParanoid2: fast, accurate, and comprehensive orthology inference with machine learning and language models"<br><br>- `sonic-manuscript-master.zip` contains all the scripts to reproduce the study, including those for generating the figures and tables included in the manuscript.</p> <p>- <a href="../api/records/11361985/draft/files/sonicparanoid2.wiki.tar.xz/content" target="_blank" rel="noopener noreferrer">sonicparanoid2.wiki.tar.xz</a> contains a snapshot fo the wiki for SonicParanoid2 as of May 30, 2024</p>
(supplementary material) Fine-Tuning and Prompt Engineering for Large Language Models-based Code Review Automation
<div> <div> <div> <div>Supplementary material for paper <strong>"Fine-Tuning and Prompt Engineering for Large Language Models-based Code Review Automation"</strong></div> <div> </div> <div> <div> <div>The script for the paper can be found in this GitHub repository: https://github.com/awsm-research/LLM-for-code-review-automatiton</div> </div> </div> </div> </div> </div>
Incentivizing news consumption on social media platforms using large language models and realistic bot accounts
<p>This project examines how to enhance users' exposure to and engagement with verified and ideologically balanced news in an ecologically valid setting. We rely on a large-scale two-week long field experiment on 28,457 Twitter users. We created 28 bots utilizing GPT-2 that replied to users tweeting about sports, entertainment, or lifestyle with a contextual reply containing two hardcoded elements: a URL to the topic-relevant section of quality news organization and an encouragement to follow its Twitter account. Treated users were randomly assigned to receive responses by bots presented as female or male. We examine whether our intervention enhances the following of news media organization, the sharing/liking of news content and the tweeting/liking of political content. We find that the treated users followed more news accounts and the users in the female bot treatment were more likely to like news content than the control.</p>
Results outputs for "Beyond the Hype: Identifying and Analyzing Math Word Problem-Solving Challenges for Large Language Models"
<p>The provided files contain outputs generated by various Large Language Models (LLMs) for solving problems in the SVAMP dataset. Additionally, they include tagged statements of problems that LLMs incorrectly resolved.</p> <p>This repository includes the following two files:</p> <ul> <li>all_data.json --> Contains the generated samples for the SVAMP dataset.</li> <li>df_combined.pkl --> Contains the tagged SVAMP statements of problems that CodeLlama failed to resolve.</li> </ul>
SDG Mapping Results for 1,000 Publications from Large Language Models: GPT-4o, Mixtral, Llama 2, Llama 3, Gemma 2, Qwen 2 and GPT-4o-mini
<p>Randomly selected 1,000 publications from the Swinburne University of Technology research bank were used for SDG mapping tasks with the large language model GPT-4o and the open-source models Mixtral, Llama 2, Llama 3, Gemma 2, Qwen 2 and GPT-4o-mini.</p> <p>The input to each model consisted of the publication’s title and abstract.</p> <p>The designed prompt is as follows: </p> <p>PROMPT = '''Analyze the publication and determine the SDGs it aligns with. Evaluate against all the 17 SDGs provide the reason for alignment. In the end summarise the confidence levels(%) for each assigned SDG in JSON format. For example: {example}. Title: {title} Description: {description}'''</p> <p>example = '''{ 'Goal 6': 0.67, 'Goal 11': 0.50, 'Goal 3': 0.25}'''</p>
Spanish Language Benchmark for Artificial Intelligence Models (TELEIA)
<h1>TELEIA Datasets</h1> <p>The dataset contains test questions to evaluate LLMs in Spanish</p> <ul> <li>TELEIA_Cervantes_AVE.csv: test questions on vocabulary and grammatical structures, following the format of the Cervantes AVE exam. The column "question" contains a sentence with a gap that must be filled in. The columns "option_a", "option_b", "option_c", "option_d" and "correct_answer" (A,B,C or D) represent the different possibilities for the gap and the correct one.</li> <li>TELEIA_PCE.csv: test on morphology and semantics resembling the style of the PCE exam, consisting of short questions or sentences to be completed. It is composed of the following columns: "question", "option_a", "option_b", "option_c" and "correct_answer" (A, B or C).</li> <li>TELEIA_SIELE.csv: different texts with questions related to them, based on the reading comprehension task of the SIELE exam. It is composed of the following columns: "question", "option_a", "option_b", "option_c" and "correct_answer" (A, B or C). </li> </ul>
Supplementary Material for "Investigating Software Development Teams Members' Perceptions of Data Privacy in the Use of Large Language Models (LLMs)"
<h3>ABSTRACT<strong>: </strong></h3> <p><strong>Context</strong>: Large Language Models (LLMs) have revolutionized natural language generation and understanding. However, they raise significant data privacy concerns, especially when sensitive data is processed and stored by third parties. <br><strong>Goal</strong>: This paper investigates the perception of software development teams members regarding data privacy when using LLMs in their professional activities. Additionally, we examine the challenges faced and the practices adopted by these practitioners. <br><strong>Method</strong>: We conducted a survey with 78 ICT practitioners from five regions of the country. <br><strong>Results</strong>: Software development teams members have basic knowledge about data privacy and LGPD, but most have never received formal training on LLMs and possess only basic knowledge about them. Their main concerns include the leakage of sensitive data and the misuse of personal data. To mitigate risks, they avoid using sensitive data and implement anonymization techniques. The primary challenges practitioners face are ensuring transparency in the use of LLMs and minimizing data collection. Software development teams members consider current legislation inadequate for protecting data privacy in the context of LLM use. <br><strong>Conclusions</strong>: The results reveal a need to improve knowledge and practices related to data privacy in the context of LLM use. According to software development teams members, organizations need to invest in training, develop new tools, and adopt more robust policies to protect user data privacy. They advocate for a multifaceted approach that combines education, technology, and regulation to ensure the safe and responsible use of LLMs.</p>
Replication Kit: "Skill Models for Programming Language Concepts"
<p><strong>Structure</strong></p> <ul> <li><strong>data</strong>: contains the data we used for our case study <ul> <li><strong>skillmodels</strong>: data sets generated from the raw data in the database</li> <li><strong>raw</strong>: raw data collected in SmartAPE [1] containing the source code of the students as well as the assessment results of the system</li> </ul> </li> <li><strong>results</strong>: contains the complete results of our case study <ul> <li>Results of AUC and RMSE for each meta-parametrization and each skill model in .csv and .Rda format</li> <li>boxplots of our AUC and RMSE distributions for each skill model and each meta-parameter in .pdf format</li> </ul> </li> <li>calculation scripts: <ul> <li><strong>sm_trainer.R</strong>: script to fit different models for different meta-parametrizations and test them using different performance metrics. Uses <em>data/skillmodels </em>as input</li> <li><strong>comparison.R</strong>: script that performs statistical tests to compare meta-parameters. Uses <em>results//results_pfa.Rda, results//results_afm.Rda, and results/results_prop.Rda</em> as input</li> </ul> </li> </ul> <p><strong>References</strong></p> <p>[1] Albrecht, Ella et al. “Experiences in Introducing Blended Learning in an Introductory Programming Course.” <em>ECSEE</em> (2018).</p>
Addiitonal Files for The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling
<p>Additional Files 1-3 for "The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling"</p> <p> </p>
Leaderboard Spanish Language Benchmark for Artificial Intelligence Models (TELEIA)
<h1>TELEIA Datasets Leaderboard</h1> <p>These dataset contains the answers of different LLMs to the <a href="../doi/10.5281/zenodo.12571762">TELEIA (Spanish Language Benchmark for Artificial Intelligence Models)</a> dataset.<br><br><span>LLMs evaluated:</span></p> <ul> <li>Yi-6B-Chat</li> <li>Meta-Llama-3-8B-Instruct</li> <li>Llama-2-7b-chat-hf</li> <li>gemma-7b-it</li> <li>Mistral-7B-Instruct-v0.1</li> <li>occiglot-7b-es-en-instruct</li> <li>GPT3.5</li> <li>GPT4</li> </ul> <p><span>Files:</span></p> <ul> <li><em>TELEIA_Cervantes_AVE_results.xlsx: </em>vocabulary and grammatical structures, following the format of the Cervantes AVE exam</li> <li><em>TELEIA_PCE_results.xlsx: </em>test on morphology and semantics resembling the style of the PCE exam, consisting of short questions or sentences to be completed</li> <li><em>TELEIA_SIELE_results.xlsx: </em>different texts with questions related to them, based on the reading comprehension task of the SIELE exam</li> </ul> <p>Each .xlsx contains a sheet with the results of each model and the following columns:</p> <ul> <li><em>question: </em>question from TELEIA</li> <li><em>option_a:</em> possible answer from TELEIA </li> <li><em>option_b:</em> possible answer from TELEIA<em> </em></li> <li><em>option_c: </em>possible answer from TELEIA <em> </em></li> <li><em>option_d: </em>possible answer from TELEIA<em> </em></li> <li><em>correct_answer:</em> correct answer form TELEIA </li> <li><em>llm_question:</em> complete question made to the LLM </li> <li><em>tokens_in:</em> list of tokens that compound the question </li> <li><em>tokens_in_count:</em> number of tokens that compound the question </li> <li><em>llm_answer:</em> raw answer from the LLM </li> <li><em>llm_answer_filtered:</em> answer in format {A,B,C,D} from the LLM </li> <li><em>tokens_out :</em> list of tokens that compound the raw answer </li> <li><em>tokens_out_count:</em> number of tokens that compound the raw answer </li> <li><em>word_count :</em> number of words that compound the raw answer </li> </ul> <p> </p>
LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models - Replication Package
<p>LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models</p> <p>This is the replication package associated with the paper "LTM: Scalable and Black-box Similarity-based Test Suite Minimization based on Language Models".</p> <p><strong>Replication Package Contents:</strong></p> <p>This replication package contains all the necessary data and code required to reproduce the results reported in the paper. We provide the results of the Fault Detection Rate (FDR), Total Minimization Time (MT), Time Saving Rate (TSR) , statistical tests for all the minimization budgets (i.e., 25%, 50%, and 75%), results for the preliminary study, results for UniXcoder/Cosine with preprocessed code on 16 projects.</p> <p><strong>Data:</strong></p> <p>We provide in the <em><strong>Data</strong></em> directory the data used in our experiments, which is the source code of test cases (Java test methods) of 17 projects collected from Defects4J.</p> <p><strong>Code:</strong></p> <p>We provide in the<em> <strong>Code</strong> </em>directory the code (Python) and bash files required to run the experiments and reproduce the results.</p> <p><strong>Results:</strong></p> <p>We provide in the<em> <strong>Results </strong></em>directory the detailed results for our approach (called LTM). We also provide the summarized results of LTM and a baseline (ATM) for comparison purposes. Additional technical details about ATM can be found at https://zenodo.org/record/7455766.</p> <p><strong>_________________________________</strong></p> <p><strong>LTM's Similarity Measurement:</strong></p> <p>The source code of this step is in the <strong><em>Code/LTM/Similarity</em></strong> directory.</p> <p><strong>Requirements:</strong></p> <p>To run this step, Python 3 is required (we used Python 3.10). Also, the required libraries in the <em><strong>Code/LTM/Similarity/requirements.txt</strong></em> file should be installed, as follows:</p> <p>cd Code/LTM/Similarity</p> <p>pip install -r requirements.txt</p> <p><strong>Input:</strong></p> <ul> <li>Data/LTM/TestMethods</li> </ul> <p><strong>Output:</strong></p> <ul> <li>Data/LTM/similarity_measurements</li> </ul> <p><strong>Running the experiment:</strong></p> <p>To measure the similarity between all pairs of test cases, the following bash script should be executed:</p> <p>bash measure_similarity.sh</p> <p>The source code of test methods of each project in the <strong><em>Data/LTM/TestMethods</em></strong> is parsed to generate pairs of test cases. This steps includes test methods tokenization, test methods embeddings extraction and similarity calculation. Then, all similarity scores are stored in <em><strong>Data/LTM/similarity_measurements</strong></em> folder. Due to the large size of the calculated similarity scores (60 GB), they were not uploaded on Zenodo, but they can be available upon request.</p> <p><strong>LTM's Test Suite Minimization:</strong></p> <p>The source code of this step is in the Code/LTM/Search directory.</p> <p><strong>Requirements:</strong></p> <p>To run this step, Python 3 is required (we used Python 3.10). Also, the required libraries in the <strong><em>Code/LTM/Search/requirements.txt</em></strong> file should be installed, as follows:</p> <p>cd Code/LTM/Search</p> <p>pip install -r requirements.txt</p> <p><strong>Input:</strong></p> <p>Data/LTM/similarity_measurements</p> <p><strong>Output:</strong></p> <p>Results/LTM/minimization_results</p> <p><strong>Running the experiments:</strong></p> <p>To minimize the test suite for each project version, the following bash script should be executed:</p> <p>bash minimize.sh</p> <p>The similarity scores of all test case pairs per project version are parsed by the search algorithm (Genetic Algorithm). Each experiment runs ten times using three minimization budgets (25%, 50%, and 75%). The results are stored in the <em><strong>Results/LTM/minimization_results</strong></em> directory.</p> <p><strong>LTM's Evaluation:</strong></p> <p>To evaluate the minimization results for each version and each project, the following bash script should be executed:</p> <p>cd Code/LTM/Evaluation</p> <p>bash evaluate_per_version.sh</p> <p>cd Code/LTM/Evaluation</p> <p>bash evaluate_per_project.sh</p> <p>This will evaluate the FDR, MT and TSR results for each version and each project for each minimization budget. These results are stored in the <em><strong>Results/LTM</strong></em> directory.</p> <p>Note that for each version, the FDR is either 1 or 0. For each project, the FDR ranges from 0 to 1.</p>
LLMs4OL 2024 Datasets: Toward Ontology Learning with Large Language Models
<p>Ontology learning (OL) from unstructured data has evolved significantly, with recent advancements integrating large language models (LLMs) to enhance various aspects of the process. The LLMs4OL 2024 datasets, were developed to benchmark and advance research in OL using LLMs. This dataset as a key component of the LLMs4OL Challenge, targets three primary OL tasks: Term Typing, Taxonomy Discovery, and Non-Taxonomic Relation Extraction. It encompasses seven domains, i.e. lexosemantics and biological functions, offering a comprehensive resource for evaluating LLM-based OL approaches Each task within the dataset is carefully crafted to facilitate both Few-Shot (FS) and Zero-Shot (ZS) evaluation scenarios, allowing for robust assessment of model performance across different knowledge domains to address a critical gap in the field by offering standardized benchmarks for fair comparison for evaluating LLM applications in OL. </p>
Geoparsing with Large Language Models: Leveraging the linguistic capabilities of generative AI to improve geographic information extraction
<h2>Geoparsing with Large Language Models</h2> <p>The .zip file included in this repository contains all the code and data required to reproduce the results from our paper. Note, however, that in order to run the OpenAI models, users will required an OpenAI API key and sufficient API credits.</p> <div> <h3>Data</h3> <p>The data used for the paper are in the <code>datasetst</code> and <code>results</code> folders.</p> <ul> <li> <p>**Datasets: **This contains the XML files (LGL and Geovirus) and Json files (News2024) used to benchmark the models. It also contains all the data used to fine-tune the gpt-3.5 model, the prompt templates sent to the LLMs, and other data used for mapping and data creation.</p> </li> <li> <p>**Results: **This contains the results for the models on the three datastes. The folder is separated by dataset, with a single <code>.csv</code> file giving the results for each model on each dataset separately. The <code>.csv</code> file is structured so that each row contains either a predicted toponym and an associated true toponym (along with assigned spatial coordinates), if the model correctly identified a toponym; otherwise the true toponym columns are empty for false positives and the predicted columns are empty for false negatives.</p> </li> </ul> <h3>Code</h3> <p>The code is split into two seperate folders <code>gpt_geoparser</code> and <code>notebooks</code>.</p> <ul> <li>**GPT_Geoparser: **this contains the classes and methods used process the XML and JSON articles (<code>data.py</code>), interact with the Nominatim API for geocoding (<code>gazetteer.py</code>), interact with the OpenAI API (<code>gpt_handler.py</code>), process the outputs from the GPT models (<code>geoparser.py</code>) and analyse the results (<code>analysis.py</code>).</li> <li><strong>Notebooks</strong>: This series of notebooks can be used to reproduce the results given in the paper. The file names a reasonably descriptive of what they do within the context of the paper.</li> </ul> <h3>Code/software</h3> <h3>Requirements</h3> <ul> <li>Numpy</li> <li>Pandas</li> <li>Geopy</li> <li>Scitkit-learn</li> <li>lxml</li> <li>openai</li> <li>matplotlib</li> <li>Contextily</li> <li>Shapely</li> <li>Geopandas</li> <li>tqdm</li> <li>huggingface_hub</li> <li>Gnews</li> </ul> <h3>Access information</h3> <p>Other publicly accessible locations of the data:</p> <ul> <li>The LGL and GeoVirus datasets can also be obtained <a href="https://github.com/milangritta/Pragmatic-Guide-to-Geoparsing-Evaluation" target="_blank" rel="noopener">here<span> (opens in new window)</span></a>.</li> </ul> <h3>Abstract</h3> <div> <p>Geoparsing- the process of associating textual data with geographic locations - is a key challenge in natural language processing. The often ambiguous and complex nature of geospatial language make geoparsing a difficult task, requiring sophisticated language modelling techniques. Recent developments in Large Language Models (LLMs) have demonstrated their impressive capability in natural language modelling, suggesting suitability to a wide range of complex linguistic tasks. In this paper, we evaluate the performance of four LLMs - GPT-3.5, GPT-4o, Llama-3.1-8b and Gemma-2-9b - in geographic information extraction by testing them on three geoparsing benchmark datasets: GeoVirus, LGL, and a novel dataset, News2024, composed of geotagged news articles published outside the models' training window. We demonstrate that, through techniques such as fine-tuning and retrieval-augmented generation, LLMs significantly outperform existing geoparsing models. The best performing models achieve a toponym extraction F1 score of 0.985 and toponym resolution accuracy within 161 km of 0.921. Additionally, we show that the spatial information encoded within the embedding space of these models may explain their strong performance in geographic information extraction. Finally, we discuss the spatial biases inherent in the models' predictions and emphasize the need for caution when applying these techniques in certain contexts.</p> </div> <h3>Methods</h3> <div> <p>This contains the data and codes required to reproduce the results from our paper. The LGL and GeoVirus datasets are pre-existing datasets, with references given in the manuscript. The News2024 dataset was constructed specifically for the paper. </p> <p>To construct the News2024 dataset, we first created a list of 50 cities from around the world which have population greater than 1000000. We then used the GNews python package <a href="https://pypi.org/project/gnews/" target="_blank" rel="noopener">https://pypi.org/project/gnews/<span> (opens in new window)</span></a> to find a news article for each location, published between 2024-05-01 and 2024-06-30 (inclusive). Of these articles, 47 were found to contain toponyms, with the three rejected articles referring to businesses which share a name with a city, and which did not otherwise mention any place names.</p> <p>We used a semi autonmous approach to geotagging the articles. The articles were first processed using a Distil-BERT model, fine tuned for named entity recognicion. This provided a first estimate of the toponyms within the text. A human reviewer then read the articles, and accepted or rejected the machine tags, and added any tags missing from the machine tagging process. We then used OpenStreetMap to obtain geographic coordinates for the location, and to identify the toponym type (e.g. city, town, village, river etc). We also flagged if the toponym was acting as a geo-political entity, as these were reomved from the analysis process. In total, 534 toponyms were identified in the 47 news articles. </p> </div> </div>
Do Current Language Models Support Code Intelligence for R Programming Language?
<p>This is the dataset used in the paper: Do Current Language Models Support Code Intelligence for Programming Language?</p> <p> </p> <p>This dataset contains code snippets from R programming language repositories on GitHub, paired with their corresponding natural language (NL) descriptions. It was created for research in software engineering tasks like code summarization and code search. The data was collected using the GitHub REST API and includes over 1,500 public R repositories. To ensure quality, only active, well-structured R packages with proper documentation were included. Roxygen2, a popular documentation framework, was used to extract both the code and its matching NL descriptions.</p> <p>The dataset is organized into three parts: base R functions (Base), functions from the tidyverse (Tidy), and a combined set (RCombine). The dataset follows the CodeSearchNet format, with a split for training, validation, and testing data, ensuring no duplicate functions.</p>
Dataset from TableLabler: Scalable Labeling of Data Tables with Language Models for Tabular Dataset Creation [Scalable Data Science]
<p>TableLabler: Scalable Labeling of Data Tables with Language Models for Tabular Dataset Creation [Scalable Data Science]</p> <p>Pre-publicatoin upload for VLDB review.</p> <p>Code is at https://github.com/RelationalAI/annotated-tables/. The Github repository name is from a previous draft version and cannot be changed. It is code for TableLabler.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.