Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

95

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

95 results for “Sentence”

Learn how ShareScore rates datasets ↗
OpenNeuro48/100

L2_SENTENCES

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo48/100

The Audio Database of Hatoma Example Sentences

<p>This is a set of sound files of Hatoma Language, Southern Ryukyuan, spoken on Hatoma island, Okinawa.</p> <p>The database has 37611 sentences in Hatoma, included in Hatoma-Japanese Dictionary.</p> <p>See the audio database of Hatoma lexicon (Southern Ryukyuan) for a set of lexicons. https://doi.org/10.5281/zenodo.4560935</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

CrowdTruth Corpus for Open Domain Relation Extraction from Sentences

<p>This repository contains a ground truth corpus for open domain relation extraction from sentences, acquired with crowdsourcing and processed with <strong><a href="http://crowdtruth.org/">CrowdTruth</a></strong> metrics that capture ambiguity in annotations by measuring inter-annotator disagreement.</p> <p>The dataset contains annotations for 4,100 sentences sampled from Angeli et al. (1) and Riedel et al. (2), over 16 relations, with each sentence annotated by 15 workers. The sentences have been pre-processed with Distant Supervision (3) using the Freebase knowledge base, in order to identify the term pairs in each sentence that are likely to express a relation. The crowdsourced data was collected from <a href="http://figure-eight.com/">Figure Eight</a> and <a href="https://www.mturk.com/">Amazon Mechanical Turk</a>.</p> <p>This corpus has been discussed in the following papers:</p> <ul> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1809.00537">Crowdsourcing Semantic Label Propagation in Relation Classification</a></strong>. <a href="http://fever.ai/">FEVER</a> Workshop at <a href="http://emnlp2018.org/">EMNLP 2018</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1711.05186">False Positive and Cross-relation Signals in Distant Supervision Data</a></strong>. <a href="http://www.akbc.ws/">AKBC</a> Workshop at <a href="http://nips.cc/">NIPS 2017</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="http://crowdtruth.org/wp-content/uploads/2017/03/collint17-open-domain.pdf">Disagreement in Crowdsourcing and Active Learning for Better Distant Supervision Quality</a></strong>. <a href="http://collectiveintelligenceconference.org/">Collective Intelligence 2017</a>.</li> </ul> <p>Sentence-level data is available in file: <code>|--data/output/aggregated_sentences.csv</code></p> <p>Worker-level data is available in file: <code>|--data/output/aggregated_workers.csv</code></p> <p>Raw crowdsourcig data is available in folder: <code>|--data/input/</code></p> <p>Results of the relation classification model are available in folder: <code>|--data/model_results/</code></p> <p>&nbsp;</p> <p>References</p> <p>(1) Angeli, Gabor, et al. &quot;Combining distant and partial supervision for relation extraction.&quot; Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 2014.</p> <p>(2) Riedel, Sebastian, et al. &quot;Relation extraction with matrix factorization and universal schemas.&quot; Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2013.</p> <p>(3) Mintz, Mike, et al. &quot;Distant supervision for relation extraction without labeled data.&quot; Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2-Volume 2. Association for Computational Linguistics, 2009.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo44/100

Children speech recording (English, spontaneous speech + pre-defined sentences)

<p>The dataset contains audio recordings (lossless WAV) of 11 young children (age M=4.9 years old; 5 females, 6 males).</p> <p>Recordings include:</p> <ul> <li>free speech (retelling a picture book, ‘Frog, Where Are You?’ by Mercer Mayer)</li> <li>repeating 5 pre-defined short sentences (like 'the horse is in the stable')</li> <li>telling the numbers from 1 to 10</li> </ul> <p>The recordings are in English and the participants include both native and non-native speakers.</p> <p>Each sample is recorded from 3 sources:</p> <ul> <li>A studio-grade microphone (Rode NT1-A)</li> <li>A portable microphone (Zoom H1)</li> <li>The two front microphones of the Aldebaran NAO robot</li> </ul> <p>(note that, due to technical issues, a few (sample/microphone) combinations are missing).</p> <p> </p> <p>For the free-speech recording, a manual segmentation of the utterances is provided as well.</p>

opencc-by-4.0Dec 2016View details →
zenodo44/100

Webis-Simple-Sentences-17 Corpus

<p>A corpus of 471,085,690 English sentences extracted from the ClueWeb12 Web Crawl. The sentences were sampled from a larger corpus to achieve a level of sentence complexity similar to the one of sentences that humans make up as a memory aid for remembering passwords. Sentence complexity was determined by syllables per word.</p> <p>The corpus is split in training and test set as it is used in the associated publication.&nbsp; The test set is extracted from part 00 of the ClueWeb12, while the training set is extracted from the other parts.</p> <p>More information on the corpus can be found on the corpus web page at our university (listed under documented by).</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

MediCause Dataset of Causal Sentences with Annotated Entities

<p>The MediCause dataset contains 1202 causal sentences from medical publications where the entities involved in the causal relations have been annotated according to the MediCause ontological model for causal relations. The entities are annotated according to the Inside-Outside-Beginning (IOB) format. The labels used for the annotation are B-C (Cause), B-VC (Causal Variable), B-CS (Beginning Causal Specifier), I-CS (Inside Causal Specifier), B-CON (Beginning Connective), I-CON (Inside Connective), B-EF (Effect), B-VE (Effect Variable), B-ES (Beginning Effect Specifier), I-ES (nside Effect Specifier), O (Outside).</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

MaSS - Multilingual corpus of Sentence-aligned Spoken utterances

<p><strong>Abstract</strong></p> <p>The CMU Wilderness Multilingual Speech Dataset is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for potentially 700 languages. However, the fact that the source content (the Bible), is the same for all the languages is not exploited to date. Therefore, this article proposes to add multilingual links between speech segments in different languages, and shares a large and clean dataset of 8,130 para-lel spoken utterances across 8 languages (56 language pairs).We name this corpus MaSS (Multilingual corpus of Sentence-aligned Spoken utterances). The covered languages (Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish) allow researches on speech-to-speech alignment as well as on translation for syntactically divergent language pairs. The quality of the final corpus is attested by human evaluation performed on a corpus subset (100 utterances, 8 language pairs).</p> <p><a href="https://arxiv.org/pdf/1907.12895.pdf">Paper </a>| <a href="https://github.com/getalp/mass-dataset">GitHub Repository</a>&nbsp;containing&nbsp;the scripts needed to build the data set from scratch (if needed)</p> <p><strong>Project structure</strong></p> <p>This repository contains 8 Numpy files, one for each featured language, pickled with Python 3.6. Each line corresponds to the spectrogram of the file mentioned in the file <em>verses.csv</em>. There is a direct mapping between the ID of the verse and its index in the list (thus verse with ID 5634 is located at index 5634 in the Numpy file). Verses not available for a given language (as stated by the value &quot;Not Available&quot; in the CSV file) are represented by empty lists in the Numpy files, thus ensuring a perfect verse-to-verse alignement between each file.</p> <p>Spectrogram were extracted using Librosa with the following parameters:</p> <pre><code>Pre-emphasis = 0.97 Sample rate = 16000 Window size = 0.025 Window stride = 0.01 Window type = 'hamming' Mel coefficients = 40 Min frequency = 20</code></pre> <p>&nbsp;</p>

openmit-licenseJul 2019View details →
zenodo44/100

Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval

<p>This submission includes all pretrained models, test data and prediction files&nbsp;for the EMNLP 2019 paper &quot;Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval&quot;. Please follow the instructions in the emnlp bran&nbsp;at the&nbsp;<a href="https://github.com/castorini/birch/tree/emnlp">Birch repo</a>&nbsp;to reproduce the results.</p>

opencc-by-4.0Aug 2019View details →
zenodo44/100

Shakespeare: Julius Caesar 1.2.30-187, sentences with Maximum Similarty Score. Appendix Table for "Innovation and Repetiton in Dramatic Texts"

<p>This is an illustrative table for the study "Innovation and Repetition in Dramatic Texts", published in the journal JCLS by Botond Szemes and Mih&aacute;ly Nagy. The table contains a scene from <em>Julius Caesar</em> 1.2.30-187' , assigning to each utterance the most similar sentence and the degree of similarity based on an S-BERT model.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Similar Sentences from English Wikipedia 20150304

<p>Similar sentences from a dump of enwiki from March 04, 2015 detected using the Wikiduper software (https://github.com/seweissman/wikiduper).</p>

opencc-by-4.0Jun 2015View details →
zenodo40/100

A crowdsourced sentence-bound chemical-induced disease relationship corpus

<p>A set of 3000 abstracts from PubMed were annotated for sentence-bound chemical-induced disease relationships in order to train a machine learning algorithm for the BioCreative V challenge.</p>

opencc-zeroAug 2015View details →
zenodo40/100

Supplementary material to: Russian verbal aspect and the activation of event knowledge: Processing typical and atypical location adverbials in perfective and imperfective sentences

<p>Supplementary material for a self-paced reading experiment:</p> <ol> <li>Material: contains the verbal stimuli (experimental and filler sentences)</li> <li>raw_data.zip: E-Prime output for all 50 participants (tab-separated txt-files)</li> <li>spr_aspect_respinf.txt: Basic information on the participants; tab-separated txt-file</li> <li>spr_aspect_preprocessing_subm2_fin.R: data preprocessing and calculation of correct responses per speaker in R; R script<br>Input: <br>- Files from the "raw_data"-folder<br>- spr_aspect_respinf.txt<br>Output: <br>- spr_aspect_respinf_corr_resp.txt: same as spr_aspect_respinf.txt + number of correct responses per participant; tab-separated txt-file<br>- spr_aspect_all.txt: relevant data from all participants in one file; tab-separated txt-file<br>- spr_aspect_all_without-outlier.txt: same as spr_aspect_all.txt, but outliers set to NA; tab-separated txt-file</li> <li>spr_aspect_figures-stats_subm2_fin.R: plots figures and calculates statistics (descriptive statistics and GLMM) in R, R script<br>Input: <br>- spr_aspect_all_without-outlier.txt<br>Output:<br>- spr_aspect_avg.txt: Mean RT and SD per condition; tab-separated txt-file<br>- Figures</li> </ol>

opencc-by-4.0Dec 2024View details →
dryad40/100

[Stimulus Set] Evoking the N400 event-related potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)

<p>During speech comprehension, the ongoing context of a sentence is used to predict sentence outcome by limiting subsequent word likelihood. Neurophysiologically, violations of context-dependent predictions result in amplitude modulations of the N400 event-related potential (ERP) component. While N400 is widely used to measure semantic processing and integration, no publicly-available auditory stimulus set is available to standardize approaches across the field. Here, we developed an auditory stimulus set of 442 sentences that utilized the semantic anomaly paradigm, provided cloze probability for all stimuli, and was developed for both children and adults. With 20 neurotypical adults, we validated that this set elicits robust N400's, as well as two additional semantically-related ERP components: the recognition potential (~250 ms) and the late positivity component (~600 ms). This stimulus set (<a href="https://doi.org/10.5061/dryad.9ghx3ffkg">https://doi.org/10.5061/dryad.9ghx3ffkg</a>) and the 20 high-density (128-channel) electrophysiological datasets (<a href="https://doi.org/10.5061/dryad.6wwpzgmx4">https://doi.org/10.5061/dryad.6wwpzgmx4</a>) are made publicly available to promote data sharing and reuse. Future studies that use this stimulus set to investigate sentential semantic comprehension in both control and clinical populations may benefit from the increased comparability and reproducibility within this field of research.</p>

opencc-zeroMay 2022View details →
zenodo40/100

Extract from the PEAPL framework : Modelling "Writing sentences" competency (French as the schooling language)

<p>Competences, skills and knowledges that make up &quot;writing sentences&quot; competency, based on the linguistic praxeological organization of French as the schooling language (https://doi.org/10.5281/zenodo.4001381). This is an extract of the general framework, some of the visible objects are linked to other objects in other main competences. Orange links show how pedagogic ressources (game levels) are linked to framework objects.</p> <p>This framework is used for the PEAPL (peapl.eu)&nbsp;project&nbsp;to link activities in the GamesHub platform (for example activities using &quot;L&#39;Orthodyss&eacute;e des Gram&quot; grammar online game, available at&nbsp;https://www.lafamillegram.ch/#)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Mandarin matrix sentence test recordings: Lombard and plain speech with different speakers

<p>This dataset was recorded within the Deutsche Forschungsgemeinschaft (DFG) project: Experiments and models of speech recognition across tonal and non-tonal language systems (EMSATON, Projektnummer 415895050).</p> <p>The Lombard effect or Lombard reflex is the involuntary tendency of speakers to increase their vocal effort when speaking in loud noise to enhance the audibility of their voice. Up to date, the phenomena of Lombard effects were observed in different languages. The present database aimed at providing recordings for studying the Lombard effect with Mandarin speech.</p> <p>Eleven native-Mandarin talkers (6 female and 5 male) were recruited, both Lombard/plain speech were recorded from the same talker in the same day.&nbsp;</p> <p>All speakers produced fluent standard Mandarin speech (North China). All listeners were normal-hearing with pure tone thresholds of 20 dB hearing level or better at audiometric octave frequencies between 125 and 8000 Hz. All listeners provided written informed consent, approved by the Ethics Committee of Carl von Ossietzky University of Oldenburg.&nbsp;Listeners received an hourly compensation for their participation.</p> <p>The recording sentences were same as the official Mandarin Chinese matrix sentence test (<a href="#_ENREF_2">CMNmatrix, Hu et al. 2018</a>).&nbsp;</p> <p>&nbsp;One hundred sentences (ten base lists of ten sentences) of the CMNmatrix were recorded from each speaker in both plain and Lombard speaking styles (each base list containing all 50 words). The 100 sentences were divided into 10 blocks of 10 sentences each, and the plain and Lombard blocks were presented in an alternating order. The recording took place in a double-walled, sound-attenuated booth fulfilling ISO 8253-3 (ISO 8253-3, 2012), using a Neumann 184 microphone with a cardioid characteristic (Georg Neumann GmbH, Berlin, Germany) and a Fireface UC soundcard (with a sampling rate of 44100 Hz and resolution of 16 bits). The recording procedure generally followed the procedures of Alghamdi et al. (2018). A Mandarin-native speaker and a phonetician participated in the recording session and listened to the sentences to control the pronunciations, intonation, and speaking rate. During the recording, the speaker was instructed to read the sentence presented on a frontal screen. In case of any mispronunciation or change in the intonation, the speaker was asked via the screen to repeat the sentence again, and on average, each sentence was recorded twice. In Lombard conditions the speaker was regularly asked via a prompt to repeat a sentence, to keep the speaker in the Lombard communication situation. For the plain-speech recording blocks, the speakers were asked to pronounce the sentences with natural intonation and accentuation, and at an intermediate speaking rate, which was facilitated by a progress bar on the screen. Furthermore, the speakers were asked to keep the speaking effort constant and to avoid any exaggerated pronunciations that could lead to unnatural speech cues. For the Lombard speech recording blocks, speakers were instructed to imagine a conversation to another person in a pub-like situation. During the whole recording session, speakers wore headphones (Sennheiser HDA200) that provided the audio signal of the speaker.. In the Lombard condition, the stationary speech-shaped noise ICRA1 (Dreschler et al., 2001) was mixed with the speaker&rsquo;s audio signal at a level of 80 dB SPL (calibrated with a Br&uuml;el &amp; Kj&aelig;r (B&amp;K) 4153 artificial ear, a B&amp;K 4134 0.5-inch inch microphone, a B&amp;K 2669 preamplifier, and a B&amp;K 2610). Previous studies showed that this level induced a robust Lombard speech without the danger of inducing hearing damage (Alghamdi et al., 2018).</p> <p>The sentences were cut from the recording, high-pass filtered (60 Hz cut-off frequency) and set to the average root-mean-square level of the original speech material of the Mandarin Matrix test (Hu et al., 2018). Then the best version of each sentence was chosen by native-Mandarin speakers regarding pronunciation, tempo, and intonation.</p> <p>For more detailed information, please contact&nbsp;hongmei.hu@uni-oldenburg,&nbsp; sabine.hochmuth@uni-oldenburg.de.&nbsp;</p> <p><em>Hu H, Xi X, Wong LLN, Hochmuth S, Warzybok A, Kollmeier B (2018) Construction and evaluation of the mandarin chinese matrix (cmnmatrix) sentence test for the assessment of speech recognition in noise. International Journal of Audiology 57:838-850. https://doi.org/10.1080/14992027.2018.1483083</em></p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
dryad40/100

Evoking the N400 Event-Related Potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)

<p class="MsoNoSpacing"><span>During speech comprehension, the ongoing context of a sentence is used to predict sentence outcome by limiting subsequent word likelihood. Neurophysiologically, violations of context-dependent predictions result in amplitude modulations of the N400 event-related potential (ERP) component. While N400 is widely used to measure semantic processing and integration, </span><span>no publicly-available auditory stimulus set is available to standardize approaches across the field. Here, we developed an auditory stimulus set of 442 sentences that utilized the semantic anomaly paradigm, provided cloze probability for all stimuli, and was developed for both children and adults. With 20 neurotypical adults, we validated that this set elicits robust N400's, as well as two additional semantically-related ERP components: the recognition potential (~250 ms) and the late positivity component (~600 ms). This stimulus set (<a href="https://doi.org/10.5061/dryad.9ghx3ffkg">https://doi.org/10.5061/dryad.9ghx3ffkg</a>) and the 20 high-density (128-channel) electrophysiological datasets (<a href="https://doi.org/10.5061/dryad.6wwpzgmx4">https://doi.org/10.5061/dryad.6wwpzgmx4</a>) </span><span>are made publicly available to promote data sharing and reuse. Future studies that use this stimulus set to investigate sentential semantic comprehension in both control and clinical populations may benefit from the increased comparability and reproducibility within this field of research.</span></p>

opencc-zeroOct 2022View details →
zenodo40/100

The SICK (Sentences Involving Compositional Knowledge) dataset for relatedness and entailment

<p>The SICK data set consists of about 10,000 English sentence pairs, generated starting from two existing sets: the&nbsp;<a href="http://nlp.cs.illinois.edu/HockenmaierGroup/data.html">8K ImageFlickr data set</a>&nbsp;and the&nbsp;<a href="http://www.cs.york.ac.uk/semeval-2012/task6/index.php?id=data">SemEval 2012 STS MSR-Video Description data set</a>. We randomly selected a subset of sentence pairs from each of these sources and we applied a 3-step generation process: first, the original sentences were normalized to remove unwanted linguistic phenomena; the normalized sentences were then expanded to obtain up to three new sentences with specific characteristics suitable to CDSM evaluation; as a last step, all the sentences generated in the expansion phase were paired with the normalized sentences in order to obtain the final data set.</p> <p>Each sentence pair was annotated for relatedness and entailment by means of crowdsourcing techniques. The&nbsp;<strong>sentence relatedness score</strong>&nbsp;(on a 5-point rating scale) provides a direct way to evaluate CDSMs, insofar as their outputs are meant to quantify the degree of semantic relatedness between sentences; the categorizations in terms of the&nbsp;<strong>entailment relation between the two sentences</strong>&nbsp;(with&nbsp;<em>entailment, contradiction</em>, and&nbsp;<em>neutral</em>&nbsp;as gold labels) is also a crucial aspect to consider, since detecting the presence of entailment is one of the traditional benchmarks of a successful semantic system.</p> <p>In the final set, gold scores for relatedness and entailment were distributed as follows: the relatednes scoring resulted in 923 pairs within the [1,2) range, 1373 pairs within the [2,3) range, 3872 pairs within the [3,4) range, and 3672 pairs within the [4,5] range; the entailment annotation led to 5595&nbsp;<em>neutral</em>&nbsp;pairs, 1424&nbsp;<em>contradiction</em>&nbsp;pairs, and 2821&nbsp;<em>entailment</em>&nbsp;pairs.</p> <p><strong>Files</strong></p> <ul> <li>SICK.zip (main file)</li> <li>SICK_Annotated.zip (a&nbsp;version of the data set annotated for the expansion rule which was used in each case)</li> <li>SICK_subsets.zip (a&nbsp;Indexes specifying further classifications, used in the JLRE 2016 publication)</li> </ul> <p>&nbsp;</p>

opencc-by-nc-sa-3.0May 2014View details →
zenodo40/100

Figure 1: The ECG model-MAPPING BETWEEN SEMANTIC GRAPHS AND SENTENCES IN GRAMMAR INDUCTION SYSTEM

<p>The following Figure 1 shows a sample semantic graph that describes a<br> simple test world.<br> During the processing of the ECG, the base units of the graph are the ECG<br> atoms. An ECG atom corresponds to a primitive statements related to one<br> predicate. It has a structure of one-level deep tree, where the root of the tree<br> is the predicate and the concepts linked to it are the leaves. The child concept<br> of the root predicate may be not only a single concept but it can be another<br> ECG atom.</p>

opencc-by-4.0Jun 2010View details →
zenodo40/100

Figure 7. The mental spaces set up by the sentence If John buys the car, he will drive to Berlin.-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>In the sentence (4) there are three proper names which constitute elements of the base space<br> or reality space (R). The CNSB if sets up the hypothetical space (H) with elements identical to those<br> of reality space (Figure 7). Every space&rsquo;s internal structure is presented in the boxes next to them.<br> In the next part, we will see how MultiNet&rsquo;s built-in meaning representation mechanisms are<br> capable of representing the basic principles of mental space building outlined above.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 6. The mental spaces set up by the sentence Mary thinks that John smokes.-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>The proper nouns Mary and John setup a base space (B). By the help of background<br> knowledge and activated frames we know that they are names of female and male humans. Not<br> having access to the previous discourse, we also consider their existence presupposed. The SVSB<br> Mary thinks that sets up a belief space (L) relative to space B (Figure 6). The identity connector<br> maintains the referential link between elements a and a ׳ both referring to the same person.</p>

opencc-by-4.0Dec 2010View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record