Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
859
datasets available to search
ShareScore release 0.9.0
Dataset results
859 results for “Speeches”
Segmented DAPS (Device and Produced Speech) Dataset
<p>This is a modified version of a subset of the Device and Produced Speech (DAPS) dataset. The original dataset can be found <a href="https://zenodo.org/record/4660670#.YKuxgKhKhPZ">here</a>. This dataset contains text-aligned audio of the first script of the "clean" partition of the DAPS dataset for all 20 speakers. Phoneme and word alignments are provided as JSON files. We segment the audio and alignments into single sentences. For each sentence, we additionally provide the raw text in a txt file. Audio is provided as 44.1 kHz WAV files.</p> <p>If you use this work as part of an academic publication, please cite the paper corresponding to the original dataset:</p> <blockquote> <p>Gautham J. Mysore, <a href="https://ieeexplore.ieee.org/document/6981922">“Can We Automatically Transform Speech Recorded on Common Consumer Devices in Real-World Environments into Professional Production Quality Speech? - A Dataset, Insights, and Challenges”</a>, in the IEEE Signal Processing Letters, Vol. 22, No. 8, August 2015</p> </blockquote>
Enhanced RAVDESS Speech Dataset
<p>This is a modified version of the speech audio contained within the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) dataset. The original dataset can be found <a href="https://zenodo.org/record/1188976#.YKvCMKhKhPY">here</a>. The unmodified version of just the speech audio used as source material for this dataset can be found <a href="https://www.kaggle.com/uwrfkaggler/ravdess-emotional-speech-audio">here</a>. This dataset performs speech enhancement and bandwidth extension on the original speech using HiFi-GAN. HiFi-GAN produces high-quality speech at 48 kHz that contains significantly less noise and reverb relative to the original recordings.</p> <p>If you use this work as part of an academic publication, please cite the papers corresponding to both the original dataset as well as HiFi-GAN:</p> <blockquote> <p>Livingstone SR, Russo FA (2018) The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English. PLoS ONE 13(5): e0196391. <a href="https://doi.org/10.1371/journal.pone.0196391">https://doi.org/10.1371/journal.pone.0196391</a>.</p> <p>Su, Jiaqi, Zeyu Jin, and Adam Finkelstein. "HiFi-GAN: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks." <em>Proc. Interspeech</em>. October 2020.</p> </blockquote> <p>Note that there are two recent papers with the name "HiFi-GAN". Please be sure to cite the correct paper as listed here.</p>
Understanding degraded speech leads to perceptual gating of a brainstem reflex in human listeners
<p>The ability to navigate "cocktail-party" situations by focussing on sounds of interest over irrelevant, background sounds is often considered in terms of cortical mechanisms. However, subcortical circuits such as the pathway underlying the medial olivocochlear (MOC) reflex modulate the activity of the inner ear itself, supporting the extraction of salient features from auditory scene prior to any cortical processing. To understand the contribution of auditory subcortical nuclei and the cochlea in complex listening tasks, we made physiological recordings along the auditory pathway while listeners engaged in detecting non(sense)-words in lists of words. Both naturally spoken and intrinsically noisy, vocoded speech—filtering that mimics processing by a cochlear implant—significantly activated the MOC reflex, but this was not the case for speech in background noise, which more engaged midbrain and cortical resources. A model of the initial stages of auditory processing reproduced specific effects of each form of speech degradation, providing a rationale for goal-directed gating of the MOC reflex based on enhancing the representation of the energy envelope of the acoustic waveform. Our data reveals the co-existence of two strategies in the auditory system that may facilitate speech understanding in situations where the signal is either intrinsically degraded or masked by extrinsic acoustic energy. Whereas intrinsically degraded streams recruit the MOC reflex to improve representation of speech cues peripherally, extrinsically masked streams rely more on higher auditory centres to de-noise signals.</p>
FestCat speech synthesis dataset in Catalan, raw at 48kHz, 16bits/sample
<p>This dataset contains audio recordings from the 10 speakers of the FestCat project and an additional speaker "uri".</p> <p>The recordings are available in the "-raw" compressed archives and are provided in the following format:</p> <pre><code class="language-bash">SAMPLERATE=48000 # Hz BITSPERSAMPLE=16 ENCODING="SIGNED_INTEGER" AUDIOHEADER="RAW" # (no header) # For the prompts and utts: TEXT_ENCODING="ISO-8859-15"</code></pre> <p>These sampling conditions were downsampled from the original recordings captured at 96kHz and using 24 bits/sample for all the FestCat speakers. The additional speaker "uri" had recordings 16kHz and were here upsampled to 48kHz.</p> <p>The recordings have been automatically segmented. The segmentation results are available at the "-utts" archives.</p> <p>The text prompts are available in the "-prompts" archive.</p> <p>The text information is encoded using ISO-8859-15.</p> <p> </p>
Limitations of Audiovisual Speech on Robots for Second Language Pronunciation Learning
<p>Additional material accompanying the HRI'23 publication found at <a href="https://doi.org/10.1145/3568162.3578633">https://doi.org/10.1145/3568162.3578633</a>.</p> <p>Simulations showing the audiovisual speech performed by Furhat during the second language tutoring experiment.</p> <p>Each video consists of the Furhat saying a Dutch word, followed by saying the Japanese translation twice.</p> <p>Videos M1-M10 correspond to the Matched condition, M11-M20 to Mismatched, and M21-M30 to the Computer-generated lip movement.</p> <p>In the experiment, participants interacted with a physical Furhat, but as it is difficult to make video recordings of the Furhat, simulations are shown here to allow for high-fidelity comparison between conditions.</p>
Binaural detection thresholds and audio quality of speech and music signals in complex acoustic environments
<p>Every-day acoustical environments are often complex, typically comprising one attended target sound in the presence of interfering sounds (e.g., disturbing conversations) and reverberation. Here we assessed binaural detection thresholds and (supra-threshold) binaural audio quality ratings of four distortions types: spectral ripples, non-linear saturation, intensity and spatial modifications applied to speech, guitar, and noise targets in such complex acoustic environments (CAEs). The target and (up to) two masker sounds were either co-located as if contained in a common audio stream, or were spatially separated as if originating from different sound sources. The amount of reverberation was systematically varied. Masker and reverberation had a significant effect on the distortion-detection thresholds of speech signals. Quality ratings were affected by reverberation, whereas the effect of maskers depended on the distortion. The results suggest that detection thresholds and quality ratings for distorted speech in anechoic conditions are also valid for rooms with mild reverberation, but not for moderate reverberation. Furthermore, for spectral ripples, a significant relationship between the listeners’ individual detection thresholds and quality ratings was found. The current results provide baseline data for detection thresholds and audio quality ratings of different distortions of a target sound in CAEs, supporting the future development of binaural auditory models.</p>
SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis
<p>This is the public release of the Samsung Open Mean Opinion Scores (SOMOS) dataset for the evaluation of neural text-to-speech (TTS) synthesis, which consists of audio files generated with a public domain voice from trained TTS models based on bibliography, and numbers assigned to each audio as quality (naturalness) evaluations by several crowdsourced listeners.<br><br><strong>Description</strong><br><br>The SOMOS dataset contains 20,000 synthetic utterances (wavs), 100 natural utterances and 374,955 naturalness evaluations (human-assigned scores in the range 1-5). The synthetic utterances are single-speaker, generated by training several Tacotron-like acoustic models and an LPCNet vocoder on the LJ Speech voice public dataset. 2,000 text sentences were synthesized, selected from Blizzard Challenge texts of years 2007-2016, the LJ Speech corpus as well as Wikipedia and general domain data from the Internet.<br>Naturalness evaluations were collected via crowdsourcing a listening test on Amazon Mechanical Turk in the US, GB and CA locales. The records of listening test participants (workers) are fully anonymized. Statistics on the reliability of the scores assigned by the workers are also included, generated through processing the scores and validation controls per submission page.</p> <p>To listen to audio samples of the dataset, please see <a href="https://innoetics.github.io/publications/somos-dataset/index.html">our Github page</a>.</p> <p>The dataset release comes with a carefully designed train-validation-test split (70%-15%-15%) with unseen systems, listeners and texts, which can be used for experimentation on MOS prediction.</p> <p><em>This version also contains the necessary resources to obtain the transcripts corresponding to all dataset audios.</em></p> <p><strong>Terms of use</strong></p> <ul> <li>The dataset may be used for <strong>research</strong> purposes only, for <strong>non-commercial</strong> purposes only, and may be distributed with the same terms.</li> <li>Every time you produce research that has used this dataset, please <strong>cite </strong>the dataset appropriately.</li> </ul> <p>Cite as:</p> <pre><code>@inproceedings{maniati22_interspeech, author={Georgia Maniati and Alexandra Vioni and Nikolaos Ellinas and Karolos Nikitaras and Konstantinos Klapsas and June Sig Sung and Gunu Jho and Aimilios Chalamandaris and Pirros Tsiakoulis}, title={{SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis}}, year=2022, booktitle={Proc. Interspeech 2022}, pages={2388--2392}, doi={10.21437/Interspeech.2022-10922} } </code></pre> <p><br><strong>References of resources & models used</strong></p> <p>Voice & synthesized texts:<br>K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017.</p> <p>Vocoder:<br>J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. ICASSP, 2019.<br>R. Vipperla, S. Park, K. Choo, S. Ishtiaq, K. Min, S. Bhattacharya, A. Mehrotra, A. G. C. P. Ramos, and N. D. Lane, “Bunched lpcnet: Vocoder for low-cost neural text-to-speech systems,” in Proc. Interspeech, 2020.</p> <p>Acoustic models:<br>N. Ellinas, G. Vamvoukakis, K. Markopoulos, A. Chalamandaris, G. Maniati, P. Kakoulidis, S. Raptis, J. S. Sung, H. Park, and P. Tsiakoulis, “High quality streaming speech synthesis with low, sentence-length-independent latency,” in Proc. Interspeech, 2020.<br>Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al., “Tacotron: Towards End-to-End Speech Synthesis,” in Proc. Interspeech, 2017.<br>J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al., “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” in Proc. ICASSP, 2018.<br>J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, “Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling,” arXiv preprint arXiv:2010.04301, 2020.<br>M. Honnibal and M. Johnson, “An Improved Non-monotonic Transition System for Dependency Parsing,” in Proc. EMNLP, 2015.<br>M. Dominguez, P. L. Rohrer, and J. Soler-Company, “PyToBI: A Toolkit for ToBI Labeling Under Python,” in Proc. Interspeech, 2019.<br>Y. Zou, S. Liu, X. Yin, H. Lin, C. Wang, H. Zhang, and Z. Ma, “Fine-grained prosody modeling in neural speech synthesis using ToBI representation,” in Proc. Interspeech, 2021.<br>K. Klapsas, N. Ellinas, J. S. Sung, H. Park, and S. Raptis, “WordLevel Style Control for Expressive, Non-attentive Speech Synthesis,” in Proc. SPECOM, 2021.<br>T. Raitio, R. Rasipuram, and D. Castellani, “Controllable neural text-to-speech synthesis using intuitive prosodic features,” in Proc. Interspeech, 2020.</p> <p>Synthesized texts from the Blizzard Challenges 2007, 2008, 2009, 2010, 2011, 2012, 2013, 2016:<br>M. Fraser and S. King, "The Blizzard Challenge 2007," in Proc. SSW6, 2007.<br>V. Karaiskos, S. King, R. A. Clark, and C. Mayo, "The Blizzard Challenge 2008," in Proc. Blizzard Challenge Workshop, 2008.<br>A. W. Black, S. King, and K. Tokuda, "The Blizzard Challenge 2009," in Proc. Blizzard Challenge, 2009.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2010," 2010.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2011," 2011.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2012," 2012.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2013," 2013.<br>S. King and V. Karaiskos, "The Blizzard Challenge 2016," 2016.</p> <p><strong>Contact</strong></p> <p>Alexandra Vioni - a.vioni@samsung.com</p> <ul> <li>If you have any questions or comments about the dataset, please feel free to write to us.</li> <li>We are interested in knowing if you find our dataset useful! If you use our dataset, please email us and tell us about your research.</li> </ul>
LJ Speech - Aligned IPA transcriptions
<p>Files:</p> <ul> <li> <p><code>grids.zip</code></p> <ul> <li>contains TextGrids for all audio files containing three tiers <code>words</code>, <code>phonemes</code> and <code>transcription</code> <ul> <li><code>words</code> contains the aligned normalized English words</li> <li><code>phonemes</code> contains IPA pronunciations transcribed using CMU dictionary which then were aligned with Montreal Forced Aligner. The pronunciations were then mapped from ARPAbet to IPA and duration marks were applied (without punctuation)</li> <li><code>transcription</code> contains unaligned phonemes including punctuation and word boundary labels (SIL0)</li> </ul> </li> </ul> </li> <li> <p><code>preview.png</code></p> <ul> <li>preview of the first TextGrid opened in Praat</li> </ul> </li> <li> <p><code>words-vocabulary.txt</code></p> <ul> <li>contains all words from tier <code>words</code></li> </ul> </li> <li> <p><code>phonemes-vocabulary.txt</code></p> <ul> <li>contains all phonemes from tier <code>phonemes</code></li> </ul> </li> <li> <p><code>transcription-vocabulary.txt</code></p> <ul> <li>contains all phonemes/punctuation from tier <code>transcription</code></li> </ul> </li> <li> <p><code>phonemes-durations.pdf</code></p> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code></li> </ul> </li> <li> <p><code>phonemes-durations-simple.pdf</code></p> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers are ignored</li> </ul> </li> <li> <p><code>pronunciations.dict</code></p> <ul> <li>contains the pronunciations for each word including punctuation and weights (occurrence)</li> </ul> </li> <li> <p><code>script.sh</code></p> <ul> <li>contains the script to reproduce all results</li> </ul> </li> </ul> <p>Phoneme duration marker:</p> <ul> <li><code>˘</code> -> [0, 20) percentile</li> <li><code>ˑ</code> -> [80, 90) percentile</li> <li><code>ː</code> -> [90, inf) percentile</li> </ul> <p>Silence marker:</p> <ul> <li><code>SIL0</code> -> no silence</li> <li><code>SIL1</code> -> [0, 33.33) percentile</li> <li><code>SIL2</code> -> [33.33, 66.66) percentile</li> <li><code>SIL3</code> -> [66.66, inf) percentile</li> </ul> <p> </p>
Raw data for manuscript Semantic context can mask intelligibility declines at above-conversational speech levels in normal-hearing listeners
<p>Raw data for the manuscript in doc file. <br> Copied from the Matlab .m file. used for the analysis.</p> <p>To be updated.</p> <p>For details, contact me at mfer@health.sdu.dk</p>
Supplementary code and data for the paper `From stage to page: language independent bootstrap measures of distinctiveness in fictional speech`
<p>The repository provides full data and processing / analysis pipeline for the paper <strong>'From stage to page: language independent bootstrap measures of distinctiveness in fictional speech</strong>'<br> <br> Rendered notebooks are also available through Github:</p> <p>1) <a href="https://github.com/perechen/difs-character-voices/blob/master/data/all_stars_clean.ipynb">Preparation, energy distance and exploration</a> (main)</p> <p>2) <a href="https://github.com/perechen/difs-character-voices/blob/master/03_analysis.md">Keyword curves & formal modeling</a></p> <p> </p> <p>- `00_dracor_get_data.R`. Script uses <a href="https://dracor.org/">DraCor</a> dedicated API to get texts spoken by characters</p> <p>- `01_distinctiveness_energy.ipynb` does the heavy lifting of data wrangling, cleaning and preprocessing, plus implements energy distance bootstrapping and does exploratory analysis</p> <p>- `02_logodds_curves.R` calculates keyword curves for characters<br> <br> - `03_analysis_and_models.R` explores keyword curves and does Bayesian models</p>
Observations from the distinct neural encoding of glimpsed and masked speech in multitalker situations
<p>These data underlie the figures contained in "Distinct neural encoding of glimpsed and masked speech in multitalker situations"</p>
vowel formant representation in human speech cortex
<p>Data accompanying the paper publication.</p>
Data from: Neural correlates of multisensory enhancement in audiovisual narrative speech perception: a fMRI investigation
<p>This fMRI study investigated the effect of seeing articulatory movements of a speaker while listening to a naturalistic narrative stimulus. It had the goal to identify regions of the language network showing multisensory enhancement under synchronous audiovisual conditions. We expected this enhancement to emerge in regions known to underlie the integration of auditory and visual information such as the posterior superior temporal gyrus as well as parts of the broader language network, including the semantic system. To this end we presented 53 participants with a continuous narration of a story in auditory alone, visual alone, and both synchronous and asynchronous audiovisual speech conditions while recording brain activity using BOLD fMRI. We found multisensory enhancement in an extensive network of regions underlying multisensory integration and parts of the semantic network as well as extralinguistic regions not usually associated with multisensory integration, namely the primary visual cortex and the bilateral amygdala. Analysis also revealed involvement of thalamic brain regions along the visual and auditory pathways more commonly associated with early sensory processing. We conclude that under natural listening conditions, multisensory enhancement not only involves sites of multisensory integration but many regions of the wider semantic network and includes regions associated with extralinguistic sensory, perceptual and cognitive processing.</p>
USPDATRO: Underrepresented Speech Dataset from Romanian language Open Data
<p> USPDATRO<br> ==========</p> <p>Underrepresented Speech Dataset from Open Data: Case Study on the Romanian Language (USPDATRO) is a manually created Romanian language speech corpus.<br> It was created specifically using speech types that are underrepresented in other speech datasets.<br> Sources for this dataset are represented by open data available on multimedia platforms under a Creative Commons license.<br> The data was manually transcribed and aligned at segment level.<br> In addition to the text and audio files, we offer text annotations (lemmatization, part of speech tags, dependency parsing) in CoNLL-U Plus format.</p> <p>Each datasource is mentioned by URL in the metadata.csv file with associated license (a Creative Commons variant).</p> <p>Dataset structure:<br> - audio: Folder with audio segments in WAV format<br> - text: Folder with corresponding transcriptions<br> - conllup: Folder with corresponding token-based annotations<br> - metadata.csv: Contains information about each segment</p> <p>LICENSING</p> <p>This work (transcriptions, alignment, metadata, annotations) is provided under the license CC BY-NC-SA 4.0 (Attribution-NonCommercial-ShareAlike 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-sa/4.0/<br> and the full text here: https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode .<br> The original works considered for audio sources are available under their respective licenses (Creative Commons variants) as described in the metadata.csv file.</p> <p><br> CONTACT</p> <p>Research Institute for Artificial Intelligence "Mihai Drăgănescu", Romanian Academy<br> Web: http://www.racai.ro<br> Contact emails: vasile@racai.ro</p>
Data and codes: Speech-recognition in landlide predictive modelling
<p>This is the data and codes for the manuscript "Speech-recognition in landlide predictive modelling"</p>
Dataset for Phonological acquisition depends on the timing of speech sound
<p>Dataset for the analysis to the manuscript: Phonological acquisition depends on the timing of<br> speech sounds: Deconvolution EEG modelling across the first five years</p>
Appendix - Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition
<p>Appendix tables for the paper "Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition".</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>
Defining and identifying Discourse Markers in spontaneous speech. A corpus-based and experimental proposal
<p>The paper has a twofold goal: (i) to identify the cathegory of Discourse Markers (DM) and their different specific functions; (ii) to validate the proposal with a perceptual experiment. First, we propose how to identify DM and their different functions in spontaneous speech. Both the identification of the categorycategory of DM and that of specific DM functions are based on prosodic criteria. In order to identify a DM, speech segmentation is crucial, since DMs necessarily appear isolated in a prosodic unit. Besides prosodic isolation, DMs do not feature pragmatic and prosodic autonomy, but depend on the illocutionary unit of the utterance. Than we show that a same lexical item can fulfill more functions, while the formal cues that convey the function are prosodic ones. The last part of the chapter is devoted to presenting the methodology and the results of a perceptual experiment that tests the theoretical hypothesis by asking the listeners to recognize three different functions by means of prosodic cues only.</p>
Defining and identifying Discourse Markers in spontaneous speech: a corpus-based and experimental proposal
<p>Audio files used as examples in the paper.</p>
Data from: Electroencephalography Responses to Simplified Visual Signals Reveal Explain Differences in Speech-in-Noise Comprehension
<p>Contents and Folder Structure:</p> <p><strong>EEG Experiment</strong></p> <ul> <li><strong>EEG_stimuli</strong>: these are the videos that were presented to participants in the EEG experiment, and the code that generates them from the original corpus (link)</li> <li><strong>data_2020>split_trials</strong>: contains the raw EEG data starting 1.995s before each trial and ending 1.995s after each trial with naming convention subXx_VV_YY_N.fif where X or Xx is the subject number, YY is the modality condition (AV for audiovisual and V0 for video only), N is the trial number (between 0 and 4 inclusive), and VV is the video condition (1e for the envelope dot, 1m is the mismatched dot, 4v is the cartoon, bw is the edge detection and nh is the natural condition). <ul> <li><strong>unprocessed>raw</strong>: contains the unprocessed raw EEG data <ul> <li><strong>processed>Fs-200>BP-1-80-ASR-INTP-AVR</strong>: contains the pre-processed raw EEG data: the output of <em>run_preprocessing.m</em></li> <li><strong>processed>Fs-200>BP-1-80-ASR-INTP-AVR-ICr</strong>: contains the pre-processed raw EEG data after ICA cleaning: the output of <em>run_reject_ICs.m</em></li> <li><strong>stim>stim_dwnspl</strong>: contains the aligned 200Hz envelopes of the presented speech used as features for the time-lagged models</li> </ul> </li> </ul> </li> <li><strong>EEG_analysis_code </strong>[note: please extract the contents of this folder to match paths] <ul> <li><strong>2_ICA_filt:</strong> this folder contains the MATLAB code that performs the pre-processing of the EEG data, including filtering, downsampling, ICA cleaning etc. The main functions are: <ul> <li><em>run_preprocessing.m: downsampling, filtering, ASR cleaning</em></li> <li><em>run_reject_ICs.m: ICLabel ICA cleaning</em></li> </ul> </li> <li><strong>3_analysis:</strong> this is the Python code that performs the TRF and backward modelling on the EEG data. The main functions are: <ul> <li><em>multisensory_bw.py</em>: backwards model</li> <li><em>multisensory_fw.py</em>: forwards model</li> </ul> </li> </ul> </li> </ul> <p><strong>Behavioural Experiment</strong></p> <ul> <li><strong>behavioural</strong> <ul> <li><strong>0_dataset</strong>: these are the videos that were presented to participants in the behavioural experiment, and the code that generates them from the original corpus (AV GRID corpus)</li> <li><strong>3_analysis</strong>: behavioural data analysis script <ul> <li>main function: <em>data_grid_v3.py</em></li> </ul> </li> </ul> </li> <li><strong>behavioural_data>data_grid</strong>: behavioural results </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.