Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

454

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

454 results for “Tau”

Learn how ShareScore rates datasets ↗
zenodo48/100

Molecular mechanism for the synchronized electrostatic coacervation and co-aggregation of alpha-synuclein and tau

<p><strong><em>The following metadata refers exclusively to electron paramagnetic resonance (EPR) measurements, which represent the contribution of the PARACAT students to this work</em></strong></p> <ul> <li><strong>Data type</strong>: EPR spectroscopic measurements and simulations</li> <li>Files are in <strong>.DTA, .DSC, .m, .mat, and .xlxs, </strong>formats</li> <li>Information on <strong>origin of the data</strong>: <ul> <li>EPR spectroscopic measurements in <strong>.DTA </strong>and<strong> .DSC</strong> formats</li> <li>EPR spectroscopic simulation and analyses in .<strong>m </strong>and<strong> .mat</strong> format</li> <li>&ldquo;Ready-to-plot&rdquo;, processed EPR spectra are in <strong>.xlxs</strong> format.</li> </ul> </li> <li>The data are <strong>generated</strong> by: <ul> <li>CW-EPR measurements were performed with a Bruker ELEXSYS E580 X-band spectrometer equipped with a Bruker ER4118 SPT-N1 resonator operating at a microwave (MW) frequency of &sim;9.7 GHz. The temperature was set to 25 &deg;C and controlled by a liquid nitrogen cryostat.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>If the dataset includes multiple files that relate to each other:</strong> <ul> <li>Files in <strong>PARACAT_WP3_20221219_EPR </strong>folder includes EPR spectroscopic measurements and computer simulations/analyses, original data are in <strong>&nbsp;.DTA/.DSC</strong> formats; files in .<strong>m</strong> format were used to process the data.</li> </ul> </li> </ul> <p>NB. See the &ldquo;READ ME&rdquo; text file for more detailed information on files organization.</p> <p>&nbsp;</p> <ul> <li><strong>Information on</strong>: <ul> <li>Abbreviations: <ul> <li><strong>avg</strong> = averaged</li> <li><strong>aS_24</strong> = alpha-synuclein protein with TEMPOL spin label at position 24 of the polypeptidic chain</li> <li><strong>aS_122</strong> = alpha-synuclein protein with TEMPOL spin label at position 122 of the polypeptidic chain</li> <li><strong>pLK</strong> = poly-lysine</li> <li><strong>Tau441</strong> = Tau protein with complete amino-acid sequence</li> <li><strong>Tau_DNt</strong> = truncated Tau protein lacking N-terminal (see paper methods for further details)</li> </ul> </li> <li>Units of measurement: <ul> <li>Temperature: <strong>&deg;</strong><strong>C</strong> (Celsius)</li> <li>Microwave Frequency: <strong>GHz</strong> (Giga-Hertz), <strong>MHz</strong> (Mega-Hertz), <strong>kHz</strong> (kilo-Hertz)</li> <li>Microwave Power: <strong>mW</strong> (milli-Watt)</li> <li>Magnetic Field: <strong>mT</strong> (milli-Tesla)</li> <li>Time: <strong>s</strong> (seconds)</li> <li>Concentration: <strong>&mu;M</strong> (micro-Molar), <strong>% w/v</strong> (percentage weight-volume)</li> </ul> </li> </ul> </li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Additional TAU datasets for Wi-Fi fingerprinting-based positioning

<p><strong>1. Contents</strong></p> <p>This document describes two datasets collected at Tampere University facilities with samples taken from a Wi-Fi network interface for experiments with indoor positioning based on Wi-Fi fingerprinting.</p> <p>To reference this dataset, please use</p> <p>E.S. Lohan et al. &ldquo;Additional TAU datasets for Wi-Fi fingerprinting-based positioning&rdquo; 10.5281/zenodo.3819917</p> <p>Additional reference using these datasets</p> <p><em>Torres-Sospedra, J.; Quezada-Gaibor, D.; Mendoza-Silva, G. M.; Nurmi, J.; Koucheryavy, Y. and Huerta, J. New Cluster Selection and Fine-grained Search for k-Means Clustering and Wi-Fi Fingerprinting Proceedings of the Tenth International Conference on Localization and GNSS (ICL-GNSS), 2020.</em></p> <p><strong>Dataset format</strong></p> <p>Two independent datasets are provided, they are in different folders, namely &ldquo;Database_Building01&rdquo; and &ldquo;Database_Building02&rdquo; respectively. Each dataset includes two sets of samples:</p> <ul> <li>radio map &ndash; a set of Wi-Fi samples collected at a grid of points (reference points);</li> <li>evaluation &ndash; a set of Wi-Fi samples randomly collected in the evaluation area.</li> </ul> <p>Two files are provided for each set that include the rss vectors and the coordinates. For the radio map, the provided files have their names starting with &ldquo;rm_&rdquo;; for the evaluation, the evaluation files have their names starting with &ldquo;eval_&rdquo;. For instance, for the radio map they are:</p> <ul> <li>rm_crd.csv: holds coordinates (x,y)and floor identifier (z) where the samples were collected;</li> <li>rm_rss.csv: holds the measured RSSI values from each of the Access Points (AP) detected in each sample;</li> </ul> <p>All the file are described in the same format, and all files are CSV &ndash; Comma Separated Values plain text (UTF-8).</p> <p><strong>Coordinates:</strong> Each sample is associated to a pair of coordinates in a 2D Euclidean reference system. The origin of the reference system was chosen arbitrarily for convenience. The units are meters. Therefore, distances between points can be easy calculated. Moreover, the floor identifier is included to enable 3D positioning.</p> <p><strong>RSSI values:</strong> The RSSI values provided as read from the Wi-Fi network interface through the Android API. In each sample, a value of +100 was assigned to each AP not detected during a measurement. No information is provided about the MAC addresses of the APs. However, in the files, the same order is used for all samples, meaning that the values in each column are all associated to the same AP.</p> <p>Both datasets are independent and none of the provided files include an identifier for each sample. The values in the two provided files are associated by the line number, meaning that the coordinates and RSSI values in the same line, in each file, refer to the same sample.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

TAU-NIGENS Spatial Sound Events 2020

<p><strong>DESCRIPTION:</strong></p> <p>The <strong>TAU-NIGENS Spatial Sound Events 2020</strong> dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and distances as seen from the recording position.&nbsp;The spatialization of all sound events is based on filtering through real spatial room impulse responses (RIRs), captured in multiple rooms of various shapes, sizes, and acoustical absorption properties. Furthermore, each scene recording is delivered in two spatial recording formats, a microphone array one (<strong>MIC</strong>), and first-order Ambisonics one (<strong>FOA</strong>). The sound events are spatialized as either stationary sound sources in the room, or moving sound sources, in which case time-variant RIRs are used. Each sound event in the sound scene is associated with a trajectory of its direction-of-arrival (DoA) to the recording point, and a temporal onset and offset time. The isolated sound event recordings used for the synthesis of the sound scenes are obtained from the <a href="https://doi.org/10.5281/zenodo.2535878">NIGENS general sound events database</a>. These recordings serve as the development dataset for the <a href="http://dcase.community/challenge2020/task-sound-event-localization-and-detection">DCASE 2020 Sound Event Localization and Detection Task</a> of the <a href="http://dcase.community/challenge2020/">DCASE 2020 Challenge</a>.</p> <p><strong>REPORT &amp; REFERENCE:</strong></p> <p>If you use this dataset please cite the report on its creation, and the corresponding DCASE2020 task setup:</p> <p>Politis., Archontis, Adavanne, Sharath, &amp; Virtanen, Tuomas (2020). A Dataset of Reverberant Spatial Sound Scenes with Moving Sources for Sound Event Localization and Detection. In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020)</em>, Tokyo, Japan.</p> <p>A longer version with more detailed information can be also found <a href="https://arxiv.org/pdf/2006.01919.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The dataset includes a large number of mixtures of sound events with realistic spatial properties under different acoustic conditions, and hence it is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li>600 one-minute long sound scene recordings (development dataset).</li> <li>200 one-minute long sound scene recordings (evaluation dataset).</li> <li>Sampling rate 24kHz.</li> <li>About 700 sound event samples spread over 14 classes (see <a href="http://doi.org/10.5281/zenodo.2535878">here</a> for more details).</li> <li>8 provided cross-validation splits of 100 recordings each, with unique sound event samples and rooms in each of them.</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics&nbsp;(<strong>FOA</strong>) and tetrahedral microphone array.</li> <li>Realistic spatialization and reverberation through RIRs collected in 15 different enclosures.</li> <li>From about 1500 to 3500 possible RIR positions across the different rooms.</li> <li>Both static reverberant and moving reverberant sound events.</li> <li>Up to two overlapping sound events allowed, temporally and spatially.</li> <li>Realistic spatial ambient noise collected from each room is added to the spatialized sound events, at varying signal-to-noise ratios (SNR) ranging from noiseless (30dB) to noisy (6dB).</li> </ul> <p>The IRs were collected in Finland by staff of Tampere University between 12/2017 - 06/2018, and between 11/2019 - 1/2020. The older measurements from five rooms were also used for the earlier <a href="https://doi.org/10.5281/zenodo.2580091">development</a> and <a href="https://doi.org/10.5281/zenodo.3066124">evaluation</a> datasets&nbsp;<strong>TAU Spatial Sound Events 2019</strong>, while ten additional rooms were added for this dataset. The data collection received funding from the European Research Council, grant agreement <a href="https://cordis.europa.eu/project/id/637422">637422 EVERYSOUND</a>.</p> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model of a convolutional recurrent neural network, performing joint SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sharathadavanne/seld-dcase2020">here</a>. This implementation serves as the baseline method in the <a href="http://dcase.community/challenge2020/task-sound-event-localization-and-detection">DCASE 2020 Sound Event Localization and Detection Task</a>.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>Version 1.0 of the dataset included only the 600 development audio recordings and labels, used by the participants of Task 3 of DCASE2020 Challenge to train and validate their submitted systems. Version 1.1 included additionally the 200 evaluation audio recordings without labels, for the evaluation phase of DCASE2020. The latest version 1.2, published after the completion of the challenge, includes also the labels for the evaluation files.</p> <p>If researchers wish to compare their system against the submissions of DCASE2020 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The three files, <strong><em>foa_dev.z01</em></strong>,<strong><em> foa_dev.z02</em></strong>, and <strong><em>foa_dev.zip</em></strong>, correspond to audio data of the <strong>FOA </strong>recording format.<br> The three files, <strong><em>mic_dev.z01</em></strong>,<strong><em> mic_dev.z02</em></strong>, and <strong><em>mic_dev.zip</em></strong>, correspond to audio data of the <strong>MIC</strong> recording format.<br> The <strong><em>metadata_dev.zip</em></strong>&nbsp;is the common metadata for both formats.</p> <p>The file, <em><strong>foa_eval.zip</strong></em>, corresponds to audio data of the <strong>FOA</strong> recording format for the evaluation dataset.<br> The file, <em><strong>mic_eval.zip</strong></em>, corresponds to audio data of the <strong>MIC</strong> recording format for the evaluation dataset.<br> The <em><strong>metadata_eval.zip</strong></em> is the common metadata for both formats. An info file is included (<em>metadata_eval_info.txt</em>) which specifies which of the two evaluation folds the mix file belongs to, and what is its number of overlapping events.</p> <p>Download the zip files corresponding to the format of interest and use your favorite compression tool to unzip these split zip files. To extract a split zip archive (named as zip, z01, z02, ...), you could use, for example, the following syntax in Linux or OSX terminal:</p> <ol> <li>Combine the split archive to a single archive: <pre>zip -s 0 split.zip --out single.zip</pre> </li> <li>Extract the single archive using unzip: <pre>unzip single.zip</pre> </li> </ol>

opencc-by-nc-4.0Apr 2020View details →
zenodo44/100

HFLAV Tau averages

<h2>HFLAV results for Tau mass, lifetime and branching fractions through August 2024</h2> <p>Cite the results presented as<br><br>Sw. Banerjee et al., *Averages of b-hadron, c-hadron, and tau-lepton properties as of 2023*, [arXiv:2411.18639](https://arxiv.org/abs/2411.18639), with specific result from [doi:10.5072/zenodo.13989054](https://doi.org/10.5072/zenodo.13989054).<br><br>Alternatively use the bibtex record</p> <pre>@article{Banerjee:2024znd,<br> collaboration = "Heavy Flavor Averaging Group, HFLAV",<br>&nbsp; &nbsp;&nbsp;author = "Banerjee, Swagato and others", title = "{Averages of $b$-hadron, $c$-hadron, and $\tau$-lepton properties as of 2023}", eprint = "2411.18639", archivePrefix = "arXiv", primaryClass = "hep-ex", month = "11", year = "2024",<br> note = "{with specific result from \href{https://doi.org/10.5281/zenodo.13989054}{{\texttt{doi:10.5281/zenodo.13989054}}}}"<br>}</pre> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

STARSS22: Sony-TAu Realistic Spatial Soundscapes 2022 dataset

<p><strong>DESCRIPTION:</strong></p> <p>The **<strong>Sony-TAu Realistic Spatial Soundscapes 2022 (STARSS22)</strong>** dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of target classes. The dataset is collected in two different countries, in Tampere, Finland by the Audio Researh Group (ARG) of **<strong>Tampere University (TAU)</strong>**, and in Tokyo, Japan by **<strong>SONY</strong>**, using a similar setup and annotation procedure. The dataset is delivered in two 4-channel spatial recording formats, a microphone array one (**<strong>MIC</strong>**), and first-order Ambisonics one (**<strong>FOA</strong>**). These recordings serve as the development dataset for the&nbsp;<a href="https://dcase.community/challenge2022/task-sound-event-localization-and-detection">DCASE 2022 Sound Event Localization and Detection Task</a>&nbsp;of the&nbsp;<a href="https://dcase.community/challenge2022/">DCASE 2022 Challenge</a>.</p> <p>Contrary to the three previous datasets of synthetic spatial sound scenes of TAU Spatial Sound Events 2019 (<a href="https://zenodo.org/record/2599196">development</a>/<a href="https://zenodo.org/record/3377088">evaluation</a>), <a href="https://doi.org/10.5281/zenodo.4064792">TAU-NIGENS Spatial Sound Events 2020</a>, and <a href="https://zenodo.org/record/5476980">TAU-NIGENS Spatial Sound Events 2021</a>&nbsp;associated with the previous iterations of the DCASE Challenge, the STARS22 dataset contains recordings of real sound scenes and hence it avoids some of the pitfalls of synthetic generation of scenes. Some such key properties are:</p> <ul> <li>annotations are based on a combination of human annotators for sound event activity and optical tracking for spatial positions,</li> <li>the annotated target event classes are determined by the composition of the real scenes,</li> <li>the density, polyphony, occurences and co-occurences of events and sound classes is not random, and it follows actions and interactions of participants in the real scenes.</li> </ul> <p>The recordings were collected between September 2021 and January 2022. Collection of data from the TAU side has received funding from Google.</p> <p><strong>REPORT &amp; REFERENCE:</strong></p> <p>If you use this dataset please cite the report on its creation, and the related DCASE2022 task setup:</p> <p>Archontis Politis,&nbsp;Kazuki Shimada,&nbsp;Parthasaarathy Sudarsanam,&nbsp;Sharath Adavanne,&nbsp;Daniel Krause,&nbsp;Yuichiro Koyama,&nbsp;Naoya Takahashi,&nbsp;Shusuke Takahashi,&nbsp;Yuki Mitsufuji,&nbsp;Tuomas Virtanen (2022).&nbsp;<strong>STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events</strong>.&nbsp;In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022)</em>, Nancy, France.</p> <p>found <a href="https://dcase.community/documents/workshop2022/proceedings/DCASE2022Workshop_Politis_51.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The dataset is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li>70 recording clips of 30 sec ~ 5 min durations, with a total time of ~2hrs, contributed by SONY (development dataset).</li> <li>51 recording clips of 1 min ~ 5 min durations, with a total time of ~3hrs, contributed by TAU (development dataset).</li> <li>52 recording clips with a total time of ~2hrs, contributed by SONY&amp;TAU (evaluation dataset).</li> <li>A training-test split is provided for reporting results using the development dataset.</li> <li>40 recordings contributed by SONY for the training split, captured in 2 rooms (dev-train-sony).</li> <li>30 recordings contributed by SONY for the testing split, captured in 2 rooms (dev-test-sony).</li> <li>27 recordings contributed by TAU for the training split, captured in 4 rooms (dev-train-tau).</li> <li>24 recordings contributed by TAU for the testing split, captured in 3 rooms (dev-test-tau).</li> <li>A total of 11 unique rooms captured in the recordings, 4 from SONY and 7 from TAU (development set).</li> <li>Sampling rate 24kHz.</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array (MIC).</li> <li>Recordings are taken in two different countries and two different sites.</li> <li>Each recording clip is part of a recording session happening in a unique room.</li> <li>Groups of participants, sound making props, and scene scenarios are unique for each session (with a few exceptions).</li> <li>To achieve good variability and efficiency in the data, in terms of presence, density, movement, and/or spatial distribution of the sounds events, the scenes are loosely scripted.</li> <li>13 target classes are identified in the recordings and strongly annotated by humans.</li> <li>Spatial annotations for those active events are captured by an optical tracking system.</li> <li>Sound events out of the target classes are considered as interference.</li> <li>Occurences of up to 3 simultaneous events are fairly common, while higher numbers of overlapping events (up to 5) can occur but are rare.</li> </ul> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>SOUND CLASSES:</strong></p> <p>13 target sound event classes are annotated. The classes follow loosely the <a href="https://research.google.com/audioset/ontology/index.html">Audioset ontology</a>.</p> <p>&nbsp; 0. <strong>Female speech, woman speaking</strong><br> &nbsp; 1. <strong>Male speech, man speaking</strong><br> &nbsp; 2. <strong>Clapping</strong><br> &nbsp; 3. <strong>Telephone</strong><br> &nbsp; 4. <strong>Laughter</strong><br> &nbsp; 5. <strong>Domestic sounds</strong><br> &nbsp; 6. <strong>Walk, footsteps</strong><br> &nbsp; 7. <strong>Door, open or close</strong><br> &nbsp; 8. <strong>Music</strong><br> &nbsp; 9. <strong>Musical instrument</strong><br> &nbsp; 10. <strong>Water tap, faucet</strong><br> &nbsp; 11. <strong>Bell</strong><br> &nbsp; 12. <strong>Knock</strong></p> <p>The content of some of these classes corresponds to events of a limited range of Audioset-related subclasses. For more information see the README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model of a convolutional recurrent neural network, performing joint SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sharathadavanne/seld-dcase2022">here</a>. This&nbsp;implementation will serve as the baseline method in the&nbsp;DCASE 2022 Sound Event Localization and Detection Task.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>The current version (Version 1.1) of the dataset includes the 121 development audio recordings and labels, used by the participants of Task 3 of DCASE2022&nbsp;Challenge to train and validate their submitted systems, and the 52 evaluation audio recordings without labels, for the evaluation phase of DCASE2022.</p> <p>If researchers wish to compare their system against the submissions of DCASE2022 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The file&nbsp;<strong><em>foa_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>FOA&nbsp;</strong>recording format.<br> The file&nbsp;<strong><em>mic_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>MIC</strong>&nbsp;recording format.<br> The&nbsp;<strong><em>metadata_dev.zip</em></strong>&nbsp;is&nbsp;the common metadata for both formats.</p> <p>The file&nbsp;<strong><em>foa_eval.zip</em></strong>, corresponds to audio data of the&nbsp;<strong>FOA</strong>&nbsp;recording format for the evaluation dataset.<br> The file&nbsp;<strong><em>mic_eval.zip</em></strong>, corresponds to audio data of the&nbsp;<strong>MIC</strong>&nbsp;recording format for the evaluation dataset.</p> <p>Download the zip files corresponding to the format of interest and use your favourite compression tool to unzip these zip files.</p>

openmit-licenseMar 2022View details →
zenodo44/100

The eclipse of the V773 Tau B circumbinary disk

<p>Data related to the paper &quot;The eclipse of the V773 Tau B circumbinary disk&quot; to be published in A&amp;A. Data includes high contrast imaging using the SPHERE instrument on the VLT, photometric time series data from several ground based observatories and from the TESS satellite, and orbital walker bundles from emcee and orbitize!</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Unsupervised New Physics detection at 40 MHz: h+ -> tau nu Signal Benchmark Dataset

<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of h+ -&gt; tau nu&nbsp;&nbsp;decays produced in&nbsp;collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page:&nbsp;https://mpp-hep.github.io/ADC2021/</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Unsupervised New Physics detection at 40 MHz: h^0 -> tau tau Signal Benchmark Dataset

<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of h^0 -&gt; tau tau&nbsp;decays produced in&nbsp;collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page:&nbsp;https://mpp-hep.github.io/ADC2021/</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Unsupervised New Physics detection at 40 MHz: LQ -> b tau Signal Benchmark Dataset

<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of Leptoquarks&nbsp;-&gt; b tau &nbsp;decays produced in&nbsp;collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page:&nbsp;https://mpp-hep.github.io/ADC2021/</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Fuτure - dataset for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons

<h1>&nbsp;Data description</h1> <h2>MC Simulation</h2> <p><br>The <strong>Fu&tau;ure</strong> dataset is intended for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons. The dataset is generated with Pythia 8, with the full detector simulation being performed by Geant4 with the CLIC-like detector setup CLICdet (CLIC_o3_v14) setup. Events are reconstructed using the Marlin reconstruction framework and interfaced with Key4HEP. Particle candidates in the reconstructed events are reconstructed using the PandoraPF algorithm.</p> <p>In this version of the dataset no &gamma;&gamma; -&gt; hadrons background is included.</p> <h2>Samples</h2> <p><br>This dataset contains e+e- samples with Z-&gt;&tau;&tau;, ZH,H-&gt;&tau;&tau; and Z-&gt;qq events, with approximately 2 million events simulated in each category.</p> <p>The following processes e+e- were simulated with Pythia 8 at sqrt(s) = 380 GeV:</p> <ul> <li>p8_ee_qq_ecm380 [Z -&gt; qq events]</li> <li>p8_ee_ZH_Htautau [ZH -&gt; Ztautau]</li> <li>p8_ee_Z_Ztautau_ecm380 [ZH -&gt; Ztautau]</li> </ul> <p>The .root files from the MC simulation chain are eventually processed by the software found in&nbsp;<a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a> in order to create flat ntuples as the final product.</p> <h2><br>Features</h2> <p><br>The basis of the ntuples are the particle flow (PF) candidates from PandoraPF. Each PF candidate has four momenta, charge and particle label (electron / muon / photon / charged hadron / neutral hadron). The PF candidates in a given event are clustered into jets using generalized kt algorithm for ee collisions, with parameters p=-1 and R=0.4. The minimum pT is set to be 0 GeV for both generator level jets and reconstructed jets. The dataset contains the four momenta of the jets, with the PF candidates in the jets with the above listed properties.</p> <p>Additionally, a set of variables describing the tau lifetime are calculated using the software in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a>. As tau lifetime is very short, these variables are sensitive to true tau decays.&nbsp;In the calculation of these lifetime variables, we use a linear approximation.</p> <p>In summary, the features found in the flat ntuples are:</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>reco_cand_p4s</td> <td>4-momenta per particle in the reco jet.</td> </tr> <tr> <td>reco_cand_charge</td> <td>Charge per particle in the jet.</td> </tr> <tr> <td>reco_cand_pdg</td> <td>PDGid per particle in the jet.</td> </tr> <tr> <td>reco_jet_p4s</td> <td>RecoJet 4-momenta.</td> </tr> <tr> <td>reco_cand_dz</td> <td>Longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dz_err</td> <td>Uncertainty of the longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy</td> <td>Transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy_err</td> <td>Uncertainty of the transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>gen_jet_p4s</td> <td>GenJet 4-momenta. Matched with RecoJet within a cone of radius dR &lt; 0.3.</td> </tr> <tr> <td>gen_jet_tau_decaymode</td> <td>Decay mode of the associated genTau. Jets that have associated leptonically decaying taus are removed, so there are no DM=16 jets. If no GenTau can be matched to GenJet within dR &lt; 0.4, a fill value is used.</td> </tr> <tr> <td>gen_jet_tau_p4s</td> <td>Visible 4-momenta of the genTau. If no GenTau can be matched to GenJet within dR&lt;0.4, a fill value is used.</td> </tr> </tbody> </table> <p>The ground truth is based on stable particles at the generator level, before detector simulation. These particles are clustered into generator-level jets and are matched to generator-level &tau; leptons as well as reconstructed jets. In order for a generator-level jet to be matched to generator-level &tau; lepton, the &tau; lepton needs to be inside a cone of dR = 0.4. The same applies for the reconstructed jet, with the requirement on dR being set to dR = 0.3. For each reconstructed jet, we define three target values related to &tau; lepton reconstruction:</p> <ul> <li>&nbsp;a binary flag <strong>isTau</strong> if it was matched to a generator-level hadronically decaying &tau; lepton. <strong>gen_jet_tau_decaymode</strong> of value -1 indicates no match to generator-level hadronically decaying &tau;.</li> <li>&nbsp;the categorical decay mode of the &tau; <strong>gen_jet_tau_decaymode</strong> in terms of the number of generator level charged and neutral hadrons. Possible <strong>gen_jet_tau_decaymode</strong> are {0, 1, . . . , 15}.</li> <li>&nbsp;if matched, the visible (neglecting neutrinos), reconstructable pT of the &tau; lepton. This is inferred from the <strong>gen_jet_tau_p4s</strong></li> </ul> <h2>Contents:</h2> <ul> <li>qq_test.parquet</li> <li>qq_train.parquet</li> <li>zh_test.parquet</li> <li>zh_train.parquet</li> <li>z_test.parquet</li> <li>&nbsp;z_train.parquet</li> <li>data_intro.ipynb</li> </ul> <h2>Dataset characteristics</h2> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong># Jets</strong></td> <td><strong>Size</strong></td> </tr> <tr> <td>z_test.parquet</td> <td> <pre>870 843</pre> </td> <td>171 MB</td> </tr> <tr> <td>z_train.parquet</td> <td> <pre>3 483 369</pre> </td> <td>681 MB</td> </tr> <tr> <td>zh_test.parquet</td> <td> <pre>1 068 606</pre> </td> <td>213 MB</td> </tr> <tr> <td>zh_train.parquet</td> <td> <pre>4 274 423</pre> </td> <td>851 MB</td> </tr> <tr> <td>qq_test.parquet</td> <td> <pre>6 366 715</pre> </td> <td>1.4 GB</td> </tr> <tr> <td>qq_train.parquet</td> <td> <pre>25 466 858</pre> </td> <td>5.6 GB</td> </tr> </tbody> </table> <p>The dataset consists of 6 files of 8.9 GB in total.</p> <h2>How can you use these data?</h2> <p>The .parquet files can be directly loaded with the Awkward Array Python library.<br>An example how one might use the dataset and the features is given in&nbsp;<strong>data_intro.ipynb</strong></p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

TAU-NIGENS Spatial Sound Events 2021

<p><strong>DESCRIPTION:</strong></p> <p>The&nbsp;<strong>TAU-NIGENS Spatial Sound Events 2021</strong>&nbsp;dataset contains multiple spatial sound-scene recordings, consisting of sound events of distinct categories integrated into a variety of acoustical spaces, and from multiple source directions and distances as seen from the recording position.&nbsp;The spatialization of all sound events is based on filtering through real spatial room impulse responses (RIRs), captured in multiple rooms of various shapes, sizes, and acoustical absorption properties. Furthermore, each scene recording is delivered in two spatial recording formats, a microphone array one (<strong>MIC</strong>), and first-order Ambisonics one (<strong>FOA</strong>). The sound events are spatialized as either stationary sound sources in the room, or moving sound sources, in which case time-variant RIRs are used. Each sound event in the sound scene is associated with a single direction-of-arrival (DoA) if static, a&nbsp;trajectory DoAs if moving, and a temporal onset and offset time. The isolated sound event recordings used for the synthesis of the sound scenes are obtained from the&nbsp;<a href="https://doi.org/10.5281/zenodo.2535878">NIGENS general sound events database</a>. These recordings serve as the development dataset for the&nbsp;<a href="http://dcase.community/challenge2021/task-sound-event-localization-and-detection">DCASE 2021 Sound Event Localization and Detection Task</a>&nbsp;of the&nbsp;<a href="http://dcase.community/challenge2021/">DCASE 2021 Challenge</a>.</p> <p>This&nbsp;dataset is the third iteration of spatialized&nbsp;sound event datasets based on&nbsp;real room responses and ambient noise from multiple spaces, with each iteration introducing more challenging conditions closer to real-life. Those iterations, including the present one, are:</p> <ul> <li><strong>TAU Spatial Sound Events 2019,&nbsp;</strong><a href="https://doi.org/10.5281/zenodo.2580091">development</a>&nbsp;and&nbsp;<a href="https://doi.org/10.5281/zenodo.3066124">evaluation</a>&nbsp;datasets.<br> 5 rooms, high direct-to-reverberant&nbsp;ratios (DRR), static sources only, minimum DoA separation 10&deg;, discrete grid of DoAs, high SNR for ambient noise, maximum polyphony of 2 simultaneous events</li> <li><strong><a href="https://doi.org/10.5281/zenodo.4064792">TAU-NIGENS Spatial Sound Events 2020</a></strong>, development and evaluations datasets.<br> 13 rooms, low-to-high DRRs, static and moving sources, continuous DoAs, low-to-high SNR for ambient noise,<br> maximum polyphony of 2 simultaneous events</li> <li><strong>TAU-NIGENS Spatial Sound Events 2021</strong>.<br> Same as 2020, with the following exceptions: a more natural temporal distribution of sound events,<br> maximum polyphony of 3&nbsp;target events, <strong>inclusion of additional out-of-target-classes directional interference events</strong></li> </ul> <p>The inclusion of directional interferences is the main new challenging property of the new dataset. They are spatialized in the scene in the same way as the target events, and can be either static or moving. The interfering events are sourced from the &quot;engine&quot;, &quot;fire&quot;, and &quot;general&quot; classes of the NIGENS sound event database. The interferers are considered unknown and no activity or directional labels of them are provided with the training datasets.</p> <p><strong>REPORT &amp; REFERENCE:</strong></p> <p>If you use this dataset please cite the report on its creation, and the corresponding DCASE2020 task setup:</p> <p>Archontis Politis, Sharath Adavanne, Daniel Krause, Antoine Deleforge, Prerak Srivastava, Tuomas Virtanen (2021).<br> A Dataset of Dynamic Reverberant Sound Scenes with Directional Interferers for Sound Event Localization and Detection.&nbsp;<br> In <em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2021)</em>, Barcelona, Spain.</p> <p>available <a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Politis_43.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The dataset includes a large number of mixtures of sound events with realistic spatial properties under different acoustic conditions, and hence it is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p><strong>SPECIFICATIONS:</strong></p> <ul> <li>600 one-minute long sound scene recordings with metadata (development dataset).</li> <li>200 one-minute long sound scene recordings without metadata (evaluation dataset).</li> <li>Sampling rate 24kHz.</li> <li>About 500 sound event samples distributed over the 12 target classes (see [here](http://doi.org/10.5281/zenodo.2535878) for more details).</li> <li>About 400 sound event samples used as interference events (see [here](http://doi.org/10.5281/zenodo.2535878) for more details).</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array.</li> <li>Realistic spatialization and reverberation through multichannel RIRs collected in 13 different enclosures.</li> <li>From 1184 to 6480 possible RIR positions across the different rooms.</li> <li>Both static reverberant and moving reverberant sound events.</li> <li>Three possible angular speeds for moving sources of approximately 10, 20, or 40deg/sec.</li> <li>Up to three overlapping sound events possible, temporally and spatially.</li> <li>Simultaneous directional interfering sound events with their own temporal activities, static or moving.</li> <li>Realistic spatial ambient noise collected from each room is added to the spatialized sound events, at varying signal-to-noise ratios (SNR) ranging from noiseless (30dB) to noisy (6dB) conditions.</li> </ul> <p>The IRs were collected in Finland by staff of Tampere University between 12/2017 - 06/2018, and between 11/2019 - 1/2020.&nbsp;The data collection received funding from the European Research Council, grant agreement&nbsp;<a href="https://cordis.europa.eu/project/id/637422">637422 EVERYSOUND</a>.</p> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model of a convolutional recurrent neural network, performing joint SELD, trained and evaluated with this dataset will be provided soon. That&nbsp;implementation will serve as the baseline method in the&nbsp;<a href="http://dcase.community/challenge2021/task-sound-event-localization-and-detection">DCASE 2021 Sound Event Localization and Detection Task</a>.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>The current and final version (Version 1.2) of the dataset includes the 600 development audio recordings and labels, used by the participants of Task 3 of DCASE2021&nbsp;Challenge to train and validate their submitted systems, and the 200 evaluation audio recordings including their&nbsp;labels, used in the evaluation phase of DCASE2021.</p> <p>If researchers wish to compare their system against the submissions of DCASE2021 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The three files,&nbsp;<strong><em>foa_dev.z01</em></strong>, and&nbsp;<strong><em>foa_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>FOA&nbsp;</strong>recording format.<br> The three files,&nbsp;<strong><em>mic_dev.z01</em></strong>,&nbsp;and&nbsp;<strong><em>mic_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>MIC</strong>&nbsp;recording format.<br> The&nbsp;<strong><em>metadata_dev.zip</em></strong>&nbsp;is&nbsp;the common metadata for both formats.</p> <p>The file,<strong>&nbsp;<em>foa_eval.zip</em></strong>, corresponds to audio data of the&nbsp;FOA&nbsp;recording format for the evaluation dataset.<br> The file,&nbsp;<strong><em>mic_eval.zip</em></strong>, corresponds to audio data of the&nbsp;MIC&nbsp;recording format for the evaluation dataset.<br> The&nbsp;<strong><em>metadata_eval.zip</em></strong>&nbsp;is the common metadata for both formats.</p> <p>Download the zip files corresponding to the format of interest and use your favorite compression tool to unzip these split zip files. To extract a split zip archive (named as zip, z01, z02, ...), you could use, for example, the following syntax in Linux or OSX terminal:</p> <ol> <li>Combine the split archive to a single archive: <pre>zip -s 0 split.zip --out single.zip</pre> </li> <li>Extract the single archive using unzip: <pre>unzip single.zip</pre> </li> </ol>

opencc-by-nc-4.0Feb 2021View details →
zenodo44/100

Nuclear Magnetic resonance Dataset of 2D spectra of S100B and Tau to study their protein-protein interaction

<p>Nuclear Magnetic resonance dataset of 2D spectra corresponding to raw data of research published in Nature Communication in a communication entitled &quot;Dynamic interactions and Ca2+ 1 -binding modulate the holdase-type chaperone activity of S100B preventing tau&nbsp;aggregation and seeding&quot; by Moreira G. et al.</p> <p>Dataset corresponds to</p> <p>raw data files in Bruker format of NMR 2D spectra (ser), associated with&nbsp;files of acquisition parameters and processing parameters (pdata),</p> <p>files in .ucsf format that can be read with NMRFAM-Sparky (free download) of 2D spectra (in sub-directory pdata/1)</p> <p>files of chemical shift value lists that can be read as text files or in NMRFAM sparky together with the corresponding ucsf files.</p> <p>physico-chemical conditions are found in title in pdata\1</p> <p>Data were acquired on a Bruker 900-MHz spectrometer equipped with a triple-resonance cryogenic probe (Bruker, Karlsruhe, Germany)</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

It takes Tau to tango : Investigating the fuzzy interaction between the Tau-R2 repeat domain and the C-terminal tails of tubulins

<p>The microtubule-associated protein (MAP) tau plays a key role in the regulation of microtubule assembly and spatial organisation. Tau hyperphosphorylation affects its binding on the tubulin surface and has been shown to be involved in several pathologies such as Alzheimer disease. As the tau binding site on the microtubule lays close to the disordered and highly flexible tubulin C-terminal tails (CTTs), these are likely to impact the tau-tubulin interaction. Since the disordered tubulin CTTs are missing from the available experimental structures, we used homology modeling to build two complete models of tubulin heterotrimers with different isotypes for the &beta;-tubulin subunit (&beta;I/&alpha;I/&beta;I and &beta;III/&alpha;I/&beta;III). We then performed long timescale classical Molecular Dynamics simulations for the tauR2-tubulin assembly (in systems with and without CTTs) and analyzed the resulting trajectories to obtain a detailed view of the protein interface in the complex and the impact of the CTTs on the stability of this assembly. Additional analyses of the CTTs mobility in the presence, or in the absence, of tau also highlight how tau might modulate the CTTs activity as hooks that are involved in the recruitment of several MAPs.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

STARSS23: Sony-TAu Realistic Spatial Soundscapes 2023

<p><strong>DESCRIPTION:</strong></p> <p>The <strong>Sony-TAu Realistic Spatial Soundscapes 2023&nbsp;(STARSS23)</strong>&nbsp;dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of target classes. The dataset is collected in two different countries, in Tampere, Finland by the Audio Researh Group (ARG) of <strong>Tampere University (TAU)</strong>, and in Tokyo, Japan by <strong>SONY</strong>, using a similar setup and annotation procedure. The dataset is delivered in two 4-channel spatial recording formats, a microphone array one (<strong>MIC</strong>), and first-order Ambisonics one (<strong>FOA</strong>). These recordings serve as the development dataset for the&nbsp;<a href="https://dcase.community/challenge2023/task-sound-event-localization-and-detection-evaluated-in-real-spatial-sound-scenes">DCASE 2023 Sound Event Localization and Detection Task</a>&nbsp;of the&nbsp;<a href="https://dcase.community/challenge2023/">DCASE 2023 Challenge</a>.<br> <br> The STARSS23 dataset is a continuation of the <a href="https://zenodo.org/record/6600531">STARSS22 dataset</a>. It extends the previous version with the following:</p> <ul> <li>An additional&nbsp;<strong>additional&nbsp;2hrs 30mins&nbsp;</strong>of recordings in the development set, from&nbsp;<strong>5 new rooms</strong>&nbsp;distributed in 47 new recording clips.</li> <li>An <strong>additional 1hr 40mins</strong>&nbsp;of recordings added in the evaluation set of the dataset.</li> <li><strong>360&deg; videos</strong> spatially and temporally aligned to the audio recordings of the dataset (apart from 12 audio-only clips).</li> <li><strong>Distance labels</strong> (in cm) for the spatially annotated sound events, instead of the previous&nbsp;azimuth and elevation only labels.</li> </ul> <p>Contrary to the three previous datasets of synthetic spatial sound scenes of TAU Spatial Sound Events 2019 (<a href="https://zenodo.org/record/2599196">development</a>/<a href="https://zenodo.org/record/3377088">evaluation</a>),&nbsp;<a href="https://doi.org/10.5281/zenodo.4064792">TAU-NIGENS Spatial Sound Events 2020</a>, and&nbsp;<a href="https://zenodo.org/record/5476980">TAU-NIGENS Spatial Sound Events 2021</a>&nbsp;associated with previous iterations of the DCASE Challenge, the STARS22-23 dataset contains recordings of real sound scenes and hence it avoids some of the pitfalls of synthetic generation of scenes. Some such key properties are:</p> <ul> <li>annotations are based on a combination of human annotators for sound event activity and optical tracking for spatial positions,</li> <li>the annotated target event classes are determined by the composition of the real scenes,</li> <li>the density, polyphony, occurences and co-occurences of events and sound classes is not random, and it follows actions and interactions of participants in the real scenes.</li> </ul> <p>The first round of recordings was collected between September 2021 and January 2022. A second round of recordings was collected between&nbsp;November 2022 and February 2023.<br> <br> Collection of data from the TAU side has received funding from Google.</p> <p>A demo video combining the different modalities and spatial annotations can be found <a href="https://www.youtube.com/watch?v=ZtL-8wBYPow">here</a>.</p> <p><strong>REPORT &amp; REFERENCE:</strong></p> <p>If you use this dataset you could cite this report on its design, capturing, and annotation process:</p> <p>Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji (2023). <strong>STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events</strong>,<br> <br> found <a href="https://arxiv.org/abs/2306.09126">here</a>, and</p> <p>Archontis Politis,&nbsp;Kazuki Shimada,&nbsp;Parthasaarathy Sudarsanam,&nbsp;Sharath Adavanne,&nbsp;Daniel Krause,&nbsp;Yuichiro Koyama,&nbsp;Naoya Takahashi,&nbsp;Shusuke Takahashi,&nbsp;Yuki Mitsufuji,&nbsp;Tuomas Virtanen (2022).&nbsp;<strong>STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events</strong>.&nbsp;In&nbsp;<em>Proceedings of the Detection and Classification of Acoustic Scenes and Events 2022 Workshop (DCASE2022)</em>, Nancy, France.</p> <p>found&nbsp;<a href="https://dcase.community/documents/workshop2022/proceedings/DCASE2022Workshop_Politis_51.pdf">here</a>.</p> <p><strong>AIM:</strong></p> <p>The STARSS22-23 dataset is suitable for training and evaluation of machine-listening models for sound event detection (SED), general sound source localization with diverse sounds or signal-of-interest localization, and joint sound-event-localization-and-detection (SELD). Additionally, the dataset can be used for evaluation of signal processing methods that do not necessarily rely on training, such as acoustic source localization methods and multiple-source acoustic tracking. The dataset allows evaluation of the performance and robustness of the aforementioned applications for diverse types of sounds, and under diverse acoustic conditions.</p> <p>Specifically the STARSS23 allows evaluation of audiovisual processing methods with a spatial dimension, such as audiovisual source localization or audiovisual object recognition.</p> <p><strong>SPECIFICATIONS:</strong></p> <p>General:</p> <ul> <li>Recordings are taken in two different sites.</li> <li>Each recording clip is part of a recording session happening in a unique room.</li> <li>Groups of participants, sound making props, and scene scenarios are unique for each session (with a few exceptions).</li> <li>To achieve good variability and efficiency in the data, in terms of presence, density, movement, and/or spatial distribution of the sounds events, the scenes are loosely scripted.</li> <li>13 target classes are identified in the recordings and strongly annotated by humans.</li> <li>Spatial annotations for those active events are captured by an optical tracking system.</li> <li>Sound events out of the target classes are considered as interference.</li> <li>Occurrences of up to 3 simultaneous events are fairly common, while higher numbers of overlapping events (up to 5) can occur but are rare.</li> </ul> <p>Volume, duration, and data split:</p> <ul> <li>A total of 16 unique rooms captured in the recordings, 4 in Tokyo and 12 in Tampere (development set).</li> <li>70 recording clips of 30 sec ~ 5 min durations, with a total time of ~2hrs, captured in Tokyo (development dataset).</li> <li>98 recording clips of 40 sec ~ 9 min durations, with a total time of ~5.5hrs, captured in Tampere (development dataset).</li> <li>79 recordings clips of 40 sec ~ 7 min durations, with a total time of ~3.5hrs, captured in both sites (evaluation dataset).</li> <li>A training-testing split is provided for reporting results using the development dataset.</li> <li>40 recordings contributed by Sony for the training split, captured in 2 rooms (dev-train-sony).</li> <li>30 recordings contributed by Sony for the testing split, captured in 2 rooms (dev-test-sony).</li> <li>50 recordings contributed by TAU for the training split, captured in 7 rooms (dev-train-tau).</li> <li>48 recordings contributed by TAU for the testing split, captured in 5 rooms (dev-test-tau).</li> </ul> <p>Audio:</p> <ul> <li>Sampling rate: 24kHz.</li> <li>Bit depth: &nbsp; &nbsp; &nbsp; &nbsp; 16 bits.</li> <li>Two 4-channel 3-dimensional recording formats: first-order Ambisonics (FOA) and tetrahedral microphone array (MIC).</li> </ul> <p>Video:</p> <ul> <li>Video 360&deg; format: equirectangular</li> <li>Video resolution: 1920x960</li> <li>Video frames per second (fps): 29.97</li> <li>All audio recordings are accompanied by synchronised video recordings, apart from 12 audio recordings with missing videos (<em>fold3_room21_mix001.wav&nbsp;-&nbsp;fold3_room21_mix012.wav</em>)</li> </ul> <p>More detailed information on the dataset can be found in the included README file.</p> <p><strong>SOUND CLASSES:</strong></p> <p>13 target sound event classes are annotated. The classes follow loosely the&nbsp;<a href="https://research.google.com/audioset/ontology/index.html">Audioset ontology</a>.</p> <p>&nbsp; 0.&nbsp;<strong>Female speech, woman speaking</strong><br> &nbsp; 1.&nbsp;<strong>Male speech, man speaking</strong><br> &nbsp; 2.&nbsp;<strong>Clapping</strong><br> &nbsp; 3.&nbsp;<strong>Telephone</strong><br> &nbsp; 4.&nbsp;<strong>Laughter</strong><br> &nbsp; 5.&nbsp;<strong>Domestic sounds</strong><br> &nbsp; 6.&nbsp;<strong>Walk, footsteps</strong><br> &nbsp; 7.&nbsp;<strong>Door, open or close</strong><br> &nbsp; 8.&nbsp;<strong>Music</strong><br> &nbsp; 9.&nbsp;<strong>Musical instrument</strong><br> &nbsp; 10.&nbsp;<strong>Water tap, faucet</strong><br> &nbsp; 11.&nbsp;<strong>Bell</strong><br> &nbsp; 12.&nbsp;<strong>Knock</strong></p> <p>The content of some of these classes corresponds to events of a limited range of Audioset-related subclasses. For more information see the README file.</p> <p><strong>EXAMPLE APPLICATION:</strong></p> <p>An implementation of a trainable model performing <strong>audio-only</strong>&nbsp;joint SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sharathadavanne/seld-dcase2023">here</a>. This&nbsp;implementation will serve as the baseline method in the&nbsp;DCASE 2023 Sound Event Localization and Detection Task, under the audio-only inference track.</p> <p>Additionally, an implementation of a trainable model performing <strong>audiovisual</strong>&nbsp;SELD, trained and evaluated with this dataset is provided <a href="https://github.com/sony/audio-visual-seld-dcase2023">here</a>. This&nbsp;implementation will serve as the baseline method in the&nbsp;DCASE 2023 Sound Event Localization and Detection Task, under the audiovisual inference track.</p> <p><strong>DEVELOPMENT AND EVALUATION:</strong></p> <p>The current version (Version 1.1) of the dataset includes development audio/video&nbsp;recordings and labels and the evaluation recordings without labels, used by the participants of Task 3 of DCASE2023&nbsp;Challenge to train and validate their submitted systems (development), and produce system outputs for the challenge evaluation phase.</p> <p>If researchers wish to compare their system against the submissions of DCASE2023 Challenge, they will have directly comparable results if they use the evaluation data as their testing set.</p> <p><strong>DOWNLOAD INSTRUCTIONS:</strong></p> <p>The file&nbsp;<strong><em>foa_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>FOA&nbsp;</strong>recording format.<br> The file&nbsp;<strong><em>mic_dev.zip</em></strong>, correspond to audio data of the&nbsp;<strong>MIC</strong>&nbsp;recording format.</p> <p>The file&nbsp;<strong><em>video_dev.zip&nbsp;</em></strong>contains the common videos for both audio formats.<br> The file&nbsp;<strong><em>metadata_dev.zip</em></strong>&nbsp;contains the common metadata for both audio formats.</p> <p>The file <em><strong>foa_eval.zip</strong></em>&nbsp;corresponds to audio data of the <strong>FOA</strong>&nbsp;recording format for the evaluation dataset.<br> The file <em><strong>mic_eval.zip</strong></em>&nbsp;corresponds to audio data of the <strong>MIC</strong>&nbsp;recording format for the evaluation dataset.<br> The file <em><strong>video_eval.zip</strong></em>&nbsp;contains the common videos for both audio formats of the evaluation dataset.</p> <p>Download the zip files corresponding to the format of interest and use your favourite compression tool to unzip these zip files.</p>

openmit-licenseMar 2023View details →
zenodo40/100

Amyloid-motif-dependent tau self-assembly is modulated by isoform sequence context

<p><span>The microtubule-associated protein tau is implicated in neurodegenerative diseases characterized by amyloid formation. Mutations associated with frontotemporal dementia increase tau aggregation propensity and disrupt its endogenous microtubule-binding activity. However, the structural relationship between aggregation propensity and biological activity remains unclear. We employed a multi-disciplinary approach, including computational modeling, NMR, cross-linking mass spectrometry, and cell models to engineer tau sequences that modulate its structural ensemble. Our findings show that substitutions near the conserved 'PGGG' </span><span>&beta;</span><span>-turn motif informed by tau isoform context reduce tau aggregation in vitro and cells and can even counteract aggregation from disease-associated proline-to-serine mutations. Engineered tau sequences maintain microtubule binding and explain why 3R isoforms exhibit reduced pathogenesis compared to 4R. We propose a simple mechanism to reduce the formation of pathogenic tau species while preserving biological function, thus offering insights for therapeutic strategies aimed at reducing tau protein misfolding in neurodegenerative diseases.</span></p> <p><strong>Description of Source Data and Supplementary Data</strong>: All MD, NMR (peptide and tauRD), ThT, XL-MS, MT stabilization, MT:tau modeling, and cell-based aggregation data are available in the Source_Data directory as Data S1, Data S2, Data S3, Data S4, Data S5, Data S6, Data S7, and Data S8, respectively. Supplementary Data for raw MD trajectory files, structure files for MSM modeling and validation file, and tau:MT modeling are available as "Supplementary_Data_MD_Trajectories_Structures", "Supplementary_Data_MSM_models_Structures" and "Supplementary_Data_MT-tau_complex_models", respectively."</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

HDAC6 screening dataset using tau-based substrate in an enzymatic assay yields selective inhibitors and activators

<p><strong>Structure and information of the data file</strong></p> <p>DATA SET; Contains the information to which data set this information belongs. There are four possibilities denoted 1 to 4. Data set1: Enzymatic assay of human HDAC6 with commercial peptide substrate. Data set2: Enzymatic assay of human HDAC6 with custom peptide substrate. Data set3: Hit confirmation of the active molecules of the enzymatic assay of human HDAC6 with custom peptide substrate. Data set4: Determination of IC50 values for inhibition of enzymatic assay of human HDAC6 with custom peptide substrate.</p> <p>INTERNAL NAME; An internal name which enables identification of the compound within data sets from Fraunhofer ITMP ScreeningPort.</p> <p>TYPE; Type of data. Either &#39;inhibition&#39; for normalized inhibition values or &#39;IC50&#39; for enzymatic IC50.</p> <p>RELATION; Relation between TYPE and VALUE, always &#39;=&#39;.</p> <p>VALUE; Value of the normalized inhibition or the enzymatic IC50.</p> <p>UNITS; Unit of the value. Either &#39;%&#39; for the normalized inhibition or &#39;uM&#39; for the enzymatic IC50.</p> <p>NAME; Trade name of the chemical compound.</p> <p>SMILES; The canonical Smile of the chemical compound.</p> <p>&nbsp;</p> <p><strong>A</strong><strong>bstract</strong></p> <p>Histone deacetylase 6 (HDAC6) and HDAC10 are unique among the other HDACs as they consist of two domains instead of one. Only in the case of HDAC6 both domains are active resulting in a number of unique deacetylase reactions. Interestingly, HDAC6 can regulate the microtubule network and plays a role in the degradation of misfolded and aggregated proteins. We therefore developed a substrate (Boc-Ile-Asp-(Dimethyl)Lys-(Ac)Lys-aminoluciferin) based on a critical acetylation site of misfolded human Tau, a hallmark of Alzheimer&rsquo;s Disease. This substrate was used to screen a 5632 compound encompassing repurposing library at 10 &micro;M in a coupled, luminescence based assay. The assay was miniaturised to 10 &micro;L per enzymatic reaction. For comparison, a generic HDAC substrate (BOC-Gly-(Ac)Lys-aminoluciferin) was also used to screen the same library. Both substrates rely on a cascade of enzymatic reactions. First, HDAC6 deacetylates the substrate followed by cleavage of aminluciferin from the peptide by porcine Trypsin and conversion of the aminoluciferin using firefly Luciferase. Compounds with an activity of at least 75% inhibition against the custom human Tau based substrate were confirmed in triplicates at the screening concentration of 10 &micro;M. Confirmed hits, activity of at least 75%, where analysed in 8 point or 15 point dose response curves, depending on their activity. The data presented here encompass both primary data sets including 5632 compounds as well as 249 values from hit confirmation screening against the hTau based substrate and 151 IC<sub>50</sub> values from confirmed hits.</p> <p>&nbsp;</p> <p><strong>Methods of data generation</strong></p> <p><strong>Enzymatic assay of human HDAC6 with commercial peptide substrate. </strong></p> <p>The assay using the commercial peptide substrate (BOC-Gly-(Ac)Lys-aminoluciferin) was obtained from Promega Inc.. In the beginning the assay buffer is thawed and the lyophilized substrate is dissolved according to the technical manual (Promega Inc.) to create the substrate reagent. HDAC6 (obtained from BPS Biosciences) is dissolved in assay buffer at 0.2 nM, which is twice the final assay concentration. Compounds and controls are added to the plates using acoustic dispensing to reach a final concentration of 10 &micro;M in the assay followed by 5 &micro;l enzyme solution per well. Plates are centrifuged shortly and incubated for 10 min at RT. Afterwards, 5 &micro;L/well substrate solution are added to the wells, centrifuged shortly and incubated for 10 min prior detection of the luminescence signal on a multimode reader. Primary screening was done at one concentration (10 &micro;M) in singlicates.</p> <p>&nbsp;</p> <p><strong>Enzymatic assay of human HDAC6 with custom peptide substrate. </strong></p> <p>The assay was designed based on a commercial HDAC6 assay available from Promega Inc. This luminescence assay works by an aminoluciferin coupled HDAC6 peptide substrate. Upon deacetylation of the peptidic substrate by HDAC6 (obtained from BPS Biosciences) Trypsin (obtained from Sigma-Aldrich) can cleave the aminoluciferine from the peptide which can be converted by Luciferase (obtained from AAT Bioquest) to the detected signal. First, a twofold concentrated enzyme solution was generated, consisting of 4 nM HDAC6 and 0.1% BSA in HEPES buffer (25 mM HEPES, 137 mM NaCl, 2.7 mM KCl and 1 mM MgCl2, pH 7.0). Second, a twofold peptide solution was generated containing 100 &micro;M custom made peptide (Boc-Ile-Asp-(Dimethyl)Lys-(Ac)Lys-aminoluciferin) in HEPES buffer. Compounds and controls are added to the plates using acoustic dispensing to reach a final concentration of 10 &micro;M in the assay followed by 5 &micro;l enzyme solution per well. Plates are centrifuged shortly and 5 &micro;L/well peptide solution are added to the wells, centrifuged shortly and incubated for 30 min at RT. Afterwards, 5 &micro;L detection reagent (0.067 mg/mL Luciferase, 133.3 &micro;M ATP, 0.133 mg/mL Trypsin in HEPES buffer) were added to each well. Plates were centrifuged shortly and measured on a multimode reader after 30 min incubation at RT in the dark. Primary screening was done at one concentration (10 &micro;M) in singlicates.</p> <p>&nbsp;</p> <p><strong>Hit confirmation of the active molecules of the enzymatic assay of human HDAC6 with custom peptide substrate</strong></p> <p>The assay was designed based on a commercial HDAC6 assay available from Promega Inc. This luminescence assay works by an aminoluciferin coupled HDAC6 peptide substrate. Upon deacetylation of the peptidic substrate by HDAC6 (obtained from BPS Biosciences) Trypsin (obtained from Sigma-Aldrich) can cleave the aminoluciferine from the peptide which can be converted by Luciferase (obtained from AAT Bioquest) to the detected signal. First, a twofold concentrated enzyme solution was generated, consisting of 4 nM HDAC6 and 0.1% BSA in HEPES buffer (25 mM HEPES, 137 mM NaCl, 2.7 mM KCl and 1 mM MgCl2, pH 7.0). Second, a twofold peptide solution was generated containing 100 &micro;M custom made peptide (Boc-Ile-Asp-(Dimethyl)Lys-(Ac)Lys-aminoluciferin) in HEPES buffer. Compounds and controls are added to the plates using acoustic dispensing to reach a final concentration of 10 &micro;M in the assay followed by 5 &micro;l enzyme solution per well. Plates are centrifuged shortly and 5 &micro;L/well peptide solution are added to the wells, centrifuged shortly and incubated for 30 min at RT. Afterwards, 5 &micro;L detection reagent (0.067 mg/mL Luciferase, 133.3 &micro;M ATP, 0.133 mg/mL Trypsin in HEPES buffer) were added to each well. Plates were centrifuged shortly and measured on a multimode reader after 30 min incubation at RT in the dark. Hit confirmation was done at one concentration (10 &micro;M) in triplicates.</p> <p>&nbsp;</p> <p><strong>Determination of IC50 values for inhibition of enzymatic assay of human HDAC6 with custom peptide substrate</strong></p> <p>The assay was designed based on a commercial HDAC6 assay available from Promega Inc. This luminescence assay works by an aminoluciferin coupled HDAC6 peptide substrate. Upon deacetylation of the peptidic substrate by HDAC6 (obtained from BPS Biosciences) Trypsin (obtained from Sigma-Aldrich) can cleave the aminoluciferine from the peptide which can be converted by Luciferase (obtained from AAT Bioquest) to the detected signal. First, a twofold concentrated enzyme solution was generated, consisting of 4 nM HDAC6 and 0.1% BSA in HEPES buffer (25 mM HEPES, 137 mM NaCl, 2.7 mM KCl and 1 mM MgCl2, pH 7.0). Second, a twofold peptide solution was generated containing 100 &micro;M custom made peptide (Boc-Ile-Asp-(Dimethyl)Lys-(Ac)Lys-aminoluciferin) in HEPES buffer. Compounds and controls are added to the plates using acoustic dispensing to reach a final concentration of 10 &micro;M in the assay followed by 5 &micro;l enzyme solution per well. Plates are centrifuged shortly and 5 &micro;L/well peptide solution are added to the wells, centrifuged shortly and incubated for 30 min at RT. Afterwards, 5 &micro;L detection reagent (0.067 mg/mL Luciferase, 133.3 &micro;M ATP, 0.133 mg/mL Trypsin in HEPES buffer) were added to each well. Plates were centrifuged shortly and measured on a multimode reader after 30 min incubation at RT in the dark. IC50 values were determined using 7 point dose response curves (DRCs) between 20 &micro;M and 312 nM. In case inhibition values were not below 50% additional 7 point DRCs were measured, starting at 312 nm with a dilution factor of 2. All DRCs were recorded in triplicates.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Dataset related to article. "Nonphosphorylated tau slows down Aβ1–42 aggregation, binds to Aβ1–42 oligomers, and reduces Aβ1–42 toxicity"

<p>Figures presented in the paper and their data can be found in the corresponding prism files (www.graphpad.com).<br> The raw data and the picture files for most of the figures can be found in the folder with the corresponding figure name.</p>

opencc-by-4.0Apr 2021View details →
zenodo40/100

An Aligned Orbit for the Young Planet V1298 Tau b

<p>Code, data, and MCMC chains associated with the article &quot;An Aligned Orbit for the Young Planet V1298 Tau b,&quot; by Johnson et al. 2022 (accepted to The Astronomical Journal). Pre-print available at <a href="https://arxiv.org/abs/2110.10707">arXiv</a>.</p> <p>&nbsp;</p>

openmit-licenseMar 2022View details →
zenodo40/100

Dataset: Alpha Tau Medical Ltd. (DRTSW) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Alpha Tau Medical Ltd. (DRTS) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record