Skip to main content
zenodoopen

Places Audio Captions (Japanese) 100k

<p>The Places Audio Caption (Japanese) 100K Corpus contains approximately 100,000 Japanese&nbsp;spoken captions for natural images drawn from the Places 205 image dataset.</p> <p>This&nbsp;speech&nbsp;corpus&nbsp;was&nbsp;collected to investigate the learning of spoken language (words, sub-word units, higher-level semantics, etc.) from visually-grounded speech.&nbsp;For a description of the corpus, see:</p> <pre><code>@INPROCEEDINGS{Ohishi2020trilingual, author={Ohishi, Yasunori and Kimura, Akisato and Kawanishi, Takahito and Kashino, Kunio and Harwath, David and Glass, James}, booktitle={ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, title={Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms}, year={2020}, pages={4352-4356}, }</code></pre> <p>The&nbsp;corpus only includes audio recordings, and not the associated images. You will need to separately download the Places image dataset&nbsp;<a href="http://places.csail.mit.edu/">here</a>.</p> <p>The data&nbsp;is&nbsp;distributed under the Creative Commons Attribution-ShareAlike (CC BY-SA) license&nbsp;<a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">(link)</a>.</p> <p>If you use this data in your own publications, please cite the paper above.</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics