Places Audio Captions (Japanese) 100k
<p>The Places Audio Caption (Japanese) 100K Corpus contains approximately 100,000 Japanese spoken captions for natural images drawn from the Places 205 image dataset.</p> <p>This speech corpus was collected to investigate the learning of spoken language (words, sub-word units, higher-level semantics, etc.) from visually-grounded speech. For a description of the corpus, see:</p> <pre><code>@INPROCEEDINGS{Ohishi2020trilingual, author={Ohishi, Yasunori and Kimura, Akisato and Kawanishi, Takahito and Kashino, Kunio and Harwath, David and Glass, James}, booktitle={ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, title={Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms}, year={2020}, pages={4352-4356}, }</code></pre> <p>The corpus only includes audio recordings, and not the associated images. You will need to separately download the Places image dataset <a href="http://places.csail.mit.edu/">here</a>.</p> <p>The data is distributed under the Creative Commons Attribution-ShareAlike (CC BY-SA) license <a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">(link)</a>.</p> <p>If you use this data in your own publications, please cite the paper above.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0