Skip to main content
zenodoopen

COALA voice data and transcripts Italian

<p>This dataset contains audio files and transcripts in Italian and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut f&uuml;r Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years&nbsp; (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: &quot;Have you ever worked in a factory, manufacturing setting?&quot; (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: &quot;If yes, for how many years?&quot; (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
8

Topics