Skip to main content
zenodoopen

Verbal derivational suffixes in Hungarian: -(s)Odik and -(s)Ul

<p>This is an open-source dataset containing more than 1.1 million corpus occurrences of the Hungarian verbal derivational suffixes -(s)Odik and -(s)Ul. Both suffixes are used to create intransitive verbs from nominal bases, and they both mean &#39;to become [adjective/noun]&#39;. This dataset is suited for quantitative investigations into the subtle differences regarding how and when these suffixes are used. It consists of the following columns:</p> <ul> <li>1 <em>id</em>: ID</li> <li>2 <em>form</em>: lowercase word form</li> <li>3 <em>lemma</em>: word without inflectional suffixes; if the verb has a (separated) preverb, there is a + sign between the preverb and the verb stem</li> <li>4 <em>prev</em>: preverb associated with the verb</li> <li>5 <em>prevtype</em>: PFX if the preverb is prefixed to the verb, SEP if the preverb is separated</li> <li>6 <em>verb</em>: verb lemma; in each case without preverb</li> <li>7 <em>root</em>: adjective or noun serving as the base of verb formation</li> <li>8 <em>suffix</em>: derivational suffix: <em>-ul/&uuml;l/sul/s&uuml;l</em> endings are represented by -(s)Ul, <em>-odik/edik/&ouml;dik/sodik/sedik/s&ouml;dik</em> endings are represented by -(s)Odik</li> <li>9 <em>w2v_cluster</em>: the cluster ID of the root, based on word2vec embedding</li> <li>10 <em>argframe_cases</em>: arguments of the verb, represented by case-endings</li> <li>11 <em>argframe_long</em>: arguments of the verb, represented by lemma + case-ending combinations</li> <li>12 <em>doc_year</em>: the year of writing or the year of publication, 0 if unknown</li> <li>13 <em>doc_style</em>: document style</li> <li>14 <em>doc_id</em>: document identifier</li> <li>15 <em>left_context</em>: text preceding the hit</li> <li>16 <em>kwic</em>: the hit</li> <li>17 <em>right_context</em>: text following the hit</li> <li>18 <em>freqsum</em>: token frequency of the verb lemma; occurrences with and without preverbs are counted together</li> <li>19 <em>prev_vs_all</em>: token frequency of the verb lemma with any preverb, divided by the &#39;freqsum&#39; value</li> <li>20 <em>actprev_vs_allprev</em>: token frequency of the specific preverb + verb lemma combination, divided by the &#39;prev_vs_all&#39; value</li> </ul> <p>The first row stands for the header. If a cell&#39;s value is unspecified, it is marked with underscore (_).</p>

ShareScore

24/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
8
Reuse readiness
0
Engagement
4

Topics