Verbal derivational suffixes in Hungarian: -(s)Odik and -(s)Ul
<p>This is an open-source dataset containing more than 1.1 million corpus occurrences of the Hungarian verbal derivational suffixes -(s)Odik and -(s)Ul. Both suffixes are used to create intransitive verbs from nominal bases, and they both mean 'to become [adjective/noun]'. This dataset is suited for quantitative investigations into the subtle differences regarding how and when these suffixes are used. It consists of the following columns:</p> <ul> <li>1 <em>id</em>: ID</li> <li>2 <em>form</em>: lowercase word form</li> <li>3 <em>lemma</em>: word without inflectional suffixes; if the verb has a (separated) preverb, there is a + sign between the preverb and the verb stem</li> <li>4 <em>prev</em>: preverb associated with the verb</li> <li>5 <em>prevtype</em>: PFX if the preverb is prefixed to the verb, SEP if the preverb is separated</li> <li>6 <em>verb</em>: verb lemma; in each case without preverb</li> <li>7 <em>root</em>: adjective or noun serving as the base of verb formation</li> <li>8 <em>suffix</em>: derivational suffix: <em>-ul/ül/sul/sül</em> endings are represented by -(s)Ul, <em>-odik/edik/ödik/sodik/sedik/södik</em> endings are represented by -(s)Odik</li> <li>9 <em>w2v_cluster</em>: the cluster ID of the root, based on word2vec embedding</li> <li>10 <em>argframe_cases</em>: arguments of the verb, represented by case-endings</li> <li>11 <em>argframe_long</em>: arguments of the verb, represented by lemma + case-ending combinations</li> <li>12 <em>doc_year</em>: the year of writing or the year of publication, 0 if unknown</li> <li>13 <em>doc_style</em>: document style</li> <li>14 <em>doc_id</em>: document identifier</li> <li>15 <em>left_context</em>: text preceding the hit</li> <li>16 <em>kwic</em>: the hit</li> <li>17 <em>right_context</em>: text following the hit</li> <li>18 <em>freqsum</em>: token frequency of the verb lemma; occurrences with and without preverbs are counted together</li> <li>19 <em>prev_vs_all</em>: token frequency of the verb lemma with any preverb, divided by the 'freqsum' value</li> <li>20 <em>actprev_vs_allprev</em>: token frequency of the specific preverb + verb lemma combination, divided by the 'prev_vs_all' value</li> </ul> <p>The first row stands for the header. If a cell's value is unspecified, it is marked with underscore (_).</p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 0
- Engagement
- 4