Skip to main content
zenodoopen

Sequence Tagging of FA-KES Dataset

<p>We used the BIOE Sequence Tagging&nbsp;strategy which was utilized in OpenTag [1] in the aim of getting every word in the dataset<br> associated with a label called &lsquo;tag&rsquo;. A tag consists of one of these letters B, I, O, or E, that stand respectively for beginning, inside, outside, or end of an attribute, followed by a &lsquo;-&rsquo; sign, followed by three letters that represent the type of information that was initially&nbsp;extracted. Tokens&nbsp;were labeled with one of the following tags: &lsquo;B-LOC&rsquo;, &lsquo;I-LOC&rsquo;, &lsquo;E-LOC&rsquo;, &lsquo;B-CIV&rsquo;, &lsquo;I-CIV&rsquo;, &lsquo;E-CIV&rsquo;, &lsquo;B-NCV&rsquo;, &lsquo;I-NCV&rsquo;, &lsquo;E-NCV&rsquo;, &lsquo;B-WMN&rsquo;, &lsquo;I-WMN&rsquo;, &lsquo;E-WMN&rsquo;, &lsquo;B-CHD&rsquo;, &lsquo;I-CHD&rsquo;, &lsquo;E-CHD&rsquo;, &lsquo;B-ACT&rsquo;, &lsquo;I-ACT&rsquo;,&lsquo;E-ACT&rsquo;, &lsquo;B-COD&rsquo;, &lsquo;I-COD&rsquo;, &lsquo;E-COD&rsquo;, &lsquo;B-DAT&rsquo;, &lsquo;I-DAT&rsquo;, &lsquo;E-DAT&rsquo;, or &lsquo;O&rsquo; (where O stands for words outside the scope, LOC for&nbsp;the incident location, CIV for the number of civilians dead, NCV for the number of non-civilians dead, WMN for the number of women targeted, CHD for the number of children killed, ACT for actor/authority responsible for the incident, COD for the cause of death, and DAT for date of incident). This was done by creating a parser that would automatically tag each word with the appropriate tag.</p> <p>We created three subsets of the FA-KES dataset. The first one consists of the articles&#39; titles, the second one&nbsp;of the articles&#39; titles concatenated with the articles&#39; first paragraphs, and the third one consists of the articles&#39; titles along with their contents. The first column in each of the three CSV files&nbsp;represents the article number in the dataset, the second column contains the sequence of&nbsp;words for each article and the third one holds&nbsp;the tags linked to&nbsp;the tokens&nbsp;of&nbsp;the previous column.</p> <p>[1]:&nbsp;G. Zheng, S. Mukherjee, X. L. Dong, and F. Li, &ldquo;Opentag: Open attribute value extraction from product profiles,&rdquo; CoRR, vol. abs/1806.01264, 2018.</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
16
Reuse readiness
0
Engagement
0

Topics