Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

14

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

14 results for “entity linking”

Learn how ShareScore rates datasets ↗
zenodo48/100

French Entity-Linking dataset between annotated tweets collected during major crises in France and French Wikipedia corpus

<p>Most of the available datasets are not particularly adapted to our target application: geolocate natural disasters from social networks. First, social media posts are largely underrepresented in these datasets, and the only Twitter dataset lacks Entity-Linking annotations. Second, none of the datasets focuses on a crisis or natural disaster event.</p> <p>To mitigate these issues, we extracted a collection of French tweets written during earthquakes and major floods that have occurred in France in recent years. We set up Label-Studio in order to annotate these tweets. A total of 4617 tweets were annotated, including 1678 tweets posted during earthquakes and 2939 during floods. For each annotated tweet, mentions were annotated using the set of labels described earlier in the paper as well as, when possible, the target Wikipedia title.</p> <p>Named &ldquo;R&eacute;SoCIO&rdquo; in reference to the research project in which it was carried out, the dataset resulting from this work contains a total of 12 828 annotated mentions and 1 513 distinct Wikipedia entities. 85% of mentions were associated with a Wikipedia page and 94 % if we ignore the RISKNAT and DAMAGES labels, which are often difficult to map to an existing entity.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entities</strong></td> </tr> <tr> <td>PERSON</td> <td>315</td> <td>263</td> <td>136</td> </tr> <tr> <td>ORG</td> <td>863</td> <td>790</td> <td>281</td> </tr> <tr> <td>GEOLOC</td> <td>4375</td> <td>4234</td> <td>701</td> </tr> <tr> <td>TRANSPORT</td> <td>250</td> <td>203</td> <td>101</td> </tr> <tr> <td>EVENT</td> <td>35</td> <td>21</td> <td>16</td> </tr> <tr> <td>FACILITY</td> <td>129</td> <td>94</td> <td>49</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>128</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>223</td> <td>200</td> <td>46</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>12828</strong></td> <td><strong>1322</strong></td> <td><strong>1513</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the Twitter dataset. #Mentions&nbsp;shows the total number of mentions per label, #Linked the number of mentions linked&nbsp;to an entity and #Entities the number of distinct entities per label present in the&nbsp;dataset.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entitie</strong>s</td> </tr> <tr> <td>PERSON</td> <td>1100102</td> <td>1098406</td> <td>557697</td> </tr> <tr> <td>ORG</td> <td>750925</td> <td>749504</td> <td>130394</td> </tr> <tr> <td>GEOLOC</td> <td>2729702</td> <td>2728296</td> <td>215924</td> </tr> <tr> <td>TRANSPORT</td> <td>161539</td> <td>160487</td> <td>53405</td> </tr> <tr> <td>EVENT</td> <td>798433</td> <td>798251</td> <td>86471</td> </tr> <tr> <td>FACILITY</td> <td>258835</td> <td>258513</td> <td>109867</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>127</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>4340621</td> <td>4339658</td> <td>682458</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>10146795</strong></td> <td><strong>10138230</strong></td> <td><strong>1836399</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the full dataset. #Mentions shows&nbsp;the total number of mentions per label, #Linked the number of mentions linked to an&nbsp;entity and #Entities the number of distinct entities per label present in the dataset.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Dataset for the paper "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia"

<p>Dataset for the EMNLP'24 Main conference paper titled "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia".</p>

opencc-by-sa-4.0Oct 2024View details →
zenodo40/100

EvaNIL: silver standard dataset for large-scale NIL entity linking evaluation

<p>The EvaNIL dataset can be used to train or evaluate approaches developed for NIL entity linking. It was built from several Biomedical and Life Sciences corpora:</p> <ul> <li>PubMed DS</li> <li>CRAFT corpus</li> <li>MedMentions</li> </ul> <p>These corpora contain entities associated with knowledge base concepts. To build the EvaNIL dataset, we assumed that those knowledge base concepts did not exist in the respective knowledge bases, so each entity is associated instead with the direct ancestors of those original concepts.</p> <p>The EvaNIL dataset is divided into 6 partitions including annotations from several knowledge bases:</p> <ul> <li>&quot;medic&quot; (CTD-MEDIC)</li> <li>&quot;ctd_anatomy&quot; (CTD-Anatomy)</li> <li>&quot;ctd_chemicals&quot; (CTD-Chemicals)</li> <li>&quot;chebi&quot; (ChEBI)</li> <li>&quot;go_bp&quot; (GO-Biological Process)</li> <li>&quot;hp&quot; (HPO)</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Dataset for Named Entity Recognition and Entity Linking from Greek Wikipedia Events

<p>An automated benchmark dataset for (Named Entity Recognition) NER and (Named Entity Linking) NEL tools, based on Greek Wikipedia events pages.</p> <p>Note: This data includes data from the following sources:<br> - Wikipedia el.wikipedia.org</p> <p><strong>Description</strong></p> <p>The dataset is provided in the form of&nbsp; three&nbsp; JSON-formatted subsets i.e.,&nbsp; train, validation and test in an analogy of 70-20-10. The current version of the dataset contains 18,617 events annotated with 40,798 entity mentions and 36,189 links to elWikipedia (and wikidata ids). The dataset contains annotations belonging to 8 entity types: person, organization, location, gpe, event, facility, product and work of art.</p> <table> <caption>Overall dataset statistics</caption> <thead> <tr> <th scope="col">&nbsp;</th> <th scope="col">Docs</th> <th scope="col">Tokens</th> <th scope="col">Sentences</th> <th scope="col">Surface Mentions</th> <th scope="col">Valid Links</th> <th scope="col">Red Links</th> </tr> </thead> <tbody> <tr> <td><strong>Train</strong></td> <td>13,031</td> <td>332,077</td> <td>16,927</td> <td>28,593</td> <td>25,365</td> <td>3,228</td> </tr> <tr> <td><strong>Validation</strong></td> <td>3,722</td> <td>94,746</td> <td>4,844</td> <td>8,168</td> <td>7,240</td> <td>928</td> </tr> <tr> <td><strong>Test</strong></td> <td>1,862</td> <td>47,450</td> <td>2,427</td> <td>4,037</td> <td>3,584</td> <td>453</td> </tr> <tr> <td><strong>Total</strong></td> <td>18,617</td> <td>474,361</td> <td>24,200</td> <td>40,798</td> <td>36,189</td> <td>4,609</td> </tr> </tbody> </table> <p><strong>Example</strong></p> <p>A record example is given below.</p> <p>{</p> <p>&quot;json_file&quot;: &quot;February 2012_39_0 events&quot;,<br> &quot;text&quot;: &quot;Sudan and South Sudan sign non-aggression pact.&quot;,<br> &quot;ground_truth_mentions&quot;: [<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {&quot;start&quot;: 0, &quot;end&quot;: 4, &quot;surface_mention&quot;: &quot;Sudan&quot;, &quot;mention_type&quot;: &quot;GPE&quot;},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {&quot;start&quot;: 10, &quot;end&quot;: 20, &quot;surface_mention&quot;: &quot;South Sudan&quot;, &quot;mention_type&quot;: &quot;GPE&quot;}<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ],<br> &quot;ground_truth_links&quot;: [<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {&quot;enwiki&quot;: &quot;Sudan&quot;,&quot;wikidata&quot;: &quot;Q1049&quot;},<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {&quot;enwiki&quot;: &quot;South_Sudan&quot;, &quot;wikidata&quot;: &quot;Q958&quot;}<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ]<br> }</p> <p><strong>Code</strong></p> <p><a href="https://gitlab.isl.ics.forth.gr/debatelab/elwiki_events_benchmark">https://gitlab.isl.ics.forth.gr/debatelab/elwiki_events_benchmark</a></p> <p><strong>Acknowledgments</strong></p> <p>This work has received funding from the Hellenic Foundation for Research and Innovation (HFRI) and the General Secretariat for Research and Technology (GSRT), under grant agreement No 4195.</p>

opencc-by-3.0Dec 2022View details →
zenodo40/100

Berlin State Library (2023). Named Entity Disambiguation German Database for the Named Entity Linking System of the Berlin State Library (SBB)

<p>This database is a part of a BERT-based entity recognition and three-stage entity linking (EL) system. Its components consist of three models as well as three related databases, one of which is published here.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Berlin State Library (2023). Named Entity Disambiguation French Database for the Named Entity Linking System of the Berlin State Library (SBB)

<p>This database is a part of a BERT-based entity recognition and three-stage entity linking (EL) system. Its components consist of three models as well as three related databases, one of which is published here.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Berlin State Library (2023). Named Entity Disambiguation English Database for the Named Entity Linking System of the Berlin State Library (SBB)

<p>This database is a part of a BERT-based entity recognition and three-stage entity linking (EL) system. Its components consist of three models as well as three related databases, one of which is published here.</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Evaluation results of the xMEN entity linking toolkit for multiple benchmark datasets

Open the record for dataset details and reuse information.

publicDec 2024View details →
zenodo36/100

Reddit Entity Linking

<p>An entity linking dataset created from the social media website,&nbsp;Reddit. The dataset contains&nbsp;619 posts and 1,243 corresponding comments that were selected and given to human annotators. Three different human annotators were used to annotate each grouping of text.&nbsp;&nbsp;The resulting mentions and entities are included with a breakdown of the inter-annotator agreement between the various mention-entity pairs.</p> <p>The mention-entity pairs collected are broken into groups based on the level of inter-annotator agreement.&nbsp;</p> <p>Gold annotations - all three agree</p> <p>Silver annotations - two out of three annotators agree</p> <p>Bronze annotations - an individual annotator&#39;s annotation that the other two did not have</p> <p>In total the dataset contains 1,342&nbsp;gold annotations, 2,723&nbsp;silver annotations, and 7,038&nbsp;bronze annotations.</p> <p>A readme file is provided that describes the structure of the files and the information within&nbsp;each one.</p>

opencc-by-4.0May 2020View details →
zenodo36/100

Benchmark for the evaluation of Named Entity Linking over ancient documents

<p><strong>Benchmark for the evaluation of Named Entity Linking over ancient documents</strong><br> Elvys Linhares Pontes, Ahmed Hamdi, Nicolas Sidere, and Antoine Doucet<br> University of Avignon: elvys.linhares-pontes@univ-avignon.fr; University of La Rochelle: {elvys.linhares_pontes,ahmed.hamdi,nicolas.sidere,antoine.doucet}@univ-lr.fr</p> <p>These are the supplementary materials for the ICADL 2019 paper <strong><em>Impact of OCR Quality on Named Entity Linking</em></strong>. If you end up using whole or parts of this resource, please use the following citation:</p> <ul> <li>Linhares Pontes, E., Hamdi, A., Sidere, N., and Doucet, A. (2019). Impact of OCR Quality on Named Entity Linking. In Proceedings of 21st International Conference on Asia-Pacific Digital Libraries ICADL 2019, Kuala Lumpur, Malaysia.</li> </ul> <p>or alternatively use the following `bib`:</p> <pre><code>@inproceedings{linhares2019icadl, title="Impact of OCR Quality on Named Entity Linking.", author={Linhares Pontes, Elvys, and Hamdi, Ahmed, and Sidere, Nicolas, and Doucet, Antoine}, year={2019}, booktitle={Proceedings of 21st International Conference on Asia-Pacific Digital Libraries ICADL 2019} }</code></pre> <p><strong>Files</strong><br> This archive contains six folders -- one per dataset -- as well as this README. The folders contain the degraded images, the noisy texts extracted by the OCR and their aligned version with clean data. This work is licensed under a [Creative Commons Attribution-ShareAlike 4.0 International License](http://creativecommons.org/licenses/by-sa/4.0/).</p> <p><strong>Acknowledgments</strong><br> This work has been supported by the European Union&#39;s Horizon 2020 research and innovation programme under grant 770299 [NewsEye](https://www.newseye.eu/).</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Datasets for Out-of-KB Mention Discovery with Entity Linking

<p>The repository contains datasets for out-of-KB mention discovery from texts, documented in the work, <em>Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking</em>, on arXiv: <a href="https://arxiv.org/abs/2302.07189">https://arxiv.org/abs/2302.07189</a> (CIKM 2023).</p> <p>Each data setting (as a sub-folder) contains train, valid, and test files and also 100 random sample files for each data split for debugging.</p> <p>Data folder names with &ldquo;syn_full&rdquo; at the end are synonym augmented data (each synonym as an entity) for the setting.</p> <p>Ontology .jsonl files have two versions for each, &quot;syn_attr&quot; setting treats synonyms are attributes, &quot;syn_full&quot; setting treats synonyms as entities.</p> <p>&nbsp;</p> <p>Data scripts are available at <a href="https://github.com/KRR-Oxford/BLINKout#data-scripts">https://github.com/KRR-Oxford/BLINKout#data-scripts</a></p> <p>&nbsp;</p> <p>Acknowledgement of the data sources below:</p> <p>ShARe/CLEF 2013 dataset is from <a href="https://physionet.org/content/shareclefehealth2013/1.0/">https://physionet.org/content/shareclefehealth2013/1.0/</a></p> <p>MedMention dataset is from <a href="https://github.com/chanzuckerberg/MedMentions">https://github.com/chanzuckerberg/MedMentions</a></p> <p>UMLS (versions 2012AB, 2014AB, 2017AA) is from <a href="https://www.nlm.nih.gov/research/umls/index.html">https://www.nlm.nih.gov/research/umls/index.html</a></p> <p>SNOMED CT (corresponding versions) is from <a href="https://www.nlm.nih.gov/healthit/snomedct/index.html">https://www.nlm.nih.gov/healthit/snomedct/index.html</a></p> <p>NILK dataset is from <a href="https://zenodo.org/record/6607514">https://zenodo.org/record/6607514</a></p> <p>WikiData 2017 dump is from <a href="https://archive.org/download/enwiki-20170220/enwiki-20170220-pages-articles.xml.bz2">https://archive.org/download/enwiki-20170220/enwiki-20170220-pages-articles.xml.bz2</a></p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Tough Tables: Carefully Evaluating Entity Linking for Tabular Data

<p>Tough Tables (2T) is a dataset designed to evaluate table annotation approaches in solving the CEA and CTA tasks.<br> The dataset is compliant with the data format used in <a href="https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2019/index.html">SemTab 2019</a>, and it can be used as an additional dataset without any modification. The target knowledge graph is DBpedia 2016-10.<br> Check out the <a href="https://github.com/vcutrona/tough-tables">2T GitHub repository</a> for more details about the dataset generation.</p> <p><strong>New in v3.0:</strong> We release the updated version of 2T! The target knowledge graphs are DBpedia <a href="http://downloads.dbpedia.org/wiki-archive/downloads-2016-10.html">2016-10</a> and Wikidata <a href="https://zenodo.org/record/6643443">20220521</a>. Starting from this version, the dataset is split into valid and test sets.</p> <p>This work is based on the following paper:</p> <blockquote> <p>Cutrona, V., Bianchi, F., Jimenez-Ruiz, E. and Palmonari, M. (2020). Tough Tables: Carefully Evaluating Entity Linking for Tabular Data. ISWC 2020, LNCS 12507, pp. 1&ndash;16.</p> </blockquote> <p>Note on License: This dataset includes data from the following sources. Refer to each source for license details:<br> - Wikipedia https://www.wikipedia.org/<br> - DBpedia https://dbpedia.org/<br> - Wikidata https://www.wikidata.org/<br> - SemTab 2019 https://doi.org/10.5281/zenodo.3518539<br> - GeoDatos https://www.geodatos.net<br> - The Pudding https://pudding.cool/<br> - Offices.net https://offices.net/<br> - DATA.GOV https://www.data.gov/</p> <p>THIS DATA IS PROVIDED &quot;AS IS&quot;, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.<br> <br> <strong>Changelog:</strong></p> <p><em><strong>v3.0</strong></em></p> <ul> <li>Both datasets require SemTab2020 CEA format (tab_id, row_id, col_id, entity). <ul> </ul> </li> <li>Tables IDs and artificial noise values differ from previous versions.</li> <li>Datasets are split into Valid and Test sets of tables.</li> <li>New GT for ToughTables-WD (2T_WD) <ul> <li>Entities&nbsp;Q23772518 and&nbsp;Q7327323 have been removed because they are no longer represented in WD</li> <li>Updated ancestor/descendant hierarchies to evaluate CTA.</li> </ul> </li> <li>Evaluation scripts are provided with the data sets.</li> </ul> <p><em><strong>v2.0</strong></em></p> <ul> <li>New GT for 2T_WD <ul> <li>A few entities have been removed from the CEA GT, because they are no longer represented in WD (e.g., dbr:Devont&eacute; points to wd:Q21155080, which does not exist)</li> <li>Tables codes and values differ from the previous version, because of the random noise.</li> <li>Updated ancestor/descendant hierarchies to evaluate CTA.</li> </ul> </li> </ul> <p><em><strong>v1.0</strong></em></p> <ul> <li>New Wikidata version (2T_WD)</li> <li>Fix header for tables CTRL_DBP_MUS_rock_bands_labels.csv and CTRL_DBP_MUS_rock_bands_labels_NOISE2.csv (column 2 was reported with id 1 in target - NOTE: the affected column has been removed from the SemTab2020 evaluation)</li> <li>Remove duplicated entries in tables</li> <li>Remove rows with wrong values (e.g., the Kazakhstan entity has an empty name &quot;&#39;&#39;&quot;)</li> <li>Many rows and noised columns are shuffled/changed due to the random noise generator algorithm</li> <li>Remove row &quot;Florida&quot;,&quot;Floorida&quot;,&quot;New York, NY&quot; from TOUGH_WEB_MISSP_1000_us_cities.csv (and all its NOISE1 variants)</li> <li>Fix header of tables: <ul> <li>CTRL_WIKI_POL_List_of_current_monarchs_of_sovereign_states.csv</li> <li>CTRL_WIKI_POL_List_of_current_monarchs_of_sovereign_states_NOISE2.csv</li> <li>TOUGH_T2D_BUS_29414811_2_4773219892816395776___videogames_developers.csv</li> <li>TOUGH_T2D_BUS_29414811_2_4773219892816395776___videogames_developers_NOISE2.csv</li> </ul> </li> </ul> <p><em><strong>v0.1-pre</strong></em></p> <ul> <li>First submission. It contains only tables, without GT and Targets.</li> </ul>

opencc-by-4.0Nov 2020View details →
zenodo28/100

SynEL: A Synthetic and Semi-Synthetic Benchmark for Entity Linking in the Customer Support Domain

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo16/100

TweetNERD - End to End Entity Linking Benchmark for Tweets

<p>TweetNERD - End to End Entity Linking Benchmark for Tweets</p> <p><a href="https://arxiv.org/abs/2210.08129">Paper</a>&nbsp;- <a href="https://www.youtube.com/watch?v=H5ypIHterWQ">Video</a>&nbsp;- <a href="https://neurips.cc/virtual/2022/poster/55647">Neurips Page</a></p> <p>This is the dataset described in the paper <a href="https://arxiv.org/abs/2210.08129"><strong>TweetNERD - End to End Entity Linking Benchmark for Tweets</strong></a> (accepted to <a href="https://datasets-benchmarks-proceedings.neurips.cc/paper/2022">Thirty-sixth Conference on Neural Information Processing Systems (Neurips) Datasets and Benchmarks Track</a>).</p> <blockquote> <p>Named Entity Recognition and Disambiguation (NERD) systems are foundational for information retrieval, question answering, event detection, and other natural language processing (NLP) applications. We introduce TweetNERD, a dataset of 340K+ Tweets across 2010-2021, for benchmarking NERD systems on Tweets. This is the largest and most temporally diverse open sourced dataset benchmark for NERD on Tweets and can be used to facilitate research in this area.</p> </blockquote> <p><strong>UPDATE: The new version contains an additional ~125K Tweets leading to a total dataset size of ~465K Tweets.</strong></p> <p>TweetNERD dataset is released under <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0)</a> LICENSE.</p> <p>The license only applies to the data files present in this dataset. See <strong>Data usage policy</strong> below.</p> <p>Check out more details at <a href="https://github.com/twitter-research/TweetNERD">https://github.com/twitter-research/TweetNERD</a></p> <p><strong>Usage</strong></p> <p>We provide the dataset split across the following tab seperated files:</p> <ul> <li><strong>OOD.public.tsv</strong>: OOD split of the data in the paper.</li> <li><strong>Academic.public.tsv</strong>: Academic split of the data described in the paper.</li> <li><code>part_*.public.tsv</code>: Remaining data split into parts in no particular order.</li> </ul> <p>Each file is tab separated and has has the following format:</p> <table> <thead> <tr> <th>tweet_id</th> <th>phrase</th> <th>start</th> <th>end</th> <th>entityId</th> <th>score</th> </tr> </thead> <tbody> <tr> <td>22</td> <td>twttr</td> <td>20</td> <td>25</td> <td>Q918</td> <td>3</td> </tr> <tr> <td>21</td> <td>twttr</td> <td>20</td> <td>25</td> <td>Q918</td> <td>3</td> </tr> <tr> <td>1457198399032287235</td> <td>Diwali</td> <td>30</td> <td>38</td> <td>Q10244</td> <td>3</td> </tr> <tr> <td>1232456079247736833</td> <td>NO_PHRASE</td> <td>-1</td> <td>-1</td> <td>NO_ENTITY</td> <td>-1</td> </tr> </tbody> </table> <p>For tweets which don&#39;t have any entity, their column values for <code>phrase, start, end, entityId, score</code> are set <code>NO_PHRASE, -1, -1, NO_ENTITY, -1</code> respectively.</p> <p>Description of file columns is as follows:</p> <table> <thead> <tr> <th>Column</th> <th>Type</th> <th>Missing Value</th> <th>Description</th> </tr> </thead> <tbody> <tr> <td>tweet_id</td> <td>string</td> <td>&nbsp;</td> <td>ID of the Tweet</td> </tr> <tr> <td>phrase</td> <td>string</td> <td>NO_PHRASE</td> <td>entity phrase</td> </tr> <tr> <td>start</td> <td>int</td> <td>-1</td> <td>start offset of the phrase in text using <code>UTF-16BE</code> encoding</td> </tr> <tr> <td>end</td> <td>int</td> <td>-1</td> <td>end offset of the phrase in the text using <code>UTF-16BE</code> encoding</td> </tr> <tr> <td>entityId</td> <td>string</td> <td>NO_ENTITY</td> <td>Entity ID. If not missing can be NOT FOUND, AMBIGUOUS, or Wikidata ID of format Q{numbers}, e.g. Q918</td> </tr> <tr> <td>score</td> <td>int</td> <td>-1</td> <td>Number of annotators who agreed on the phrase, start, end, entityId information</td> </tr> </tbody> </table> <p>In order to use the dataset you need to utilize the <code>tweet_id</code> column and get the Tweet text using the <a href="https://developer.twitter.com/en/docs/twitter-api">Twitter API</a> (See <strong>Data usage policy</strong> section below).</p> <p>Data stats</p> <table> <thead> <tr> <th>Split</th> <th>Number of Rows</th> <th>Number unique tweets</th> </tr> </thead> <tbody> <tr> <td>OOD</td> <td>34102</td> <td>25000</td> </tr> <tr> <td>Academic</td> <td>51685</td> <td>30119</td> </tr> <tr> <td>part_0</td> <td>11830</td> <td>10000</td> </tr> <tr> <td>part_1</td> <td>35681</td> <td>25799</td> </tr> <tr> <td>part_2</td> <td>34256</td> <td>25000</td> </tr> <tr> <td>part_3</td> <td>36478</td> <td>25000</td> </tr> <tr> <td>part_4</td> <td>37518</td> <td>24999</td> </tr> <tr> <td>part_5</td> <td>36626</td> <td>25000</td> </tr> <tr> <td>part_6</td> <td>34001</td> <td>24984</td> </tr> <tr> <td>part_7</td> <td>34125</td> <td>24981</td> </tr> <tr> <td>part_8</td> <td>32556</td> <td>25000</td> </tr> <tr> <td>part_9</td> <td>32657</td> <td>25000</td> </tr> <tr> <td>part_10</td> <td>32442</td> <td>25000</td> </tr> <tr> <td>part_11</td> <td>32033</td> <td>24972</td> </tr> <tr> <td>part_12</td> <td>76559</td> <td>25000</td> </tr> <tr> <td>part_13</td> <td>67240</td> <td>24920</td> </tr> <tr> <td>part_14</td> <td>67745</td> <td>25000</td> </tr> <tr> <td>part_15</td> <td>67652</td> <td>25000</td> </tr> <tr> <td>part_16</td> <td>65739</td> <td>25000</td> </tr> </tbody> </table> <p>Data usage policy</p> <p>Use of this dataset is subject to you obtaining lawful access to the <a href="https://developer.twitter.com/en/docs/twitter-api">Twitter API</a>, which requires you to agree to the <a href="https://developer.twitter.com/en/developer-terms/">Developer Terms Policies and Agreements</a>.</p> <p>Please cite the following if you use TweetNERD in your paper:</p> <pre>@dataset{TweetNERD_Zenodo_2022_6617192, author = {Mishra, Shubhanshu and Saini, Aman and Makki, Raheleh and Mehta, Sneha and Haghighi, Aria and Mollahosseini, Ali}, title = {{TweetNERD - End to End Entity Linking Benchmark for Tweets}}, month = jun, year = 2022, note = {{Data usage policy Use of this dataset is subject to you obtaining lawful access to the [Twitter API](https://developer.twitter.com/en/docs /twitter-api), which requires you to agree to the [Developer Terms Policies and Agreements](https://developer.twitter.com/en /developer-terms/).}}, publisher = {Zenodo}, version = {0.0.0}, doi = {10.5281/zenodo.6617192}, url = {https://doi.org/10.5281/zenodo.6617192} } @inproceedings{TweetNERDNeurips2022, author = {Mishra, Shubhanshu and Saini, Aman and Makki, Raheleh and Mehta, Sneha and Haghighi, Aria and Mollahosseini, Ali}, booktitle = {Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks}, pages = {}, title = {TweetNERD - End to End Entity Linking Benchmark for Tweets}, volume = {2}, year = {2022}, eprint = {arXiv:2210.08129}, doi = {10.48550/arXiv.2210.08129} } </pre>

restrictedJun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record