Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

878

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

878 results for “Corpus”

Learn how ShareScore rates datasets ↗
zenodo44/100

Spanish Biomedical Crawled Corpus

<p>The largest Spanish biomedical and heath corpus to date gathered from a massive Spanish health domain crawler over more than 3,000 URLs were downloaded and preprocessed. All the collected data have been preprocessed to produce the CoWeSe (Corpus Web Salud Espa&ntilde;ol) resource, a large-scale and high-quality corpus intended for biomedical and health NLP in Spanish.</p> <p>Enlarged version with less restrictive document and sentence deduplication.</p> <p><strong>Citation</strong></p> <p>If you use this resource in your work, please cite our paper:</p> <pre>@misc{carrino2021spanish, title={Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models}, author={Casimiro Pio Carrino and Jordi Armengol-Estap&eacute; and Ona de Gibert Bonet and Asier Guti&eacute;rrez-Fandi&ntilde;o and Aitor Gonzalez-Agirre and Martin Krallinger and Marta Villegas}, year={2021}, eprint={2109.07765}, archivePrefix={arXiv}, primaryClass={cs.CL} } </pre> <p>Copyright (c) 2022 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Argument Aspect Corpus - Nuclear Energy

<p>The Argument Aspect Corpus&ndash;Nuclear Energy (AAC-NE) contains English-language sentences with aspect annotations describing the content of arguments on the topic of nuclear energy.</p> <p>It was introduced in this paper:</p> <blockquote> <p>Jurkschat, L., Wiedemann, G., Heinrich, M., Ruckdeschel, M., &amp; Torge, S. (2022). Few-Shot Learning for Argument Aspects of the Nuclear Energy Debate. In Proceedings of the 13th International Conference on Language Resources and Evaluation (LREC 2022). European Language Resources Association (ELRA).</p> </blockquote> <p>The AAC-NE corpus is based on a subset of all argumentative sentences contained in the UKP SAM dataset [1] for which a majority vote of three annotators could be achieved during the annotation of the main argument aspect of each sentence.</p> <p>The CSV files contain one of nine aspect labels per argumentative sentence split into training, dev, and test set.</p> <table> <thead> <tr> <th><strong>aspect</strong></th> <th><strong>train</strong></th> <th><strong>dev</strong></th> <th><strong>test</strong></th> <th><strong>Sum</strong></th> <th><strong>Kripp. Alpha</strong></th> </tr> </thead> <tbody> <tr> <td>alternatives</td> <td>100</td> <td>16</td> <td>21</td> <td>137</td> <td>0.69</td> </tr> <tr> <td>costs</td> <td>98</td> <td>17</td> <td>29</td> <td>144</td> <td>0.72</td> </tr> <tr> <td>environment</td> <td>209</td> <td>27</td> <td>64</td> <td>300</td> <td>0.74</td> </tr> <tr> <td>innovation</td> <td>33</td> <td>2</td> <td>8</td> <td>43</td> <td>0.38</td> </tr> <tr> <td>reactor safety</td> <td>112</td> <td>17</td> <td>43</td> <td>172</td> <td>0.59</td> </tr> <tr> <td>reliability</td> <td>47</td> <td>5</td> <td>10</td> <td>62</td> <td>0.36</td> </tr> <tr> <td>waste</td> <td>87</td> <td>5</td> <td>26</td> <td>118</td> <td>0.80</td> </tr> <tr> <td>weapons</td> <td>52</td> <td>11</td> <td>15</td> <td>78</td> <td>0.77</td> </tr> <tr> <td>other</td> <td>120</td> <td>23</td> <td>29</td> <td>172</td> <td>0.49</td> </tr> <tr> <td><strong>all</strong></td> <td><strong>858</strong></td> <td><strong>123</strong></td> <td><strong>245</strong></td> <td><strong>1226</strong></td> <td><strong>0.62</strong></td> </tr> <tr> <td>pro</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>706</td> <td>&nbsp;</td> </tr> <tr> <td>cons</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>520</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>Additionally, it contains 2000 unlabeled sentences with presumably argumentative content sampled from the newspaper &ldquo;The Guardian&rdquo;.</p> <p>[1] Stab, C., Miller, T., Schiller, B., Rai, P., &amp; Gurevych, I. Cross-topic Argument Mining from Heterogeneous Sources. In E. Riloff, D. Chiang, J. Hockenmaier, &amp; J. Tsujii (Eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 3664&ndash;3674). Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/D18-1402">https://doi.org/10.18653/v1/D18-1402 </a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Corpus of the Epigraphy of the Italian Peninsula in the 1st Millennium BCE

<p>The Corpus of the Epigraphy of the Italian Peninsula in the 1st Millennium BCE, or CEIPoM, is a linguistic database focusing on the Italian peninsula in the first millennium BCE. Currently, it covers Messapic, Venetic, the Sabellic languages and epigraphic Latin up to about 100 BCE.</p>

opencc-by-sa-4.0Sep 2021View details →
zenodo44/100

German Climate Change Tweet Corpus (GerCCT)

<p>First release&nbsp;of the GerCCT Corpus, a German tweet resource annotated for argument components, argument properties, sarcasm and toxic language.</p> <p>The corpus consists of 1,200 tweets and its annotations. Each tweet is associated with its respective source tweet, i.e. the tweet it replies to. Source tweets were used to provide annotators with additional context. The annotations refer to the reply tweet, i.e. NOT to the source tweet. For copyright reasons we cannot distribute the actual tweet content. Instead we share the source and reply tweet IDs and the annotations.</p> <p>The current version includes class annotations on the document level, i.e. on the tweet level. We are working on creating the respective span annotations.</p>

opencc-by-sa-4.0Apr 2022View details →
zenodo44/100

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria&mdash;Hausa, Igbo, Nigerian-Pidgin, and Yor&ugrave;b&aacute;&mdash;consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Metadata and annotation data for the XSample corpus on German academic language

<p>The XSample corpus war created in the project <em>XSample</em> (https://www.izus.uni-stuttgart.de/fokus/fdm-projekte/xsample/) at Universit&auml;t Stuttgart in 2021 by Melanie Andresen and Axel Pichler. It contains 135 German academic journal articles, 45 each from the disciplines linguistics, literary studies and philosophy. The texts themselves cannot be made public for copyright reasons. However, metadata and some annotation data are published here.</p> <p><strong>xsample-metadata.csv</strong><br> This file contains metadata on the texts in the corpus, like journal, title, authors, text length, and the URL to the original paper. It also contains two analytical metrics, &#39;past-ratio&#39; and &#39;temp-expr-ratio&#39;, that are based on the annotations in the other two files. The variable &#39;past-ratio&#39; expresses the proportion of verbs in past tense relative to all finite verbs in the text. The variable &#39;temp-expr-ratio&#39; gives the number of temporal expressions per 1000 token.</p> <p><strong>xsample-heidel.csv</strong><br> This file contains all temporal expressions found and classified by the annotation tool <em>HeidelTime</em> (https://github.com/HeidelTime/heideltime, V. 2.2.1, Str&ouml;tgen &amp; Gertz 2013 ). The variable &#39;position&#39; expresses the position of the first character of the temporal expression in the text in characters.</p> <p><strong>xsample-sticker2.csv</strong><br> This file contains all finite verbs found and classified by the annotation tool <em>sticker2</em> (https://github.com/stickeritis/sticker2). The variable &#39;position&#39; expresses the position of the first character of the finite verb in the text in characters.</p> <p><strong>References</strong><br> Str&ouml;tgen, Jannik &amp; Michael Gertz. 2013. Multilingual and cross-domain temporal tagging. <em>Language Resources and Evaluation</em>. Springer 47(2). 269&ndash;298. <a href="https://doi.org/10.1007/s10579-012-9179-y">https://doi.org/10.1007/s10579-012-9179-y</a>.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

CafeteriaSA corpus: Scientific abstracts annotated across different food semantic resources

<p>In the last decades, a great amount of work has been done in predictive modeling of issues related to human and environmental health. Resolution of issues related to healthcare is made possible by the existence of several biomedical vocabularies and standards, which play a crucial role in understanding health information, together with a large amount of health-related data. However, despite the large number of available resources and work done in the health and environmental domains, there is a lack of semantic resources that can be utilized in the food and nutrition domain, as well as their interconnections. For this purpose, in an European Food Safety Authority-funded project CAFETERIA, we have developed the first annotated corpus of 500 scientific abstracts that consists of 6,407 annotated food entities with regard to Hansard taxonomy, 4,299 for FoodOn, and 3,623 for SNOMED-CT.&nbsp; The CafeteriaSA corpus will enable further development of natural language processing methods for food information extraction from textual data that will allow extracting of food information from scientific textual data.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - Italian Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at&nbsp;<a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - Spanish Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at&nbsp;<a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - German Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - French Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at&nbsp;<a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - Dutch Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at&nbsp;<a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Polifonia Corpus - Books Module Metadata - English Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

CafeteriaFCD corpus: Food consumption data annotated with regard to different food semantic resources

<p>The FoodBase curated version which contains 1,000 manually evaluated recipes, annotated with the appropriate semantic tags from the Hansard Taxonomy, FoodON and SNOMED-CT.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

The Threatening English Language (TEL) Corpus

<p>TEL is the Threatening English Language corpus. It is a collection of 309 written texts compiled from the publicly-available portion of CTARC (the Communicated Threat Assessment Research Corpus, compiled by Tammy Gales), MFT (the Malicious Forensic Texts corpus, compiled by Andrea Nini), and the written portion of CoJO (the Corpus of Judicial Opinions, compiled by Julia Muschalik). Additional texts are from ForensicLing.com (the forensic linguistic data site hosted by Tammy Gales and Dakota Wing). Basic metadata is supplied for each text where known from the original case research. We wish to thank our graduate student fellows who helped compile the texts and metadata: Nicole Harris, Annina van Riper, Zara Rabinko, and Zachary Boudreaux.</p> <p>Total texts: 309<br> Total estimated authors: 203<br> Total word count: 54,167</p> <p>METADATA KEY</p> <p>TG = Tammy Gales (public portion of CTARC)<br> AN = Andrea Nini (MFT)<br> JM = Julia Muschalik (written portion of CoJo)<br> FL = ForensicLing.com (Tammy Gales and Dakota Wing)</p> <p>Name###_## = file name, case number, text number within case<br> File name might be threat recipient or author; remaining info is about the author, where known</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Corpus and list of keywords from Improving sustainable crop protection using population genetics concepts

<p>Corpus extracted in April 2021 from the ISI Web of Science portal (https://www.webofscience.com) with the following request: &lsquo;Plant AND Resistan* AND Durab*&rsquo;. A first corpus of 2522 articles was built considering all publication years for this extraction. This collection was then refined by categories to remove articles outwith the scope of our search (e.g. related to durable resistant materials for constructions). We also kept only articles cited at least once. The final corpus was composed of 1783 articles from 1979 to 2021:</p> <ul> <li>CORPUS_plant_resistance_durability.zip</li> </ul> <p>List of keywords used for the network presented in the article:</p> <ul> <li>keywords_list.csv</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Corpus de revistas Proyecto Digitization and Analysis of Cultural Transfers in Colombian Literary Magazines (1892–1950)

<p>Coprpus de revistas Proyecto Digitization and Analysis of Cultural Transfers in Colombian Literary Magazines (1892&ndash;1950)</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Experimental AI corpus from OpenAlex

<p>A corpus of AI research from OpenAlex. Includes:</p> <ul> <li>A works table with metadata about AI papers</li> <li>An authors&nbsp;table with information about the authors</li> <li>An institutions table with information about institutions</li> <li>A concepts table with information about concepts in works</li> <li>A MeSH table with information about MeSH terms in works</li> <li>A concepts json with the OpenAlex concept taxonomy</li> <li>An abstracts json with deinverted abstracts</li> <li>A citations json with citations from papers</li> </ul> <p>See `ai_openalex_description.md` for data dictionaries.</p> <p>See `ai_openalex_methodology.md` for a description of the method used to create the dataset.</p> <p>See here for additional information:&nbsp;<a href="https://github.com/nestauk/ai_genomics">https://github.com/nestauk/ai_genomics</a></p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

The COUGHVID crowdsourcing dataset: A corpus for the study of large-scale cough analysis algorithms

<p><strong>Overview</strong></p> <p>Cough audio signal classification has been successfully used to diagnose a variety of respiratory conditions, and there has been significant interest in leveraging Machine Learning (ML) to provide widespread COVID-19 screening. The COUGHVID dataset provides over 30,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses. Furthermore, experienced pulmonologists labeled more than 2,000 recordings to diagnose medical abnormalities present in the coughs, thereby contributing one of the largest expert-labeled cough datasets in existence that can be used for a plethora of cough audio classification tasks.&nbsp;As a result, the COUGHVID dataset contributes a wealth of cough recordings for training ML models to address the world&rsquo;s most urgent health crises.</p> <p><strong>Private Set and Testing Protocol</strong></p> <p>Researchers interested in testing their models on the private test dataset should contact us at coughvid@epfl.ch, briefly explaining the type of validation they wish&nbsp;to make, and their obtained results obtained through&nbsp;cross-validation with the public data. Then, access to the unlabeled recordings will be provided, and&nbsp;the researchers should&nbsp;send the predictions of their models on these recordings. Finally,&nbsp;the&nbsp;performance metrics of the predictions will be sent to the researchers. The private testing data is not included in any file within our Zenodo record, and it can only be accessed by contacting the COUGHVID team at the aforementioned e-mail address.</p> <p><strong>New Semi-Supervised Labeling</strong></p> <p>The third version of the COUGHVID dataset contains thousands of additional recordings obtained through October 2021. Additionally, the recordings containing coughs were re-labeled according to a semi-supervised learning algorithm that combined the user labels with those of the expert physicians, which were&nbsp;modeled using ML and expanded on the previously unlabeled data. These labels can be found in the &quot;status_SSL&quot; column of the &quot;metadata_compiled.csv&quot; file.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Corpus of Decisions: Permanent Court of International Justice (CD-PCIJ)

<p>&nbsp;</p> <p><strong>Overview</strong></p> <p>The <strong>Corpus of Decisions: Permanent Court of International Justice (CD-PCIJ)</strong> collects and presents for the first time in human- and machine-readable formats all documents of PCIJ Series A, B and A/B of the <a href="https://www.icj-cij.org/en/pcij">Permanent Court of International Justice (PCIJ)</a>. Among these are judgments, advisory opinions, orders, appended minority opinions, annexes, applications instituting proceedings and requests for an advisory opinion. The International Court of Justice, the successor of the PCIJ, has kindly made available these documents <a href="https://www.icj-cij.org/en/pcij">on its website</a>.</p> <p>The <a href="https://www.icj-cij.org/en/pcij">Permanent Court of International Justice (PCIJ)</a> was the primary judicial organ of the League of Nations, the ill-fated predecessor of the United Nations, which existed from 1920 to 1946. Nonetheless, as the first international court with general thematic jurisdiction, the PCIJ influenced international law in profound ways that are still felt today. Every lawyer who sets out on the path of international law encounters epoch-defining opinions such as the <em>Lotus</em> and <em>Factory at Chorz&oacute;w</em> decisions, but the Court&#39;s lesser-known jurisprudence and the appended minority opinions offer many more ideas and legal principles which are seldom appreciated today.</p> <p>This data set is designed to be complementary to and fully compatible with the <strong><a href="https://doi.org/10.5281/zenodo.3826445">Corpus of Decisions: International Court of Justice (CD-ICJ)</a></strong>, which is also available open access.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>A peer-reviewed academic paper describing the construction and relevance of the data set entitled <strong><a href="https://doi.org/10.1111/jels.12313">&#39;Introducing Twin Corpora of Decisions for the International Court of Justice (ICJ) and the Permanent Court of International Justice (PCIJ)&#39;</a></strong> was published open access in the <a href="https://doi.org/10.1111/jels.12313">Journal of Empirical Legal Studies (JELS)</a>. It is also available in print at JELS 2022, Vol. 19, No. 2, pp. 491-524.</p> <p>If you use the data set for academic work, please cite both the JELS paper and the precise version of the data set you used for your analysis.</p> <p>&nbsp;</p> <p><strong>NEW in Version 1.1.0</strong></p> <ul> <li>Full recompilation of data set</li> <li>CHANGELOG and README converted to external markdown files</li> <li>Display of version number on Codebook and Compilation Report title pages fixed; correctly display semantic versioning</li> <li>The ZIP archive of source files includes the TEX files</li> <li>Config file converted to TOML format</li> <li>All R packages are version-controlled with {renv}</li> <li>Data set creation process cleans up all files from previous runs before a new data set is created</li> <li>Remove redundant color from violin plots</li> </ul> <p>&nbsp;</p> <p><strong>Updates</strong></p> <p>The CD-PCIJ will only be updated if <a href="https://github.com/SeanFobbe/cd-pcij/issues">errors are discovered</a>, enhancements are developed or in the unlikely event that the Court publishes additional documents within the collection ambit of the data set (PCIJ Series A, B and A/B).</p> <p>Notifications regarding new and updated data sets will be published on my academic website at <a href="https://seanfobbe.com/">www.seanfobbe.com</a> or via Mastodon at <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <p><strong>Recommended Variants</strong></p> <table summary="Recommended variants"> <thead> <tr> <th scope="col">Target Audience</th> <th scope="col">Recommended Variant</th> </tr> </thead> <tbody> <tr> <td><em>Practitioners</em></td> <td>PDF_ENHANCED_MajorityOpinions</td> </tr> <tr> <td><em>Traditional Scholars</em></td> <td>PDF_ENHANCED_FULL</td> </tr> <tr> <td><em>Quantitative Analysts</em></td> <td>CSV_TESSERACT_FULL</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Please refer to the Codebook regarding the relative merits of each variant. All variants are available in either English or French. Unless you have very specific needs you should only use the variants denoted &#39;ENHANCED&#39; or &#39;TESSERACT&#39; for serious work.</p> <p>&nbsp;</p> <p><strong>Features</strong></p> <ul> <li>Fully compatible with the <a href="https://doi.org/10.5281/zenodo.3826445">Corpus of Decisions: International Court of Justice (CD-ICJ)</a></li> <li>29 variables</li> <li>Public Domain (CC-Zero 1.0)</li> <li>Open and platform independent file formats (PDF, TXT, CSV)</li> <li>Extensive Codebook</li> <li><a href="https://zenodo.org/record/7051937/files/CD-PCIJ_1-1-0_CompilationReport.pdf?download=1">Compilation Report</a> explains construction and validation of the data set in detail</li> <li>Large number of diagrams for all purposes (see the &#39;ANALYSIS&#39; archive)</li> <li>Diagrams are available as PDF (for printing) and PNG (for web display), tables are available as CSV for easy readability by humans and machines</li> <li>Secure cryptographic signatures</li> <li><a href="https://doi.org/10.5281/zenodo.7051937">Publication of full source code (Open Source)</a></li> </ul> <p>&nbsp;</p> <p><strong>Key Metrics</strong></p> <p><em>Version:</em> 1.1.0</p> <p><em>Temporal Coverage</em>: 22 May 1922 &ndash; 26 February 1940</p> <p><em>Documents:</em> 259 (English) / 261 (French)</p> <p><em>Tokens:</em>&nbsp;1,296,536 (English) /&nbsp;1,262,184 (French)</p> <p><em>Formats:</em> PDF, TXT, CSV</p> <p>&nbsp;</p> <p><strong>Source Code and Compilation Report</strong></p> <p>With every compilation of the full data set an <a href="https://zenodo.org/record/7051937/files/CD-PCIJ_1-1-0_CompilationReport.pdf?download=1">extensive Compilation Report</a> is created in a professionally layouted PDF format (comparable to the Codebook). The Compilation Report includes the Source Code, comments and explanations of design decisions, relevant computational results, exact timestamps and a table of contents with clickable internal hyperlinks to each section. The Compilation Report and Source Code are published under the same DOI: <a href="https://doi.org/10.5281/zenodo.7051937">https://doi.org/10.5281/zenodo.7051937</a></p> <p>For details of the construction and validation of the data set please refer to the Compilation Report.</p> <p>&nbsp;</p> <p><strong>Disclaimer</strong></p> <p>This data set has been created by Mr Se&aacute;n Fobbe using documents available on the website of the International Court of Justice (<a href="https://www.icj-cij.org">https://www.icj-cij.org</a>). It is a personal academic initiative and is not associated with or endorsed by the International Court of Justice or the United Nations.</p> <p>The Court accepts no responsibility or liability arising out of my use, or that of third parties, of the documents and information produced, used or published on the Zenodo website. Neither the Court nor its staff members nor its contractors may be held responsible or liable for the consequences, financial or otherwise, resulting from the use of these documents and information.</p> <p>&nbsp;</p> <p><strong>Academic Publications (Fobbe)</strong></p> <p>Website &mdash; <a href="https://www.seanfobbe.com">www.seanfobbe.com</a></p> <p>Open Data &mdash; <a href="https://zenodo.org/communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data</a></p> <p>Code Repository &mdash; <a href="https://zenodo.org/communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code</a></p> <p>Regular Publications &mdash; <a href="https://zenodo.org/communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications</a></p> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>Did you discover any errors? Do you have suggestions on how to improve the data set? You can either post these to the <a href="https://github.com/SeanFobbe/cd-pcij/issues">Issue Tracker on GitHub</a> or write me an e-mail at <a href="mailto:fobbe-data@posteo.de">fobbe-data@posteo.de</a></p> <p>&nbsp;</p>

opencc-zeroFeb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record