Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

219

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

219 results for “wikipedia”

Learn how ShareScore rates datasets ↗
zenodo44/100

Translated Wikipedia Biographies (English-Catalan)

<p>A professional translation of the Translated Wikipedia Biographies dataset, designed to analyze common gender errors in machine translation like incorrect gender choices in anaphora resolutions, possessives and gender agreement, commissioned by BSC LangTech Unit.</p> <p>License type: CC BY-SA 4.0</p> <p>This work was funded by the Departament de la Vicepresid&egrave;ncia i de Pol&iacute;tiques Digitals i Territori de la Generalitat de Catalunya within the framework of Projecte AINA.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Ludopaedia ! Atelier contributif Wikipedia autour du jeu et des jouets dans l'Antiquité. Université de Fribourg/ERC Locus Ludi, 23.10.2019

<p>23.10.2019. Universit&eacute; de Fribourg. Formation doctorale Swiss Universities/ERC Locus Ludi. Ludopaedia ! Atelier contributif Wikipedia &nbsp; autour du jeu et des jouets dans l&rsquo;Antiquit&eacute;. 23 octobre 2019. R&eacute;alisation:&nbsp;Vision is Mind Prod.&nbsp;&nbsp;Musique: Edwan&nbsp;<a href="https://youtube.com/edwanmusic">https://youtube.com/edwanmusic</a></p> <p>Les contenus &laquo;&nbsp;libres&nbsp;&raquo; disponibles en ligne &agrave; propos des jeux et des pratiques ludiques dans l&rsquo;Antiquit&eacute; sont peu nombreux et incomplets. Le but du projet &laquo;&nbsp;Ludopaedia&nbsp;&raquo; est de contribuer &agrave; cr&eacute;er un r&eacute;seau nourri d&rsquo;articles et de f&eacute;d&eacute;rer une communaut&eacute; active autour des jeux, des jouets et des pratiques sociales li&eacute;es &agrave; ces activit&eacute;s dans l&rsquo;Antiquit&eacute;. Ce premier atelier organis&eacute; en lien avec le projet ERC Locus Ludi |<em>The Cultural Fabric of Play and Games in Classical Antiquity</em> et l&rsquo;exposition &laquo; Ludique ! Jouer dans l&rsquo;Antiquit&eacute; &raquo; &agrave; Lyon jusqu&rsquo;au 1er d&eacute;cembre 2019 s&rsquo;adresse aux &eacute;tudiants et aux chercheurs sensibles aux probl&eacute;matiques de l&rsquo;Open Access en milieu acad&eacute;mique. L&rsquo;objectif est de contribuer&nbsp;ensemble &agrave; am&eacute;liorer le contenu des pages existantes en ajoutant des r&eacute;f&eacute;rences bibliographiques et des illustrations de bonne qualit&eacute; sur Wikipedia et sur Commons.</p> <p>https://youtu.be/lkmfzy0uaUk</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

Long document similarity datasets, Wikipedia excerptions for movies, video games and wine collections

<p>Three&nbsp;corpora in different domains extracted from Wikipedia.</p> <p>For all datasets, the figures and tables have been filtered out, as well as the categories and &quot;see also&quot; sections.</p> <p>The article structure, and&nbsp;particularly the sub-titles and paragraphs are kept in these datasets</p> <p>&nbsp;</p> <p><strong>Wines</strong></p> <p>Wikipedia wines dataset consists of 1635 articles from the wine domain. The extracted dataset consists of a non-trivial mixture of articles, including different wine categories, brands, wineries, grape types, and more. The ground-truth recommendations were crafted by a human sommelier, which annotated 92 source articles with ~10 ground-truth recommendations for each sample. Examples for ground-truth expert-based recommendations are&nbsp;</p> <ul> <li>Dom P&eacute;rignon - Mo&euml;t &amp; Chandon</li> <li>Pinot Meunier - Chardonnay</li> </ul> <p><strong>Movies</strong></p> <p>The Wikipedia movies dataset consists of 100385 articles describing different movies. The movies&#39; articles may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.<br> For this dataset, we have extracted a test set of ground truth annotations for 50 source articles using the &quot;<a href="https://bestsimilar.com/">BestSimilar</a>&quot;&nbsp;database. Each source articles is associated with a list of ${\scriptsize \sim}12$ most similar movies.<br> Examples for ground-truth expert-based recommendations are&nbsp;</p> <ul> <li>Schindler&#39;s List - The Pianist</li> <li>Lion King - The Jungle Book</li> </ul> <p><strong>Video games</strong></p> <p>The Wikipedia video games dataset consists of 21,935 articles reviewing video games from all genres and consoles. Each article may consist of a different combination of sections, including summary, gameplay, plot, production, etc. Examples for ground-truth expert-based recommendations are:</p> <ul> <li>Grand Theft Auto - Mafia</li> <li>Burnout Paradise - Forza Horizon 3</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Similar Sentences from English Wikipedia 20150304

<p>Similar sentences from a dump of enwiki from March 04, 2015 detected using the Wikiduper software (https://github.com/seweissman/wikiduper).</p>

opencc-by-4.0Jun 2015View details →
zenodo40/100

A word2vec model file built from the French Wikipedia XML Dump using gensim.

<p>A word2vec model file built from the French Wikipedia XML dump using gensim. The data published here includes three model files (you need all three of them in the same folder) as well as the Python script used to build the model (for documentation). The Wikipedia dump was downloaded on October 7, 2016 from https://dumps.wikimedia.org/. Before building the model, plain text was extracted from the dump. The size of that dataset is about 500 million words or 3.6 GB of plain text. The principal parameters for building the model were the following: no lemmatization was performed, tokenization was done using the "\W" regular expression (any non-word character splits tokens), and the model was built with 500 dimensions.</p>

opencc-by-4.0Oct 2016View details →
zenodo40/100

English Wikipedia

<p>This text corpus is composed of texts of English Wikipedia extracted from the Wikipedia dump of 26th September 2015 using the WikiExtractor tool (https://github.com/attardi/wikiextractor). </p>

opencc-by-4.0Jan 2017View details →
zenodo40/100

Data for: Wikipedia as a gateway to biomedical research

<p>Wikipedia has been described as a gateway to knowledge. However, the extent to which this gateway ends at Wikipedia or continues via supporting citations is unknown. This dataset was used&nbsp;to establish benchmarks for the relative distribution and referral (click) rate of citations, as indicated by presence of a Digital Object Identifier (DOI), from Wikipedia with a focus on medical citations.</p> <p>This data set includes for each day in August 2016 a listing of all DOI present in the English language version of Wikipedia and whether or not the DOI are biomedical in nature. Source Code for these data are available at: Ryan Steinberg. (2017, July 9). Lane-Library/wiki-extract: initial Zenodo/DOI release. Zenodo. http://doi.org/10.5281/zenodo.824813</p> <p>This dataset also includes a listing from Crossref DOIs that were referred from Wikipedia in August 2016 (Wikipedia_referred_DOI). Source code for these data sets is available at:&nbsp;Joe Wass. (2017, July 4). CrossRef/logppj: Initial DOI registered release. Zenodo. http://doi.org/10.5281/zenodo.822636&nbsp;</p> <p>An&nbsp;article based on this data was published in PLOS One:</p> <p>Maggio LA, Willinsky JM, Steinberg RM, Mietchen D, Wass JL, Dong T. Wikipedia as a gateway to biomedical research: The relative distribution and use of citations in the English Wikipedia. PloS one. 2017 Dec 21;12(12):e0190046.&nbsp;</p> <p>https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0190046&nbsp;</p>

opencc-zeroJul 2017View details →
dryad40/100

Causal evidence for social group sizes from Wikipedia editing data

<p>Human communities have self-organizing properties in which specific Dunbar Numbers may be invoked to explain group attachments.  By analyzing Wikipedia editing histories across a wide range of subject pages, we show that there is an emergent coherence in the size of transient groups formed to edit the content of subject texts, with two peaks averaging at around $N=8$ for the size corresponding to maximal  contention, and at around $N=4$ as a regular team. These values are consistent with the observed sizes of conversational groups, as well as the hierarchical structuring of Dunbar graphs.  We use the Promise Theory model of bipartite trust to derive a scaling law that  fits the data and may apply to all group size distributions, when based on attraction to a seeded group process.  In addition to  providing further evidence that even spontaneous communities of strangers are self-organizing, the results have important implications for the governance of the Wikipedia commons and for the security of all online social platforms and associations.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Cross-language Wikipedia link graph

<p>Wikipedia articles use Wikidata to list the links to the same article in other language versions. Therefore, each Wikipedia language edition stores the Wikidata Q-id for each article.</p> <p>This dataset constitutes a Wikipedia link graph where all the article identifiers are normalized to Wikidata Q-ids. It contains the normalized links from all Wikipedia language versions. Detailed link count statistics are attached. Note that articles that have no incoming nor outgoing links are not part of this graph.</p> <p>The format is as follows:</p> <p>Q-id of linking page (outgoing) &lt;tab&gt; Q-id of linked page (incoming) &lt;tab&gt; language version - dump date (20241101)</p> <p>This dataset was used to compute <a href="https://danker.s3.amazonaws.com/index.html">Wikidata PageRank</a>. More information can be found on the <a href="https://github.com/athalhammer/danker">danker</a> repository, where the source code of the link extraction as well as the PageRank computation is hosted.</p> <p>Example entries:<br><br>$ bzcat 2024-11-06.allwiki.links.bz2 | head</p> <p>1 &nbsp; &nbsp;107 &nbsp; &nbsp;ckbwiki-20241101<br>1 &nbsp; &nbsp;107 &nbsp; &nbsp;lawiki-20241101<br>1 &nbsp; &nbsp;107 &nbsp; &nbsp;ltwiki-20241101<br>1 &nbsp; &nbsp;107 &nbsp; &nbsp;tewiki-20241101<br>1 &nbsp; &nbsp;107 &nbsp; &nbsp;wuuwiki-20241101<br>1 &nbsp; &nbsp;111 &nbsp; &nbsp;hywwiki-20241101<br>1 &nbsp; &nbsp;11379 &nbsp; &nbsp;bat_smgwiki-20241101<br>1 &nbsp; &nbsp;11471 &nbsp; &nbsp;cdowiki-20241101<br>1 &nbsp; &nbsp;150 &nbsp; &nbsp;ckbwiki-20241101<br>1 &nbsp; &nbsp;150 &nbsp; &nbsp;lowiki-20241101</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-sa-3.0Oct 2022View details →
zenodo40/100

Wikipedia: wikipedia-lt (Lithuanian)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://lt.wikipedia.org/

opencc-by-sa-4.0Aug 2024View details →
zenodo40/100

Wikipedia: wikipedia-de (German)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>

opencc-by-sa-4.0Aug 2024View details →
zenodo40/100

Wikipedia: wikipedia-en (English)

<p>EOL harvests articles from wikipedia that are indexed as species or higher taxa.</p> <p>Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. The name Wikipedia is a portmanteau of the words wiki (a technology for creating collaborative websites, from the Hawaiian word wiki, meaning quick) and encyclopedia. Wikipedia articles provide links designed to guide the user to related pages with additional information.</p> <p>https://en.wikipedia.org</p>

opencc-by-sa-4.0Aug 2024View details →
zenodo40/100

Topic Model for English Wikipedia's Biographies with list of all 1.8M articles linked to Wikidata

<p>A Genism LDA Topic Model of English Wikipedia biographical articles with list of all 1.8M articles, and some associated Wikidata information</p> <p>The model has 150 Topics.</p> <p>This model was developed in the process of isolating a set of visual arts biographical articles, as described in &quot;Clowns in the Visual Artists: Topic Modeling Wikipedia and Wikidata&quot; in the Spring 2022 issue of <em>Art Documentation - </em><a href="https://doi.org/10.1086/719999">https://doi.org/10.1086/719999</a></p> <p>Because names, nationalities, and birthdays are so prominent in biographies, the stopwords list removed 170,000 names, surnames, city names, place names, countries, days, months and other time related words (https://github.com/mandiberg/Names-Surnames-and-Countries-for-Stopwords). &nbsp;We also directly removed each article subject&rsquo;s given and surname, which were almost always the most frequently occurring words in any given article. Otherwise, the model just produced topics based on nationality, and common names and surnames.</p> <p><strong>Files:</strong></p> <p>all_enwiki_bios_from_wikidata.csv<br> The list of all Wikidata items for humans with an enwiki page (e.g&nbsp;biographical article) was extracted from Wikidata JSON dump; list includes gender, occupation, and nationality. This was joined with the converted plaintext from an English Wikipedia dump. This data was downloaded in March 2021.</p> <p>Wikipedia Biographies LDA Topic Model human readable summary.csv<br> A human readable file with the 150 topics ranked by count of articles per topic from the 1.8M corpus. The most popular topics have categorical descriptions of the occupations of each cluster. Some are marked as not an occupation cluster.&nbsp;</p> <p>BoW_corpus.mm*<br> model_lda_full_Sep2_150Tv2*<br> These six files comprise the topic model. The code to load them is present in the python files.&nbsp;</p> <p>dict_full_Aug-28-2021<br> processed_docs_full_Aug-28-2021.txt<br> processed_docs_1000_Aug-18-2021.txt<br> These are the dictionary and processed corpuses required to build and implement the model using this code. The corpus with the first 1000 items is meant to be used for testing, as the full one is quite large and takes a long time to complete.&nbsp;</p> <p>topic-model-wikipedia-sept2021.zip<br> The code and settings used for creating and implementing this model are included in this zip and are also available here: https://github.com/mandiberg/topic-model-wikipedia</p> <p>All-Wikipedia-Biographies-with-topic1.csv<br> All-Wikipedia-Biographies-with-topic1and2.csv<br> These are the list of 1.8M biographies matched to topics. The &quot;topic1&quot; file just includes the first topic, this is a slightly larger list. The &quot;topic1and2&quot; file is slightly smaller because about 2% articles do not match to a second topic.</p> <p>Analysis-for-Clowns-Visual-Arts.zip<br> These are the raw data and final data produced for the &quot;Clowns in the Visual Artists.&quot; Please see the article for context.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Recent Revision History of Featured Articles on Wikipedia

<p>The dataset is&nbsp;composed of recent revisions (from July 21st 2020 to Dec 10th 2021) of 6,039 featured articles on Wikipedia. Each page is stored in a .json file with all associated revisions contained within, sorted from most recent to oldest. Each revision contains metadata as well as the full wikitext content.</p> <p>Extra fields were added to revisions:</p> <ul> <li>content-plaintext: Parsed wikitext content (using python&#39;s mwparserfromhell package)</li> <li>comment-plaintext: Parsed plaintext of comment</li> <li>content-diff: Sentence-level diffs of current revision&#39;s content&nbsp;and the previous revision&#39;s content, produced by the difflib&nbsp;python package (unified diff)</li> <li>content-plaintext-diff: Similar to content-diff, but with content-plaintext as input</li> <li>reverted: Indicates&nbsp;whether the revision was reverted by a later revision</li> </ul> <p>Total size after decompression: ~ 59.2 GB</p>

opencc-by-3.0Dec 2021View details →
zenodo40/100

Machine Assisted Translation of Wikipedia Articles into Low Resource Languages

<p><strong>Wikipedia is the largest encyclopedia ever assembled with the vision of enabling every human being to freely share in the sum of all knowledge. Wikipedia currently has a total of more than six million articles and over 17 billion words in its English edition. Unfortunately, millions of people cannot access this resource because it&rsquo;s not available in their language. For instance, at the moment there are only 218 Tigrinya Wikipedia and 15,018 Amharic Wikipedia articles.</strong></p> <p><strong>In this project, we investigate the problem of translating Wikipedia articles from a high resource language into low resource languages using human-in-the-loop MT systems. In particular, we investigate different approaches to translate a sample of English Wikipedia articles into Tigrinya and Amharic. Currently, this repository contains 100k English Wikipeida abstracts translated using Lesan (https://lesan.ai) into Amharic and Tigrinya.</strong><br> <br> &nbsp;</p> <p><strong>Structure of data directory:</strong></p> <p><strong>data<br> ├── human<br> └── mt<br> &nbsp; &nbsp; ├── google<br> &nbsp; &nbsp; ├── lesan<br> &nbsp; &nbsp; │&nbsp;&nbsp;&nbsp;├── am.txt<br> &nbsp; &nbsp; │&nbsp;&nbsp;&nbsp;├── en.txt<br> &nbsp; &nbsp; │&nbsp;&nbsp;&nbsp;└── ti.txt<br> &nbsp; &nbsp; └── microsoft</strong><br> &nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Dataset of first appearances of the scholarly bibliographic references on English Wikipedia articles as of 1 March 2017 and as of 1 October 2021

<p><strong>Abstract</strong></p> <p>We developed a methodology to detect the oldest scholarly reference added to Wikipedia articles by which a certain paper is uniquely identifiable as the &quot;first appearance of the scholarly reference.&quot; We identified the first appearances of 923,894 scholarly references (611,119 unique DOIs) in180,795 unique pages on English Wikipedia as of March 1, 2017, and stored them in the dataset. Moreover, we assessed the precision of the dataset, which was and it was a high precision regardless of the research field. In this version,&nbsp;it is available not only the dataset of English Wikipedia as of March 1, 2017, but also English Wikipedia as of October 1, 2021, generated by using&nbsp;the same methodology.</p> <p>&nbsp;</p> <p><strong>Data Records</strong></p> <p>The data format of the dataset is JSON lines, where each line is a single record. In this dataset, we detected the first appearance of each scholarly reference added to Wikipedia articles. If there are multiple references corresponding to the same paper on the same page, only the oldest one is collected. Sample of the record is the following.</p> <ul> <li>doi -- DOI corresponding to the paper (String), e.g., &quot;10.1006/anbe.1996.0497&quot;</li> <li>paper_type -- Document type of the paper (String), e.g., &quot;journal-article&quot;</li> <li>paper_container_title -- Journal title, book title, or proceedings title (Array of String), e.g., [&quot;Animal Behaviour&quot;]</li> <li>paper_publisher -- Publisher name (String), e.g., &quot;Elsevier BV&quot;</li> <li>paper_title -- Paper title (Array of String), e.g., [&quot;Push or pull: an experimental study on imitation in marmosets&quot;]</li> <li>paper_published_year -- Published year (String), e.g., &quot;1997&quot;</li> <li>paper_issue -- Issue number (String), e.g., &quot;4&quot;</li> <li>paper_volume -- Volume number (String), e.g., &quot;54&quot;</li> <li>paper_page -- Page numbers (String), e.g., &quot;817-831&quot;</li> <li>paper_author -- Authors information consisted of given and family names, sequences (order in author names), and affiliations (Array of JSON), e.g., [{&quot;given&quot;:&quot;THOMAS&quot;, &quot;family&quot;:&quot;BUGNYAR&quot;, &quot;sequence&quot;:&quot;first&quot;, &quot;affiliation&quot;:[]}, {&quot;given&quot;:&quot;LUDWIG&quot;, &quot;family&quot;:&quot;HUBER&quot;, &quot;sequence&quot;:&quot;additional&quot;, &quot;affiliation&quot;:[]}]</li> <li>issn -- ISSN related to the paper (Array of String), e.g., [&quot;0003-3472&quot;]</li> <li>research_field -- Research fields from ESI categories (Array of String), e.g., [&quot;PLANT &amp; ANIMAL SCIENCE&quot;]</li> <li>page_id -- Page id (String), e.g., &quot;577858&quot;</li> <li>page_title -- Page title (String), e.g., &quot;Imitation&quot;</li> <li>revision_id -- Revision id (String), e.g., &quot;203309031&quot;</li> <li>revision_timestamp -- Revision timestamp (String), e.g., &quot;2008-04-04 15:54:09 UTC&quot;</li> <li>revision_comment -- Revision comment (edit summary) (String), e.g., &quot;/* Animal Behaviour */&quot;</li> <li>editor_name -- Wikipedia editor&#39;s name (String), e.g., &quot;Nicemr&quot;</li> <li>editor_type -- Type of the editor (String), e.g., &quot;User&quot;</li> </ul> <p><strong>References</strong></p> <ul> <li>Kikkawa, J., Takaku, M. &amp; Yoshikane, F. &quot;Dataset of first appearances of the scholarly bibliographic references on Wikipedia articles&quot;, Scientific Data, Vol. 9, Article number 85, pp. 1-11, 2022. <a href="https://doi.org/10.1038/s41597-022-01190-z">https://doi.org/10.1038/s41597-022-01190-z</a>.</li> </ul> <p><strong>FUNDING</strong></p> <ul> <li>JSPS KAKENHI Grant Number <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-20K12543">JP20K12543</a></li> <li>JSPS KAKENHI Grant Number <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-21K21303">JP21K21303</a></li> </ul>

opencc-by-sa-4.0Oct 2021View details →
zenodo40/100

Wikipedia video games similarity dataset with expert annotations

<p>A video games NLP dataset extracted from Wikipedia.</p> <p>For all articles, the figures and tables have been filtered out, as well as the categories and &quot;see also&quot; sections.</p> <p>The article structure, and&nbsp;particularly the sub-titles and paragraphs are kept in these picese.</p> <p>Provided as well are 90 seeds with recommended articles, annotated by human experts.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Wikipedia Knowledge Graph dataset

<p>Wikipedia is the largest and most read online free encyclopedia currently existing. As such, Wikipedia offers a large amount of data on all its own contents and interactions around them, &nbsp;as well as different types of open data sources. This makes Wikipedia a unique data source that can be analyzed with quantitative data science techniques. However, the enormous amount of data makes it difficult to have an overview, and sometimes many of the analytical possibilities that Wikipedia offers remain unknown. In order to reduce the complexity of identifying and collecting data on Wikipedia and expanding its analytical potential, after collecting different data from various sources and processing them, we have generated a dedicated Wikipedia Knowledge Graph aimed at facilitating the analysis, contextualization of the activity and relations of Wikipedia pages, in this case limited to its English edition. We share this Knowledge Graph dataset in an open way, aiming to be useful for a wide range of researchers, such as informetricians, sociologists or data scientists.</p> <p>There are a total of 9 files, all of them in tsv format, and they have been built under a relational structure. The main one that acts as the core of the dataset is the <em><strong>page</strong></em> file, after it there are 4 files with different entities related to the Wikipedia pages (<em><strong>category</strong></em>, <em><strong>url</strong></em>, <em><strong>pub</strong></em> and <em><strong>page_property</strong></em> files) and 4 other files that act as &quot;intermediate tables&quot; making it possible to connect the pages both with the latter and between pages (<em><strong>page_category</strong></em>, <em><strong>page_url</strong></em>, <em><strong>page_pub</strong></em> and <em><strong>page_link</strong></em> files).</p> <p>The document <em><strong>Dataset_summary</strong></em>&nbsp;includes a detailed description of the dataset.</p> <p>Thanks to Nees Jan van Eck and the Centre for Science and Technology Studies (CWTS) for the valuable comments and suggestions.</p>

opencc-zeroMar 2022View details →
zenodo40/100

Wikipedia rendered as synthetic handwriting

<p>This is the synthetic handwriting data used to pre-train <strong>Dessurt</strong> (<a href="https://arxiv.org/abs/2203.16618">https://arxiv.org/abs/2203.16618</a>).</p> <p>It is text sampled from Wikipedia and generated with the method described in &quot;Text and Style Conditioned GAN for Generation of Offline Handwriting Lines&quot; (<a href="https://arxiv.org/abs/2009.00678">https://arxiv.org/abs/2009.00678</a>). More data can be quite easily obtained using this code: <a href="https://github.com/herobd/handwriting_line_generation">https://github.com/herobd/handwriting_line_generation</a></p> <p>Inside the tar is a single directory with ~800k generated handwriting line images (&quot;sample_0.png&quot;, &quot;sample_123.png&quot;,&nbsp;&quot;sample_3292524.png&quot;, etc.), and &quot;OUT.txt&quot; which has the GT for each line image.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Polifonia_Corpus_Wikipedia_Annotations_DE

<p>Polifonia_Corpus_Wikipedia_Annotations_DE</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record