Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “Belarusian”

Learn how ShareScore rates datasets ↗
zenodo44/100

Wikipedia: wikipedia-be (Belarusian)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://be.wikipedia.org/

opencc-by-sa-4.0Aug 2024View details →
zenodo36/100

Ukrainian 14-syllable verse in Belarusian poetry: the rhythm of translations and imitations (dataset)

<p>Data and source code accompanying the talk:<br> У. В. Парыцкі. Украінскі 14-складовы верш у беларускай паэзіі: рытміка перакладаў і імітацый // X Міжнародны Кангрэс даследчыкаў Беларусі, Коўна, 01.10.2022 [Vladislav Poritski. Ukrainian 14-syllable verse in Belarusian poetry: the rhythm of translations and imitations // Presented at 10th International Congress of Belarusian Studies, Kaunas, 01.10.2022]</p> <p>The empirical investigation of 14-syllable verse, presented in the talk, is based upon a sample of Ukrainian texts by Taras Shevchenko, their Belarusian translations, and original Belarusian poetry by Yanka Kupala, Yakub Kolas, Piatruś Brouka. The dataset structure is as follows:</p> <ul> <li>./0_plain &ndash; plain texts;</li> <li>./1_accentuated &ndash; accentuated texts;</li> <li>metadata_shevchenko.tsv, metadata_be_authors.tsv &ndash; metadata files describing the texts;</li> <li>make_reports.py &ndash; a Python script to generate statistic reports from the accentuated texts;</li> <li>./2_reports &ndash; programmatically generated reports;</li> <li>slides.tex &ndash; LaTeX source code of the talk&#39;s slides, where the reports are embedded as diagrams and tables;</li> <li>slides.pdf &ndash; PDF version of the slides.</li> </ul> <p>The directories ./0_plain, ./1_accentuated, ./2_reports are provided in .zip archives.</p> <p>Belarusian translations of Taras Shevchenko&#39;s poetry have been taken from the book:<br> Т. Р. Шаўчэнка. Вершы. Паэмы. Мінск: Мастацкая літаратура, 1989.<br> (scan copy available at https://files.knihi.com/Knihi/scanned/Saucenka.Viersy_paemy.djvu)<br> Each poem is stored in a separate .txt file. The file name indicates the number of the poem&#39;s first page in the scanned book, e.g.: 021.txt. Same names are used for the respective Ukrainian texts. In each pair of files, such as e.g. ./0_plain/uk/021.txt and ./0_plain/be/021.txt, the texts are aligned line by line. Poem titles in both languages, translator names, and the URLs of Ukrainian source texts are provided in metadata_shevchenko.tsv.</p> <p>Original Belarusian poetry, kept in ./0_plain/be, doesn&#39;t require any alignment, and the naming scheme is different. Poem titles, author names, and the URLs of Belarusian source texts are provided in metadata_be_authors.tsv.</p> <p>In all Ukrainian and Belarusian texts, metrically irrelevant lines are discarded, only 14-syllable verse lines are stored, each of them split graphically into 8+6 syllables. Occasional minor violations, i.e. &plusmn; one or two syllables, are allowed in the texts but ignored in the statistic reports. No spans shorter than a pair of rhyming 14-syllable lines (or, graphically, a quatraine of 8+6+8+6 syllables) were sampled from polymetric poems.</p> <p>These special characters are used:</p> <ul> <li>&quot;/&quot; to represent line break in the source edition;</li> <li>&quot;//&quot; for section break (next stanza, another character&#39;s words);</li> <li>trailing &quot;#&quot; for the inverse of line break: to recover the original 14-syllable line as printed in the source edition, one should remove the newline;</li> <li>leading &quot;#&quot; for mis-aligned lines, e.g. those missing in the Belarusian translation and added hypothetically, in order to restore the alignment.</li> </ul> <p>The procedure of accentuating Ukrainian and Belarusian texts was semi-automatic, using an opportunistic database of word accents crawled from online lexicographic resources: https://slounik.org for Belarusian, https://uk.wiktionary.org and https://slovnyk.ua/nagolos.php for Ukrainian. The database and the accentuator script are not part of this dataset. Although a fair bit of manual supervision was put into ensuring that most accents are accurate, it&#39;s likely that some errors still remain, especially in the Ukrainian data, so please be cautious.</p> <p>Accentuated texts in ./1_accentuated/uk and ./1_accentuated/be are lowercased, with all punctuation stripped off. As usual in quantitative study of East Slavic verse (see e.g. https://doi.org/10.12697/smp.2019.6.2.02 for a recent overview), we distinguish between two kinds of stresses: pronouns and certain other function words bear &quot;light&quot; stress, while content words bear &quot;heavy&quot; stress. These are the designations:</p> <ul> <li>&quot;`&quot; for light stress, to the left of the stressed vowel;</li> <li>&quot;&#39;&quot; for heavy stress, to the right of the stressed vowel (note that after consonants, &quot;&#39;&quot; is an apostrophe);</li> <li>&quot;*&quot; for variant heavy stress, as in Ukrainian <em>ба*йду*же</em>;</li> <li>&quot;_&quot; to group clitics together with stressed words, as in Ukrainian <em>і_не_привіта&#39;ла</em>.</li> </ul> <p>In rare exceptional cases, the meter may require to pronounce syllabic consonants, as in Belarusian <em>рэестр</em>. To match pronunciation, we add a vowel in square brackets: <em>рэест[а]р</em>.</p> <p>The reports summarize certain statistic properties of the dataset:</p> <ul> <li>translators.csv &ndash; a breakdown of Shevchenko&#39;s Belarusian translations into the numbers of lines contributed by each translator. 8+6 are counted as separate lines. Syllable count violations are ignored: a pair of aligned Ukrainian / Belarusian lines is not counted towards the translator&#39;s total, if the number of syllables is irrelevant (e.g. 9 and 9) or mismatched (e.g. 8 and 6).</li> <li>be_authors.csv &ndash; line counts by author in the original Belarusian poetry. Same counting rules apply, modulo the alignment.</li> <li>rhythm.csv &ndash; percentages of accents on each of the 14 syllables in various samples, grouped by author and / or translator. Rows are syllables, columns are samples. Accents in each sample are counted two ways: &quot;min&quot; &ndash; only heavy stresses, &quot;max&quot; &ndash; all stresses.</li> <li>total_accentuation.csv &ndash; average accent counts per line in Shevchenko&#39;s Ukrainian texts and Belarusian translations, separately for 8+6, separately for heavy and all stresses.</li> <li>word_boundary.csv &ndash; statistics of word boundary positions in 8-syllable 2-word heavy-stressed lines in Shevchenko&#39;s Ukrainian texts and Belarusian translations.</li> <li>trochaicity.csv &ndash; ratio of stresses that match trochaic metrical template, separately for 8+6, heavy stresses only. Rows are samples: Shevchenko&#39;s Ukrainian texts and Belarusian translations, original poetry by three Belarusian authors.</li> </ul> <p>For implementation details, see the source code of make_reports.py.</p> <p>To reproduce report generation, you will need Python. Unzip the archive 1_accentuated.zip and run:<br> python3 make_reports.py</p> <p>To rebuild the slides, you will need LaTeX:<br> xelatex -synctex=1 -interaction=nonstopmode -shell-escape slides.tex<br> If the bibliographic references are not rendered properly, rerun once again.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Annotated files of the Soviet Belarusian newspaper Zviazda (years 1938, 1939, 1940)

<p>The present data contain the texts of the Soviet Belarusian newspaper &#39;Zviazda&#39; for the years 1938, 1940 and 1940 as a ZIP archive.</p> <p>The data are in the <em>CoNLL-U</em> format, therefore morphologically and syntactically annotated.</p> <p>The data were generated from PDF files through OCR (by Tesseract) and anotated with UDPipe. These data were generated while exploring a pipeline to process Belarusian texts and they are very noisy.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Ontolex-lemon and TIAD versions of Apertium Belarusian-Russian dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> <ul> <li>2016-2017, Francis M. Tyers </li> <li>2017, John Lyell </li> <li>2016, Vladislav Kiryukhin </li> </ul>

opengpl-2.0-or-laterMar 2020View details →
zenodo32/100

Fig. 1 in New and rare for the Belarusian fauna True Bug species (Insecta: Hemiptera: Heteroptera) from the parks of Brest Region

Fig. 1. The localities of new and rare the true bugs species (Heteroptera) registered in Brest Region: 1 — fragment of the park in the former Kościuszko's Estate (Malye Sekhnovichi Village, Zhabinka District); 2 — park of the former Moloczewski's Estate (Peski Village, Kobrin District); 3 — Suvorov Park (Kobrin Town); 4 — park of the former Wyslouch's Estate (Perkovichi Village, Drogichin District); 5 — park of the former Puslowski's Estate (Peski Village, Beryoza District. Рис. 1. Места регистрации новых и редких видов настоЯЩих полужесткокрылых (Heteroptera) на территории Брестской области: 1 — фрагмент парка в имении КостюШек (д. Малые Сехновичи, Жабинковский район); 2 — парк панов Молочевских (д. Пески, Кобринский район); 3 — парк им. Суворова (г. Кобрин); 4 — парк панов Вислоухов (д. Перковичи, Дрогиченский район); 5 — парк графов Пусловских (д. Пески, БереЗовский район).

opennotspecifiedMar 2021View details →
zenodo28/100

Weak lexical equivalents in Russian → Belarusian translation: approaches to identification (dataset)

<p>Data accompanying the talk:<br> А. А. Ваўчок, У. В. Парыцкі. <a href="https://www.academia.edu/109056239">Слабыя лексічныя адпаведнікі ў руска-беларускім перакладзе: падыходы да ідэнтыфікацыі</a> // XI Міжнародны Кангрэс даследчыкаў Беларусі, Гданьск, 23.09.2023 [Oksana Volchek, Vladislav Poritski. Weak lexical equivalents in Russian &rarr; Belarusian translation: approaches to identification // Presented at 11th International Congress of Belarusian Studies, Gdańsk, 23.09.2023]</p> <p>The file <code>data.csv</code> is a test battery of 100 contexts in Russian, expected to cause difficulties when attempting to translate them into Belarusian. This is because each context contains a pair of words distinct in Russian but lacking clearly distinct Belarusian lexical equivalents. The contexts were taken, often with some simplifications, from various sources available on the web, i.e. they are representative of real usage, not artificially constructed.</p> <p>The file has four columns:</p> <ul> <li><code>w1</code>, <code>w2</code> &ndash; two words belonging to the same part of speech (typically adjective, noun, or verb), <code>w1</code> lexicographically preceding <code>w2</code>;</li> <li><code>context</code> &ndash; a snippet of text with both words (typically a single sentence, sometimes longer);</li> <li><code>inspired_by</code> &ndash; URL of a web page or file containing the original version of the context, which may have been slightly abridged or reworded for illustrative purposes. Archive copies of most URLs, with the exception of Google Books links, are available in the Wayback Machine (https://web.archive.org) or in https://archive.today.</li> </ul>

opencc-by-4.0Sep 2023View details →
geo24/100

East Eurasian ancestry in the middle of Europe: genetic footprints of Steppe nomads in the genomes of Belarusian Lipka Tatars

GEO Series GSE82309. Homo sapiens. 6 samples. Type: SNP genotyping by SNP array; Genome variation profiling by SNP array.

openGEO-OpenAug 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record