Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,502

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,502 results for “Russian”

Learn how ShareScore rates datasets ↗
edi56/100

Anatoxin concentrations, algal assemblages, and water quality data for the South Fork Eel, Salmon, and Russian Rivers in northern California, 2022-2023

We collected this data to better understand the timing of peak benthic cyanobacterial mat occurrence (specifically taxa associated with anatoxin production, Microcoleus and Anabaena) and mat anatoxin concentrations in rivers. We sampled in northern California on the South Fork Eel, Salmon, and Russian Rivers biweekly in 2022, and the Salmon River biweekly and South Fork Eel weekly in 2023. During each sampling event, we conducted benthic cover surveys, measured in-situ water quality parameters (temperature, pH, dissolved oxygen, conductivity), and collected surface water samples and targeted cyanobacteria samples. In 2022 on all rivers and in 2023 at the Salmon River, we also collected distributed non-targeted periphyton samples to characterize full-reach community compositions. All sampling was completed in 150-m reaches upstream of sensors recording continuous dissolved oxygen, conductivity, and temperature data. We analyzed surface water samples for nitrate, ammonium, soluble reactive phosphate, total dissolved carbon, and dissolved organic carbon. We also analyzed surface water samples from 2022 for major anions (Cl, SO4, Br) and cations (Na, K, Mg, Ca). Targeted-cyanobacteria and non-target periphyton samples were analyzed for anatoxins (and two other classes of toxins, microcystins and cylindrospermopsins), relative abundance of algal taxa (via microscopy), ash-free dry mass, and chlorophyll-a. To estimate mean river depth within the dissolved oxygen footprint upstream of sensors, we kayaked portions of the river and collected river depth measurements. We also measured discharge at each river excluding the Salmon River (due to high discharge) and completed pebble counts at the South Fork Eel River to obtain sediment grain size distributions. Lastly, we estimated daily reach-scale river metabolism (gross primary productivity and ecosystem respiration) using data from dissolved oxygen sensors with the "streamMetabolizer" package in R at all sensor placements.

openCC (other)Nov 2025View details →
zenodo52/100

Database of Annotated Core Arguments: English, Lao and Russian

<p>This database contains corpus examples of transitive clauses with annotated core arguments realized as syntactic subjects and objects (A and P) in English, Lao and Russian. The coding scheme&nbsp;was developed together with Alena Witzlack-Makarevich</p>

opencc-by-4.0Oct 2020View details →
zenodo48/100

Russian University Journals sample

<p>The database includes information on Russian journals published by Federal universities, National Research universities and Basic universities and covers such aspects as indexing information in different bibliographic databases, volumes of indexed papers in 2018-2022, rankings in different databases, subject categories, geographical locations, numbers of issues per year, list of founders, journals sites, etc. (in total 59 variables).</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Local Governance in Ukraine during the full-scale Russian invasion. – Merged data from online surveys of local self-government authorities by the Congress of Local and Regional Authorities of the Council of Europe in 2022 and Kyiv School of Economics in 2024.

The dataset includes responses from two waves of online surveys targeting local self-government representatives in Ukraine, with a focus on crisis governance during the ongoing Russian war. The first wave was conducted from August 30 to September 20, 2022, by the Congress of Local and Regional Authorities of the Council of Europe, yielding 241 responses (16% of all Ukrainian local communities). The second wave was conducted by Kyiv School of Economics from January 1 to March 12, 2024, with 181 responses (14% of government-controlled municipalities). Data formats include CSV and SAV files, along with an XSL codebook for both waves. The merged dataset comprises 442 responses from small, medium, and large municipalities under varied security conditions, with a total file size of approximately 4 MB.

openodc-byNov 2024View details →
zenodo44/100

Photonics4All Bookmark LED (Russian)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How can Light Emitting Diodes (LEDs) transform local food production?<br> <br> Because LEDs emit pure and specific colours they can be used to make plants grow faster and larger.  LEDs can replace sunlight or costly greenhouse lamps to grow crops in cold climates or during off-season periods. Growing food locally reduces the need for long-distance transport and lessens the environmental impact used to produce the food. All thanks to Photonics!</p> <p> </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Photonics4All Bookmark Needle (Russian)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How can light replace a needle?</p> <p>We no longer need to use a needle to monitor the level of oxygen in your blood!   We can use light emitting diodes (LEDs) attached to the top of your finger - and a light detector underneath to measure the amount of light passing through your finger.  As Hemoglobin - the proteins in red blood cells which carry oxygen - absorbs light we can determine whether you have enough oxygen in your blood. More advanced devices can also monitor your heart-rate and blood pressure.  We'll even be using light to measure your blood-sugar level in the future. All thanks to Photonics!</p> <p> </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Photonics4All Bookmark Crime (Russian)

<p>The purpose of the bookmarks for the project Photonics4all is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How can light help solve crimes?<br> <br> Light sources used to illuminate crime scenes help investigators solve crimes. Light is used by forensic detectives to help locate evidence such as latent fingerprints, bodily fluids, hair and fibres, bruises, wound patterns, shoe and foot imprints, gunshot residues or drug traces.  All of these clues fluoresce - or glow brightly - under selectively coloured light.  Photography is also vital to help collect and store this evidence. </p> <p>All thanks to Photonics! </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Photonics4All Bookmark Bubble (Russian)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> Why do soap bubbles have colour?<br> <br> Light reflects off both the inner and outer surfaces of a soap bubble. As the bubble dries out it changes thickness and the light waves reflecting off both surfaces have to travel different distances. White light is made up of all different colours – or waves of different lengths and - when light waves meet – or overlap - they create different colours.  Because reflected light travels different distances due to the different film thicknesses we see iridescence in soap bubbles. This phenomenon is used in photonics to provide anti-reflection coating on your glasses for example.</p> <p> </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Photonics4All Bookmark Chip (Russian)

<p>The purpose of the bookmarks for the project Photonics4All is to increase the public awareness of photonics and especially of the technological advances of photonics which have changed and improved everyday life (basic technology introduction).<br> <br> How light makes computers and phones smaller and faster?</p> <p>Did you know that we use light to fabricate the electronic chips in computers and mobile phones? Recent developments in photolithography where light is used to control where conductive metal is placed on the chips - have enabled us to put more transistors than there are people on earth!  Transistors are responsible for controlling the path of electricity/information through a chip.  These technological developments have led to improving the speed, size and energy consumption of our chips, making them smaller and more efficient.</p> <p> </p>

opencc-by-4.0Jul 2015View details →
zenodo44/100

Toward a Comparable Corpus of Latvian, Russian and English Tweets

<p>Twitter has become a rich source for linguistic data. Here, a possibility of building a trilingual Latvian-Russian-English corpus of tweets from Riga, Latvia is investigated. Such a corpus, once constructed, might be of great use for multiple purposes such as training machine translation models, examining cross-lingual phenomena and studying the population of Riga. This pilot study shows that it is feasible to build such a resource by building and analysing a pilot corpus, which is made publicly available and can be used to construct a large comparable corpus.</p>

opencc-by-4.0May 2017View details →
zenodo44/100

Grammatical functions, inflectional class and textual frequency in Russian nominals

<p><br>The following datasets were created for the ESRC-funded project 'Paradigms in use' (https://www.smg.surrey.ac.uk/projects/paradigms/) (RES-000-23-0082):</p> <p>The datasets are in the form of 8 EXCEL spreadsheets. The lexemes recorded in the datasets are those represented by word forms occurring in total at least five times. Lexemes occurring less than five times were excluded to avoid large standard errors in the estimates which occur when observed numbers in each category are small (Corbett, Hippisley, Brown and Marriott 2001:208).</p>

opencc-by-nc-4.0Aug 2015View details →
zenodo44/100

RUSSE'2018: Human-Annotated Sense-Disambiguated Word Contexts for Russian

<p>This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the <a href="https://russe.nlpub.org/2018/wsi/">RUSSE&#39;2018</a> shared task on Word Sense Induction and Disambiguation for the Russian language; part of the&nbsp;<em>bts-rnc</em>&nbsp;evaluation dataset. These sense identifiers are disambiguated as according to the sense inventory of the <a href="http://gramota.ru/slovari/info/bts/">Large Explanatory Dictionary of Russian</a>.</p> <p>The annotation is done on December 1, 2017, on the&nbsp;<a href="https://tolokanyandex.com/">Yandex.Toloka</a>&nbsp;crowdsourcing platform. In particular, 80 pre-annotated contexts are used for&nbsp;training the human annotators, 2562 contexts are annotated by humans such that each&nbsp;context was annotated by 9 different annotators. The annotation reliability&nbsp;is indicated by a high value of Krippendorff&#39;s&nbsp;&alpha; = 0.83. After the annotation, every context was additionally inspected (&ldquo;curated&rdquo;) by the organizers of the shared task.</p> <p>The following words are represented:&nbsp;<em>акция</em> (action / stock), <em>байка</em> (yarn / tale), <em>гвоздика</em> (carnation / nail), <em>гипербола</em> (hyperbole), <em>град</em> (avalanche), <em>гусеница</em> (grub), <em>домино</em> (domino), <em>кабачок</em> (marrow / pub), <em>капот</em> (hood), <em>карьер</em> (mine / career), <em>кок</em> (cook), <em>крона</em> (top / crown), <em>круп</em> (croup), <em>мандарин</em> (mandarine), <em>рок</em> (fate / rock), <em>слог</em> (syllable), <em>стопка</em> (glass, stack), <em>таз</em> (bowl), <em>такса</em> (rate / badger-dog), <em>шах</em> (shah / check).</p> <p>The following files are included in this dataset:</p> <ul> <li>Toloka assignments (training:&nbsp;<em>tasks-train.tsv</em>, annotation: <em>tasks-test.tsv</em>)</li> <li>Toloka output (non-aggregated: <em>assignments_01-12-2017.tsv.xz</em>, aggregated:&nbsp;<em>aggregated_results_pool_1036853__2017_12_01.tsv</em>)</li> <li>annotator agreement report (<em>agreement.txt</em>)</li> <li>curated report (<em>report-curated.tsv.xz</em> and a supplementary file&nbsp;<em>tasks-eval.tsv.xz</em>)</li> <li>the final aggregated dataset (<em>bts-rnc-crowd.tsv</em>)</li> </ul> <p>The <em>bts-rnc-crowd.tsv</em>&nbsp;file has the following format: <em>id</em>, <em>lemma</em>, <em>sense_id</em>, <em>left</em> hand side context, <em>word</em> form, <em>right</em> hand side context, list of&nbsp;<em>senses</em>. The encoding is UTF-8 and the line breaks are LF (UNIX).</p>

opencc-by-sa-4.0Jan 2018View details →
zenodo44/100

Defective Drugs in Russian Federation

<p>Using this dataset it&#39;s possible to estimate control measures over&nbsp;quality control of medicines in Russian Federation. Datasets provide information about&nbsp;events of control and destruction of defective drugs.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Electronic representation of Russian journals on Earth Sciences

<p>The dataset describes the deprth of digital archives of the most authoritative Russian journal on Earth Sciences. The five figures demostrate the results of graphical processing while preparing digital archives of Geologiya i Geofizika and Zapiski Gornogo Instituta journals, as well as the forms of metadata presentation in electronic archives.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

1117 Russian cities with city name, region, geographic coordinates and 2020 population estimate

<p>1117 Russian cities with city name, region, geographic coordinates and 2020 population estimate.</p> <p>&nbsp;</p> <p>How to use</p> <pre>from pathlib import Path import requests import pandas as pd url = (&quot;https://raw.githubusercontent.com/&quot; &quot;epogrebnyak/ru-cities/main/assets/towns.csv&quot;) # save file locally p = Path(&quot;towns.csv&quot;) if not p.exists(): content = requests.get(url).text p.write_text(content, encoding=&quot;utf-8&quot;) # read as dataframe df = pd.read_csv(&quot;towns.csv&quot;) print(df.sample(5))</pre> <p>&nbsp;</p> <p>Files:</p> <ul> <li><a href="https://github.com/epogrebnyak/ru-cities/blob/main/assets/towns.csv">towns.csv</a> - city information</li> <li><a href="https://github.com/epogrebnyak/ru-cities/blob/main/assets/regions.csv">regions.csv</a> - list of Russian Federation regions</li> <li><a href="https://github.com/epogrebnyak/ru-cities/blob/main/assets/alt_city_names.json">alt_city_names.json</a> - alternative city names</li> </ul> <p>&nbsp;</p> <p>Сolumns (towns.csv):</p> <p>Basic info:</p> <ul> <li><code>city</code> - city name (several cities have alternative names marked in <code>alt_city_names.json</code>)</li> <li><code>population</code> - city population, thousand people, Rosstat estimate as of 1.1.2020</li> <li><code>lat,lon</code> - city geographic coordinates</li> </ul> <p>Region:</p> <ul> <li><code>region_name</code> - subnational region (oblast, republic, krai or AO)</li> <li><code>region_iso_code</code> - <a href="https://en.wikipedia.org/wiki/ISO_3166-2:RU">ISO 3166 code</a>, eg <code>RU-VLD</code></li> <li><code>federal_district</code>, eg <code>Центральный</code></li> </ul> <p>City codes:</p> <ul> <li><code>okato</code></li> <li><code>oktmo</code></li> <li><code>fias_id</code></li> <li><code>kladr_id</code></li> </ul> <p>&nbsp;</p> <p>Data sources</p> <ul> <li>City list and city population collected from Rosstat publication <a href="https://rosstat.gov.ru/folder/210/document/13206">Регионы России. Основные социально-экономические показатели городов</a> and parsed from publication Microsoft Word files.</li> <li>City list corresponds to <a href="https://ru.wikipedia.org/wiki/%D0%A1%D0%BF%D0%B8%D1%81%D0%BE%D0%BA_%D0%B3%D0%BE%D1%80%D0%BE%D0%B4%D0%BE%D0%B2_%D0%A0%D0%BE%D1%81%D1%81%D0%B8%D0%B8">this Wikipedia article</a>.</li> <li>Alternative dataset is <a href="https://github.com/hflabs/city">wiki-based Dadata city dataset</a> (no population data).</li> </ul> <p>&nbsp;</p> <p>Comments</p> <p>&nbsp;</p> <p>City groups</p> <ul> <li> <p><code>Ханты-Мансийский</code> and <code>Ямало-Ненецкий</code> autonomous regions excluded to avoid duplication as parts of <code>Тюменская область</code>.</p> </li> <li> <p>Several notable towns are classified as administrative part of larger cities (<code>Сестрорецк</code> is a municpality at Saint-Petersburg, <code>Щербинка</code> part of Moscow). They are not and not reported in this dataset.</p> </li> </ul> <p>&nbsp;</p> <p>By individual city</p> <ul> <li><code>Белоозерский</code> not found in Rosstat publication, but <a href="https://github.com/epogrebnyak/ru-cities/issues/5#issuecomment-886179980">should be considered a city as of 1.1.2020</a></li> </ul> <p>&nbsp;</p> <p>Alternative city names</p> <ul> <li> <p>We suppressed letter &quot;ё&quot; <code>city</code> columns in towns.csv - we have <code>Орел</code>, but not <code>Орёл</code>. This affected:</p> <ul> <li><code>Белоозёрский</code></li> <li><code>Королёв</code></li> <li><code>Ликино-Дулёво</code></li> <li><code>Озёры</code></li> <li><code>Щёлково</code></li> <li><code>Орёл</code></li> </ul> </li> <li> <p><code>Дмитриев</code> and <code>Дмитриев-Льговский</code> are the same city.</p> </li> </ul> <p><code>assets/alt_city_names.json</code> contains these names.</p> <p>&nbsp;</p> <p>Tests</p> <pre><code>poetry install poetry run python -m pytest </code></pre> <p>&nbsp;</p> <p>How to replicate dataset</p> <p>&nbsp;</p> <p>1. Base dataset</p> <p>Run:</p> <ul> <li>download data stro rar/get.sh</li> <li>convert <code>Саратовская область.doc</code> to docx</li> <li>run make.py</li> </ul> <p>Creates:</p> <ul> <li><code>_towns.csv</code></li> <li><code>assets/regions.csv</code></li> </ul> <p>&nbsp;</p> <p>2. API calls</p> <p>Note: do not attempt if you do not have to - this runs a while and loads third-party API access.</p> <p>You have the resulting files in repo, so probably does not need to these scripts.</p> <p>Run:</p> <ul> <li><code>cd geocoding</code></li> <li>run coord_dadata.py (needs token)</li> <li>run coord_osm.py</li> </ul> <p>Creates:</p> <ul> <li>coord_dadata.csv</li> <li>coord_osm.csv</li> </ul> <p>&nbsp;</p> <p>3. Merge data</p> <p>Run:</p> <ul> <li>run merge.py</li> </ul> <p>Creates:</p> <ul> <li>assets/towns.csv</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Kremlin.ru transcripts 1999–2019, Russian

<p>Document collection scraped from the Russian governmental website kremlin.ru. Includes all items listed on&nbsp;http://kremlin.ru/events/president/transcripts and following pages (e.g. http://kremlin.ru/events/president/transcripts/2) from 31 December 1999 until the end of 2019.</p> <p>10,723 documents. One document in each row.</p> <p>Columns:</p> <p>- Id: format Kremlin-1&nbsp;&nbsp;<br> - Id_no: format 1&nbsp;&nbsp;<br> - Date: Document date, format 1999-12-31&nbsp;&nbsp;<br> - Title: Document title<br> - Text: Document text including title&nbsp;&nbsp;<br> - URL: URL from which the document is downloaded&nbsp;&nbsp;<br> - Downloaded: Date of download, format 2019-12-31</p> <p>Formats: rds and json.</p> <p>Version 1.1: edited column names.</p> <p>Kremlin.ru content is licensed&nbsp;under Creative Commons Attribution 4.0 International.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Dataset of inappropriate utterances on sensitive topics in Russian

<p>This dataset is dedicated to inappropriate messages -- the messages on a sensitive topic that can frustrate the reader and/or harm the reputation of the speaker. The concept of inappropriateness is rather close to toxicity, however, the clear toxicity itself, as well as explicit obscenity, has been intentionally dropped from this dataset.</p> <p>Not all messages related to sensitive topics are inappropriate. For example, speaking about racism you may either attack or protect someone. The main aim of this dataset is to detect appropriate and inappropriate utterances within known sensitive topics.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

A dataset with news messages from a Russian and a Ukrainian TV news channels

<p>This data was used in the following publications:</p> <p>1. Koltsova, O., &amp; Pashakhin, S. (2019). Agenda divergence in a developing conflict: Quantitative evidence from Ukrainian and Russian TV newsfeeds. Media, War &amp; Conflict, 1750635 21982987. <a href="https://doi.org/10.1177/1750635219829876">https://doi.org/10.1177/1750635219829876</a></p> <p>2.Pashakhin S. Topic Modeling for Frame Analysis of News Media // Proceedings of the AINL FRUCT 2016. С. 103-105 &ndash; URL: <a href="http://fruct.org/publications/abstract-AINL-FRUCT-2016/files/Pas.pdf">http://fruct.org/publications/abstract-AINL-FRUCT-2016/files/Pas.pdf</a></p> <p>The dataset contains 45,009 news messages collected from official websites of a Russian (Channel One) and a Ukrainian (Channel 5) TV channels. Ukrainian news items were translated into Russian.</p> <p>The dataset has six variables:</p> <ul> <li>text -- a news item;</li> <li>channel -- a source of an item (&#39;first&#39; -- Russian TV channel, &#39;five&#39; -- Ukrainian TV channel);</li> <li>date -- the date of publishing;</li> <li>url -- links to original news messages.</li> </ul>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Figs 45–52 in Immature stages and biology of the enigmatic oxyporine rove beetles, with new data on Oxyporus larvae from the Russian Far East (Coleoptera: Staphylinidae)

Figs 45–52. Third instar larva of Oxyporus (P.) melanocephalus Kirschenblatt, 1938, head morphology. 45 – head, dorsal view; 46 – head, ventral view; 47 – antenna, dorsal view; 48 – mandible, dorsal view; 49 – maxilla, dorsal view; 50 – labium, dorsal view; 51 – labium, lateral view; 52 – maxilla, ventral view.

opencc-by-4.0Mar 2020View details →
zenodo40/100

Figs 39–44 in Immature stages and biology of the enigmatic oxyporine rove beetles, with new data on Oxyporus larvae from the Russian Far East (Coleoptera: Staphylinidae)

Figs 39–44. Scanning electron micrographs of larva of Oxyporus procerus Kraatz, 1879. 39 – campaniform sensilla and setae of nasale; 40 – posterior epicranial group of sensilla; 41 – antennomeres II and III, apical sensorial complex; 42 – premental group of sensilla; 43 – campaniform sensillum, segment II of maxillary palpus; 44 – thoracic tergite I, lateral view.

opencc-by-4.0Mar 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record