Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “Mandarin Chinese”
Mandarin Chinese IDS wordlist by Hsiao-jung Yu and Yifan Wang
<p>Cite the source of the dataset as:</p> <blockquote> <p>Hsiao-jung Yu and Yifan Wang. 2021. Mandarin Chinese dictionary. In: Key, Mary Ritchie & Comrie, Bernard (eds.) The Intercontinental Dictionary Series. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at https://ids.clld.org/)</p> </blockquote>
Wbbyyr: FastText language models for Mandarin Chinese, trained on 14m Sina Weibo posts for each year in 2012-2018 (Fold 1 of 10)
<p>Wbbyyr: FastText language models for Mandarin Chinese, trained on 14,440,000 Sina Weibo posts for each year in 2012-2018.</p> <p>The 14,440,000 posts from each year are split into 10 folds. Due to Zenodo size limit, this dataset contains only the first fold from each year.</p> <p>Each model is trained for 20 iterations. Each vector is 300 dimensions long.</p>
Coded data and R scripts for the article-Toward a dynamic behavioral profile of the Mandarin Chinese temperature term re
<p>These are the coded dataset and R scripts for the article "Towards a dynamic behavioral profile of Mandarin Chinese temperature term re: A diachronic semasiological approach".</p>
Divide and Remaster v3: Mandarin Chinese Validation & Test Set
<p>Divide and Remaster v3 is a multilingual rework of the Divide and Remaster v2 dataset by Pétermann et al.</p> <p>This repository contains the <strong>metadata, validation set audio, and test set audio of the</strong> <strong>Mandarin Chinese variant </strong>of DnR v3.</p> <p>The major changes from DnR v2 are as follows:</p> <ul> <li>the dialogue stem now contains content from more than 30 languages across various language families;</li> <li>speech, vocals, and/or vocalizations have been removed from the music and effects stems;</li> <li>loudness and timing parametrization have been adjusted to approximate the distributions of real cinematic content;</li> <li>the mastering process now preserves relative loudness between stems and approximates standard industry practices.</li> </ul> <p>See the linked GitHub repository for more details.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.