Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,359
datasets available to search
ShareScore release 0.7.1
Dataset results
7,359 results for “Chinese”
Khotanese Manuscripts from Chinese Turkestan in the British Library (XML records)
<p>The file contains XML records matching the print edition of Skjaervo's catalogue, in TEI schema P4.</p> <p>The records in this file are a <strong>draft version</strong>. They have not yet been proofed and checked against physical holdings, which will be done with the next version release.</p> <p>The XML records have been produced as part of the work for the project <em>Beyond Boundaries: Religion, Region, Language and the State</em> (An ERC Synergy project from the European Research Council under the EU's 7th Framework Programme (FP7/2007-2013)/ERC grant agreement no.609823)</p>
COSN paper data (The Chinese Open Science Network (COSN): Building an Open Science community from scratch)
<p>This is the dataset for generating figure1 and figure 3 in the manuscript <em>The Chinese Open Science Network (COSN): Building an Open Science community from scratch </em>(Accepted by AMPPS). Preprint at: <a href="https://doi.org/10.31234/osf.io/ac9by">https://doi.org/10.31234/osf.io/ac9by</a>.</p> <p>All the data and codes are available in repo: <a href="https://github.com/OpenSci-CN/COSN_AMPPS_Paper">COSN_AMPPS_Paper</a> Accepted Version.</p>
Raw data to accompany the manuscript 'Data for Engineering Lipid Metabolism of Chinese Hamster Ovary (CHO) Cells for Enhanced Recombinant Protein Production' published in the Journal Data in Brief
<p>This repository consists of the raw western blot, microscopy and mass spectrometry data to accompany the manuscript 'Data for Engineering Lipid Metabolism of Chinese Hamster Ovary (CHO) Cells for Enhanced Recombinant Protein Production' published in the Journal Data in Brief and associated with the article '<a href="https://www.ncbi.nlm.nih.gov/pubmed/31805379">Engineering of Chinese hamster ovary cell lipid metabolism results in an expanded ER and enhanced recombinant biotherapeutic protein production</a>' published in the journal Metabolic Engineering (see DOI: 10.1016/j.ymben.2019.11.007). </p> <p>The western blot raw file is associated with Figure 1a and 1b of the Data in Brief manuscript.</p> <p>The confocal microscopy raw image files (x3) are associated with Figure 1c of the Data in Brief manuscript.</p> <p>The mass spectrometry files are the raw data that refers to the samples presented in Figure 5 of the Data in Brief manuscript. Files are labelled as in the Data in Brief and Metabolic Engineering manuscripts. The file name structures is as follows;</p> <p>CHO-Controlpoolai</p> <p>Where 'a' represents replicate 'a' of three biological replicates and 'i' refers to mass spectrometry technical analysis 1 of 3 technical analyses of each replicate (thus for each cell pool or line there are three biological replicates that are each analysed in triplicate such that there are 9 raw mass spectrometry files for each cell pool or line).</p> <p>All the mass spectrometry files are found in the compressed (zip) file named mass_spectrometry_raw_files_archive.zip</p>
A Finite State Tranducer that models Chinese Historical Phonology
<p>This is a finite state transducer that tries to model Chinese historical phonology from Old Chinese as reconstructed by Baxter and Sagart in <em>Old Chinese: a New Reconstruction</em> (Oxford, 2014) to Middle Chinese as presented in the system of Baxter in <em>A Handbook of Old Chinese Phonology </em>(Mouton, 1992).</p>
Members of the Chinese Students' Alliance in the United States (1912)
<p>The dataset is a list of members of the Chinese Students Alliance in the United States for the academic year 1911-1912. It is sourced from <em>The Directory of Chinese Students in the United States, 1911-1912</em>, compiled by the Chinese Students Alliance in 1912. The directory has been digitized by Google and is available in full view on <a href="https://babel.hathitrust.org/cgi/pt?id=nnc2.ark:/13960/t9574zp5q&seq=5">HathiTrust</a> and <a href="https://archive.org/details/ldpd_11381020_000">Internet Archive</a>.</p> <p>The attached table contains information on the students' names (in Chinese, English, and pinyin transliteration), gender, address in the United States, university in the United States (when available), and the alliance section (Eastern, Midwest, Western) to which they belonged (when available). Additionally, the table provides the geographical coordinates of the cities.</p> <p>The data was extracted using <a href="https://claude.ai/">Claude (AI</a>) and curated using Excel and R. The complete code for extracting and curating the data is available on <a href="https://github.com/carmand03/csa-directories">GitHub</a>. Additionally, the GitHub repository contains various statistics and visualizations, such as the distribution of students by city.</p> <p>The dataset contains 882 students (unique individuals), 802 men and 80 women. </p> <p>Distribution by sections (p.119): </p> <table> <tbody> <tr> <td><strong>Sections</strong></td> <td><strong>Members</strong></td> <td><strong>Non-Members</strong></td> <td><strong>Total</strong></td> </tr> <tr> <td>Eastern</td> <td>201</td> <td>127</td> <td>328</td> </tr> <tr> <td>Midwest</td> <td>121</td> <td>123</td> <td>244</td> </tr> <tr> <td>Western</td> <td>42</td> <td>66</td> <td>108</td> </tr> <tr> <td>Unreturned (Missing)</td> <td>101</td> <td>96</td> <td>197</td> </tr> <tr> <td>Total</td> <td>465</td> <td>412</td> <td>877</td> </tr> </tbody> </table>
Source Code Accompanying the Paper "More on network approaches in Historical Chinese Phonology (音韻學)"
<p>First version of the source code and data accompanying the paper "More on Network Approaches in Historical Chinese Phonology".</p> <p>This paper is available here:</p> <ul> <li>List, Johann-Mattis (2018): <strong>More on network approaches in Historical Chinese Phonology (音韻學)</strong>. Paper prepared for the <em>LFK Society Young Scholars Symposium</em>. Taibei: Li Fang-Kuei Society ofr Chinese Linguistics. URL: <a href="https://hal.archives-ouvertes.fr/hal-01706927">https://hal.archives-ouvertes.fr/hal-01706927</a>.</li> </ul> <pre><code>@InProceedings{List2018a, author = {List, Johann-Mattis}, title = {{More on Network Approaches in Historical Chinese Phonology (音韻學)}}, booktitle = {{LFK Society Young Scholars Symposium}}, year = {2018}, publisher = {Li Fang-Kuei Society for Chinese Linguistics}, pdf = {https://hal.archives-ouvertes.fr/hal-01706927/file/main.pdf}, url = {https://hal.archives-ouvertes.fr/hal-01706927}, address = {Taipei}, hal_id = {hal-01706927}, } </code></pre> <p>See the README.md for mor information.</p> <ul> <li> </li> </ul>
cldf-datasets/szetosinitic: Chinese Structure Dataset from Szeto et al.'s (2018) paper in CLDF-Format
<p>This is a structural dataset originally published along with a paper by Szeto et al. (2018) on Chinese dialect classification:</p> <blockquote> <p>Szeto, P. Y.; Ansaldo, U. & Matthews, S.Typological variation across Mandarin dialects: An areal perspective with a quantitative approach Linguistic Typology, 2018, 22, 233-275.</p> </blockquote>
CNBH-10 m: A first Chinese building height at 10 m resolution
<p>Building height is a crucial variable in the study of urban environments, regional climates, and human-environment interactions. However, high-resolution data on building height, especially at the national scale, are limited. Fortunately, high spatial-temporal resolution earth observations, harnessed using a cloud-based platform, offer an opportunity to fill this gap. We describe an approach to estimate 2020 building height for China at 10 m spatial resolution based on all-weather earth observations (radar, optical, and night light images) using the Random Forest (RF) model. Results show that our building height simulation has a strong correlation with real observations at the national scale (RMSE of 6.1 m, MAE = 5.2 m, <em>R</em> = 0.77). The Combinational Shadow Index (CSI) is the most important contributor (15.1%) to building height simulation. Analysis of the distribution of building morphology reveals significant differences in building volume and average building height at the city scale across China. <a href="https://www.sciencedirect.com/topics/earth-and-planetary-sciences/macao">Macau</a> has the tallest buildings (22.3 m) among Chinese cities, while Shanghai has the largest building volume (298.4 10<sup>8</sup> m<sup>3</sup>). The strong correlation between modelled building volume and socio-economic parameters indicates the potential application of building height products. The building height map developed in this study with a resolution of 10 m is open access, provides insights into the 3D morphological characteristics of cities and serves as an important contribution to future urban studies in China.</p>
List of Chinese Rotarians in Shanghai
<p>This file contains the complete list of the (113) Chinese members of the Rotary Clubs of Shanghai and Shanghai West from 1919 to 1948. It was compiled from the series of rosters available at the Archives of Rotary International (Evanston, Ill). The file contains not only membership but also biographical data extracted from contemporary <em>Who's Who</em>. For each member, we provide the following information: </p> <ol> <li><em>Membership</em>: Name (Chinese, pinyin, English/Wade-Giles transliteration), club of affiliation (Shanghai or Shanghai West), nickname, year of joining/leaving the club, age at joining the club, duration of membership, classification (the profession he represented), main affiliation (the institution he worked for), position, nature and degree of commitment to the club. </li> <li><em>Biographical Data</em>: Birth date (year), place of birth (locality, province), ancestry/native place, country of education, university of graduation, academic discipline, highest degree obtained, level of social activity (number of positions held during their career), social-professional mobility (range of sectors in which they worked), geographical mobility (number of places visited in their lifetime). </li> </ol>
Genotyping of the Chinese Spring x Renan mapping population with the TaBW280K SNP array
<p>The TaBW280K SNP array (Rimbert et al., PLoS ONE 2018) was used to genotype 430 Single Seed Descent (SSD) individuals<br> derived from a cross between Chinese Spring and Renan (CsRe; Choulet et al., Science 2014). Out of the 280,226 probesets, 85,276 were found to be polymorphic between the two parental lines and PHR on the population. Eventually, 83,721 (98.2%) SNPs were genetically mapped in 21 linkage groups corresponding to the 21 chromosomes of bread wheat, with no unlinked markers. This file contains the genotyping data of the 430 SSD lines.<br> </p>
Supplement for "Using Phylogenetic Networks to Model Chinese Dialect History"
<p>This is the supplementary material accompanying the paper "Using Phylogenetic Networks to Model Chinese Dialect History", which appeared in 2014 in "Language Dynamics and Change" (volume 4, issue 2).</p>
Chinese Transcription of Buddhist Terms in the Late Hàn Dynasty - Dataset
<p>This dataset is a compilation of Chinese transcriptions of Buddhist terms produced by translators<br> from the late Hàn period. It is a compilation of the previous works of (Coblin, 1983), (Karashima,<br> 2010), (Vetter, 2012), (Hill, Nattier, Granger, & Kollmeier, 2020) for the Chinese transcriptions. To<br> these were added phonological reconstructions of the Chinese terms for late Hàn from (Schuessler,<br> 2009) and Middle Chinese from (Baxter & Sagart, 2014a), as well as the Gandhari equivalents of<br> Sanskrit and Pāli terms from (Baums & Glass, 2002). This dataset, shared on Zenodo, aims at<br> being the new state-of-the-art dataset on Buddhist transcription material and can be used by anyone<br> working on Hàn Chinese phonology and will help better understanding the possible language sources<br> of the Chinese transcriptions, as well as the phonology of the target Chinese dialects.</p>
Wikipedia: wikipedia-zh (Chinese)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p>維基百科(Wikipedia,聆聽i/ˌwɪkᵻˈpiːdi.ə/或聆聽i/ˌwɪkiˈpiːdi.ə/)是一個自由內容、公開編輯且多語言的網络百科全書協作计划,透過Wiki技術使得包括您在內的所有人都可以簡單地使用網頁瀏覽器修改其中的內容。維基百科的名称取自於本網站核心技術「Wiki」以及具有百科全書之意的「encyclopedia」共同創造出來的新混成詞「Wikipedia」,當前維基百科由维基媒體基金會負責運營。 維基百科主要是由来自互联网上的志願者共同合作編寫而成,任何使用网络進入維基百科的用户都可以編寫和修改裡面的文章,但是在一些情况下為了避免擾亂或者破壞可能會限制編輯功能。我們可以自由選擇使用匿名、化名或者直接用真實身份來編輯維基百科。與傳統的百科全書相比,在互联网上運作的維基百科其文字和絕大部分圖片使用創用CC 姓名標示-相同方式分享 3.0協議和GNU自由文件授權條款來提供每個人自由且免費的資訊,任何人都可以成為條目的作者,以及在遵守协议并標示來源後直接複製、使用以及发布這些內容。<p></p>https://zh.wikipedia.org
Raw frequency data: Thoughts on "Reliable" Learner's Vocabularies for Classical and Literary Chinese
<p>This dataset includes the raw frequency counts (classical_chinese_learners_vocabularies_raw_frequencies.zip) used in the article Thoughts on “Reliable” Learner’s Vocabularies for Classical and Literary Chinese. </p> <p>Corpus I – Micheal Loewe (1993)’s <em>Early Chinese Texts</em><br> Corpus II – Official Histories (zhengshi 正史)<br> Corpus III Six Novels (xiaoshuo 小說), as defined in Hsia 1968</p> <p>The download includes one folder per corpus, structured as follows:</p> <ul> <li>xx_corpus.csv > list of texts and sources / used versions, token and type counts</li> <li>xx_freq_1-1.csv > unigram / character frequencies and counts</li> <li>xx_freq_1-4.csv > 1 to 4 character word frequencies and counts, "words" according to Hanyu da cidian 漢語大詞典 (Luo 1986–1994))</li> <li>xx_freq_2-4.csv > 2 to 4 character words</li> </ul> <p>Additionally, pca_zhengshi_vs_loewe_vs_xiaoshuo.html is an interactive version of the Principal Component Analysis (PCA) presented in the article, texts from the three corpora are represented using the 1.000 most frequent 1–4 character combinations from the dataset.</p>
CBFdataset: A Dataset of Chinese Bamboo Flute Performances
<p><em>CBFdataset</em> is<em> </em>a dataset of Chinese bamboo flute (CBF) performances, created for ecologically valid analysis of music playing techniques in context.</p> <p>The dataset comprises monophonic recordings of classic CBF pieces and isolated playing techniques, recorded by 10 professional CBF performers; and expert annotations of seven playing techniques: vibrato, tremolo, trill, flutter-tongue (FT), acciaccatura, portamento, and glissando. The recorded pieces include <em>Busy Delivering Harvest (BH)</em> 扬鞭催马运粮忙, <em>Jolly Meeting (JM)</em> 喜相逢, <em>Morning (Mo)</em> 早晨, <em>and Flying Partridge (FP)</em> 鹧鸪飞. All data was recorded in a professional recording studio using a Zoom H6 recorder at 44.1kHz/24-bits. The difference between different Versions 1.2, 1.1, and 1.0:</p> <ul> <li>V1.2 is the complete CBFdataset with a total duration of 2.6 hours.</li> <li>V1.1 splits the CBFdataset into two subsets according to playing technique types: CBF-periDB and CBF-petsDB. The former contains all the full-length pieces, isolated playing techniques, and annotations of four periodic modulations: vibrato, tremolo, trill, and flutter-tongue. The latter comprises the same full-length recordings, isolated playing techniques, and annotations of three pitch evolution-based techniques: acciaccatura, portamento, and glissando.</li> <li>V1.0 includes only the CBF-periDB.</li> </ul> <p>Related updates, demos, and code for reproducibility are available at <a href="http://c4dm.eecs.qmul.ac.uk/CBFdataset.html">http://c4dm.eecs.qmul.ac.uk/CBFdataset.html</a>. Any queries, please feel free to contact Changhong at changhong.wang@telecom-paris.fr. Please cite the following paper when using this dataset:</p> <p>Changhong Wang, Emmanouil Benetos, Vincent Lostanlen, and Elaine Chew, "Adaptive Scattering Transforms for Playing Technique Recognition," <em>IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)</em>, 30 (2022): 1407-1421.</p>
Indices of the supply of Chinese Higher Education
<p>The higher education indices for 31 Chinese provinces proposed by Borsi, Valerio Mendoza, and Comim (2021). </p>
CLDF dataset derived from Beijing Daxue's "Chinese Character Pronunciations" from 1962
<p>Cite the source of the dataset as:</p> <blockquote> <p>Běijīng Dàxué 北京大学 (1962): Hànyǔ fāngyán zìhuì 漢語方音字彙 [Chinese dialect character pronunciation list]. Beijing: Wenzi Gaige.</p> </blockquote>
The allocation of Chinese and Indian development finance in Nepal and its influence on local election results
<p>This realease contains the used datasets and calculations for my Bachelorthesis about Indian and Chinese allocation of Overall Development Assistance in Nepal.</p>
CLDF dataset derived from Beijing University's "Chinese Dialect Vocabularies" from 1964
<p>Cite the source of the dataset as:</p> <blockquote> <p>Běijīng Dàxué 北京大学 (1964): Hànyǔ fāngyán cíhuì 汉语方言词汇 [Chinese dialect vocabularies]. Beijing: Wenzi Gaige.</p> </blockquote>
बोधगया Bodhgaya (बिहार). Photograph of Chinese Inscription.
<p><a href="https://siddham.network/inscription/inch0001/">INCH0001</a> बोधगया Bodhgaya (बिहार). Photograph of Chinese Inscription (with date equivalent to 1021-22 CE). British Museum 1897,0528.0.31 a, presented by A. W. Franks. © Trustees of the British Museum</p> <p>[This inscription is number 1 in Cunningham's listing (Cunningham 1892, 69). Note: Cunningham (1892), 69 refers to Plate XXX, fig. 2 but this is an error; he means fig.1. The find-spot is mentioned in Cunningham (1892), 38, location Z2, just east of the main temple.]</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.