Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,359
datasets available to search
ShareScore release 0.7.1
Dataset results
7,359 results for “Chinese”
Bodhgayā, Bihar. Photograph of Chinese Inscription. British Museum 1897,0528.0.30 a
<p>Bodhgayā, Bihar. Photograph of the Chinese Inscription dated in the second year of 明道 Míngdào (CE 1032-33) of the Song emperor 真宗 Zhēnzōng (see notes for explanation). Further data in SIDDHAM <a href="https://siddham.network/object/obch0004/">OBCH0004</a>. © British Museum.</p>
CLDF dataset derived from Hóu's "Phonological Database of Chinese Dialects" from 2004
<p>Cite the source of the dataset as:</p> <blockquote> <p>Hóu, J. (2004): Xiàndài Hànyǔ fāngyán yīnkù 现代汉语方言音库 [Phonological database of Chinese dialects]. Shànghǎi: Shànghǎi Jiàoyù.</p> </blockquote>
CLDF dataset derived from Liú et al.'s "Collection of Basic Words in Chinese Dialects" from 2007
<p>Cite the source of the dataset as:</p> <blockquote> <p>Líu, L.; Wáng, H.; Bǎi, Y. (2007): Xiàndài Hànyǔ fāngyán héxīncí, tèzhēng cíjí 现代汉语方言核心词·特征词集 [Collection of basic vocabulary words and characteristic dialect words in modern Chinese dialects]. Nánjīng: Fènghuáng.</p> </blockquote>
Mandarin Chinese IDS wordlist by Hsiao-jung Yu and Yifan Wang
<p>Cite the source of the dataset as:</p> <blockquote> <p>Hsiao-jung Yu and Yifan Wang. 2021. Mandarin Chinese dictionary. In: Key, Mary Ritchie & Comrie, Bernard (eds.) The Intercontinental Dictionary Series. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at https://ids.clld.org/)</p> </blockquote>
Qualitative Data on Effects of Early and Prolonged Parent-Child Separation: Understanding Mental Health of Separated-Reunited Chinese American Children
<div> <div> <div> <div>Early and prolonged parent-child separation due to parental migration or immigration may result in attachment disruption that can threaten the long-term mental health and functioning of affected children, and these risks can persist following reunification and through adulthood. Although sending infants back to the home country for rearing is often practiced among <em>Chinese</em> immigrants, especially low-income families, research has been sparse in understanding the long-term impact of early and prolonged parent-child separation and reunification on disparities in mental health and functioning among separated-reunited children and the mechanism through which such relationships may operate.</div> <div>Funded by National Institute on Minority Health and Health Disparities (NIMH), we collected semi-structured interview data from 24 parent-child dyads who have experienced separation. The data included interview scrpits with primary coding to understand the mental health impacts, risk/protective factors, and service needs among separated-reunited <em>Chinese</em> American children.</div> </div> </div> </div> <div> </div>
Chinese Temples in Bangkok (Ver 2024-11)
<div> <div>This is a geo-referenced survey of Chinese temples in the Bangkok Metropolitan Region. For the background of the survey and how it relates to previous surveys see: <div> <div>Marcus BINGENHEIMER, Paul MCBAIN: “In the Footsteps of Wolfgang Franke – Revisiting and Surveying Chinese Temples in Bangkok” Journal of the Siam Society Vol.112-1: 49–70. [Link to Open Access Journal](https://so06.tci-thaijo.org/index.php/pub_jss/article/view/274764)</div> <div> </div> <div>For metadata regarding the column names etc., see the markdown file ("chineseTemplesBangkok_v2024-11.md").</div> </div> </div> </div>
The 30 m land cover dataset for capturing land cover changes induced by ecological restoration from 1990 to 2022 on the Chinese Loess Plateau
<p>Continuous time-series of land cover is critical for attributing runoff, sediment and carbon changes on the Chinese Loess Plateau (CLP). However, current land cover products with annal temporal resolution lack spatial identification accuracy, particularly in capturing authentic changes of cropland, forest and grassland. To address these issues, a 30 m annual land cover dataset was proposed by the Yellow River Conservancy Commission (YRCC_LPLC) for the CLP from 1990 to 2022. Different levels of land cover were classified using different combinations of spectral, monthly and annual temporal and topographic features and Random Forest classifier. Compared to other land cover products (45.64%–73.38%), the accuracy of YRCC_LPLC has a better performance with an overall accuracy of 85.16%. The YRCC_LPLC is capable of capturing not only the explicit spatial variation but also the change direction and change time of land cover, especially for the most critical conversion of cropland into forest and grassland induced by implementation of Grain to Green Program on the CLP.</p>
CLDF dataset derived from Wang's "Basic Words in Chinese Dialects" from 2004
<p>Cite the source of the dataset as:</p> <blockquote> <p>Wang, F. 2004. BCD: basic words of Chinese dialects. Unpublished dataset. [Digital version in: List, J.-M. (2015): Network perspectives on Chinese dialect history. Bulletin of Chinese Linguistics 8. 42-67.]</p> </blockquote>
Data for paper on the evolution of Chinese characters
<p>This dataset contains all image, complexity and distinctiveness data that was used for:</p> <p>Han, S. J, Kelly, P., Winters, J., & Kemp, C. (2022). Simplification is not dominant in the evolution of Chinese characters. <em>Open Mind</em>.</p> <p>The code for this project can be found <a href="https://github.com/cskemp/chinesecharacters">here</a>. The file uploaded here is intended to replace the sample data folder that is available in the code repository.</p> <p>Our dataset includes data scraped from hanziyuan.net, as well as data from the following sources:</p> <p>Sun, C. C., Hendrix, P., Ma, J., & Baayen, R. H. (2018). Chinese lexical database (CLD): A large-scale lexical database for simplified Mandarin Chinese. Behavior Research Methods, 50(6), 2606–2629.</p> <p>Wikimedia Commons. (2021). Chinese characters decomposition. <a href="https://commons.wikimedia.org/wiki/Commons:Chinese_characters_decomposition">https://commons.wikimedia.org/wiki/Commons:Chinese_characters_decomposition</a></p> <p>Liu, C.-L., Yin, F., Wang, D.-H., & Wang, Q.-F. (2011). CASIA online and offline Chinese handwriting databases. In 2011 international conference on document analysis and recognition (pp. 37–41). <a href="https://doi.org/10.1109/ICDAR.2011.17">https://doi.org/10.1109/ICDAR.2011.17</a></p> <p>Chen, P.-C. (2020). Traditional Chinese handwriting dataset. GitHub. <a href="https://github.com/AI-FREE-Team/Traditional-Chinese-Handwriting-Dataset">https://github.com/AI-FREE-Team/Traditional-Chinese-Handwriting-Dataset</a></p>
Chinese Educational Mission Dataset (1872-1881)
<p>This series of 11 datasets is drawn from Rhoads, Edward J. M. <em>Stepping Forth into the World: The Chinese Educational Mission to the United States, 1872-81</em>. Hong Kong University Press, 2011.</p> <p>They document the 120 young Chinese who participated in the pioneering Chinese Educational Mission (CEM) in the United States (1872-1881). The first 8 files are drawn directly from the tables in Rhoads: </p> <ol> <li>Table 2.1 CEM students, by detachment (p.14-17)</li> <li>Table 5.1. Initial host family assignments (p.51-54)</li> <li>Table 7.1. CEM students in middle schools (by state and locality) (p. 90-94)</li> <li>Table 7.2 CEM students in public high schools (by state and locality) (p.96-99)</li> <li>Table 7.3 CEM students in private academies (by state and locality) (p.99-100)</li> <li>Table 8.1 CEM students in colleges (by academic year of enrollment) (p.116-118)</li> <li>Table 9.1 Deaths, dismissals, and withdrawals from the CEM (by date) (p.136)</li> <li>Table 9.2 CEM students in the June 1880 census (p.138-142)</li> </ol> <p>Based on these tables, I created three synthetic datasets which can be used for statistical and network analyses: </p> <ol> <li>cem_attributes: students' vital attributes, including their multiple names and transliteration, date and place of birth, and other attribute data (one row for each individual). </li> <li>cem_host: students' host families in the United States</li> <li>cem_education: students' educational curricula </li> </ol> <p>Each file contains two tabs, one for the data (data), one for the description of variables (key). Grey columns refer to the unstructured information given in the original source. </p>
Who's Who of American Returned Students 遊美同學錄 (1917): Affiliation Data (Chinese)
<p>This dataset is derived from the <em>Whoʻs Who of American Returned Students </em>遊美同學錄<em> </em>[<em>Youmei Tongxue Lu</em>] published in<em> </em>Peking [Beijing] in 1917, compiled by the Returned Students’ Information Bureau (Liumei xuesheng tongxunchu 留美學生通訊處) established at Tsinghua School in 1915. This book is crucial for documenting the early <em>liumei</em>'s experiences during the transitional period between the late Qing dynasty and the early years of the Republic (1911-). </p> <p>The dataset records all the institutions to which the students were affiliated in the course of their lives, including the educational institutions in which they studied in China, the United States, and other countries; the public or private organizations in which they were employed; as well as their memberships in clubs and associations. The names of organizations were retrieved automatically from the Chinese biographies using <a href="https://bookdown.enpchina.eu/rpackage/HistTextRManual.html#4_Named_Entity_Recognition">named entity recognition</a> (SpaCy model), then manually cleaned, classified, and validated by the author. </p> <p>The attached file contains three tabs for (1) the list of affiliations (data); (2) the classification of organizations (class), and (3) the description of variables (key). The dataset records a total of 2,883 affiliations, linking 401 unique individuals to 1,344 unique institutions, distributed as followed: </p> <table> <tbody> <tr> <td><strong>category</strong></td> <td><strong>n</strong></td> </tr> <tr> <td>education</td> <td>565</td> </tr> <tr> <td>association</td> <td>271</td> </tr> <tr> <td>administration</td> <td>132</td> </tr> <tr> <td>business</td> <td>110</td> </tr> <tr> <td>facility</td> <td>92</td> </tr> <tr> <td>media</td> <td>66</td> </tr> <tr> <td>government</td> <td>49</td> </tr> <tr> <td>factory</td> <td>30</td> </tr> <tr> <td>other</td> <td>22</td> </tr> <tr> <td>military</td> <td>7</td> </tr> </tbody> </table>
Chinese Engineers Relational Database (CERD) Bi-monthly Export
<p><strong>This is a bi-monthly export.</strong></p> <p>CERD is a database of engineers from the Chinese Republican period (1912–1949). Based on various digitised historical sources, it is a prosopographic catalogue of individuals, their education and their employment, and the institutions connected with it. Most biographical events have geographical information attached to them. The data can be used freely by researchers to answer individual research questions.</p> <p><strong>Citation recommendation:</strong></p> <p>Pelzer, Thorben, et al., eds. (2021–2023). Chinese Engineers Relational Database (CERD) (Version 1.7.0). Zenodo. http://doi.org/10.5281/zenodo.4075601.</p> <p><strong>Changelog:</strong></p> <p>1.7.0 (February 2023): Approx. 17,600 individuals, two additional sources<br> 1.6.0 (August 2022): Approx. 17,400 entries, completed sources, minor corrections, mergers, translations<br> 1.5.0 (June 2022): Approx. 17,300 entries, added additional memberships<br> 1.4.0 (April 2022): Approx. 16,800 entries, added additional sources, schooling<br> 1.3.0 (February 2022): Approx. 16,500 entries, added documentation, additional memberships<br> 1.2.0 (December 2021): Approx. 16,300 entries, added frequent CSV exports, source annotations, 5 missing <em>minglu</em> pages, additional association memberships<br> 1.1.0 (October 2021): Approx. 15,700 entries, added selected association memberships<br> 1.0.0 (August 2021): Approx. 15,400 entries [complete <em>gongchengren minglu</em> dataset milestone]<br> 0.5.0 (June 2021): Approx. 13,000 entries<br> 0.4.0 (April 2021): Approx. 10,500 entries<br> 0.3.0 (February 2021): Approx. 7,500 entries, as well as various corrections, mergers, translations<br> 0.2.0 (December 2020): Approx. 5,000 entries, as well as various corrections, mergers, translations<br> 0.1.0 (October 2020): Early version with approx. 3,000 entries</p> <p><strong>Online Access:</strong></p> <p>Via Heurist: <a href="https://home.uni-leipzig.de/cerd/">https://home.uni-leipzig.de/cerd/</a></p>
2.8M SNPs Chinese Spring RefSeq v2.1 dataset
<p>VCF file of 2,799,166 single nucleotide polymorphism (SNP) markers positioned onto the Chinese Spring reference assembly RefSeq v2.1 developed by the International Wheat Genome Sequence Consortium (IWGSC; Zhu et al., 2021). These SNPs were lifted from the 1,000 wheat exome project, originally positioned onto RefSeq v1.0 (He et al., 2019). The SNP projection from RefSeq v1.0 onto RefSeq v2.1 was accomplished using LiftOff (Shumate and Salzberg, 2021).</p> <p>References</p> <p>He F, Pasam R, Shi F, Kant S, Keeble-Gagnere G, Kay P, Forrest K, Fritz A, Hucl P, Wiebe K, et al: <strong>Exome sequencing highlights the role of wild-relative introgression in shaping the adaptive landscape of the wheat genome.</strong> <em>Nature Genetics </em>2019, <strong>51:</strong>896-904.</p> <p>Shumate A, Salzberg SL: <strong>Liftoff: accurate mapping of gene annotations.</strong> <em>Bioinformatics </em>2021, <strong>37:</strong>1639-1643.</p> <p>Zhu T, Wang L, Rimbert H, Rodriguez JC, Deal KR, De Oliveira R, Choulet F, Keeble-Gagnère G, Tibbits J, Rogers J, et al: <strong>Optical maps refine the bread wheat Triticum aestivum cv. Chinese Spring genome assembly.</strong> <em>The Plant Journal </em>2021, <strong>107:</strong>303-314..</p>
Data of Chinese treatment group for the research work "Disentangling material, social, and cognitive determinants of human behavior and belief".
<p>This repository contains data files of Chinese treatment group for the research work "Disentangling material, social, and cognitive determinants of human behavior and belief".</p>
Predictive nano-QSAR modeling of the cytotoxicity using epithelial cells obtained from Chinese hamster ovary (CHO-K1 cell line) for hybrid TiO2-based nanomaterials
<p>Results obtained from developed model indicated that the cytotoxicity of hybrid TiO2-based nanomaterials is related to additive electronegativity (χmix) of studied nanomaterials that are indirectly related to the electron generation and ROS formation. ROS production is the most common toxicity cause as discussed in the literature in the case of nanoparticles. The high efficiency of surface modified TiO2-based semiconductors can be attributed to the involvement of TiO2 band gap (Eg) excitation and absence of noble metals at the TiO2 surface. It can be expected that noble metals (i.e. Pd/Pt) may trap holes (h+), at the same time photo-generated electrons can be then transferred from the valence band to the conduction band of TiO2 and to its surface where redox processes were initiated. Thus, observed reduction of the electron–hole pair recombination influences the reactive oxygen species (ROS) formation and the photocatalytic redox process initiation.</p> <p>Since the electronegativity was positively correlated with the cytotoxicity it can be expected that some ions are released from the TiO2 surface easier than others.</p>
UBGG-3m: Fine-grained urban blue-green-gray landscape dataset for 36 Chinese cities based on deep learning network
<p>The UBGG dataset provides easily access and leverage to researchers and analysts, which is stored in the following Zenodo repository (<a href="https://doi.org/10.5281/zenodo.8053333">https://doi.org/10.5281/zenodo.8352777</a>). The UBGG dataset consists of two main components:</p> <ul> <li><strong>UBGG-3m: the fine-grained UBGG map product of 36 metropolises in China.</strong> The UBGG-3m dataset captures the intricate urban landscape features with remarkable precision, providing a detailed representation at an impressive 3-meter resolution. Fig. 1 in User Guides shows the classification results for 36 Chinese metropolises. Researchers can delve into the nuances of the UBGG continuum, gaining invaluable insights into the interplay between the blue, green, and gray elements of urban environments in each metropolis.</li> </ul> <ul> <li><strong>UBGGset:</strong> <strong>the large-volume sample dataset to support the UBGG deep learning research.</strong> Complementing the UBGG-3m dataset, UBGGset serves as a large-volume sample dataset specifically tailored to support and foster UBGG research endeavors (Fig. 2). The UBGGset consists of 14,627 sample images (without data augmentation), with dimensions of 256 pixels in length and width, covering an urban area of approximately 2,272 km<sup>2</sup>. The UBGGset was constructed with co-registered pairs of 3 m Planet images and fine-annotated urban landscapes labeled on 1 m Google Earth image. This dataset encompasses 15 typical cities, offering researchers a rich and diverse resource to drive exploration, analysis, and innovation in the field of urban landscape studies.</li> </ul> <p> </p> <p><strong>Citation format for paper and dataset:</strong></p> <p>[1] Zhiyu Xu, Shuqing Zhao. Fine-grained urban blue-green-gray landscape dataset for 36 Chinese cities based on deep learning network. <em>Sci Data</em> 11, 266 (2024). https://doi.org/10.1038/s41597-023-02844-2</p> <p>[2] Zhiyu Xu, Shuqing Zhao, Fine-grained urban landscape mapping reveals broad-scale homogeneity in urban environments,<br>Science Bulletin, (2024). https://doi.org/10.1016/j.scib.2024.03.060</p> <p>[3] Zhiyu Xu, Shuqing Zhao. UBGG-3m: Fine-grained urban blue-green-gray landscape dataset for 36 Chinese cities based on deep learning network (v1.0) [Data set]. (2023). Zenodo. https://doi.org/10.5281/zenodo.8352777</p>
Figs 22–29 in Three new Cryptochetum Rondani, 1875 (Diptera: Cryptochetidae) from Yunnan Province, China and an identification key to Chinese species
Figs 22–29. Wings of eight species of Cryptochetum. 22. C. curvatum Yang & Yang, 1996. 23. C. deltatum Yang & Yang, 1996. 24. C. tianmuense Yang & Yang, 2001. 25. C. acutulum Yang & Yang, 1996. 26. C. zalatilabium Xi & Yang, 2015. 27. C. kunmingense Yang & Yang, 1996. 28. C. fanjingshanum Yang & Yang, 1988. 29. C. maolanum Yang & Yang, 1996. Scale bar = 0.1 mm
Figs 15–17 in Three new Cryptochetum Rondani, 1875 (Diptera: Cryptochetidae) from Yunnan Province, China and an identification key to Chinese species
Figs 15–17. Cryptochetum longilingum sp. nov., holotype (CR154), ♂. 15. Head, lateral view. 16. Head, dorsal view. 17. Wing, dorsal view. Scale bar = 0.1 mm.
Figs 8–10 in Three new Cryptochetum Rondani, 1875 (Diptera: Cryptochetidae) from Yunnan Province, China and an identification key to Chinese species
Figs 8–10. Cryptochetum glochidiatusum sp. nov., holotype, ♂ (CR122). 8. Head, lateral view. 9. Head, dorsal view. 10. Wing, dorsal view. Scale bar = 0.1 mm.
Figs 1–3 in Three new Cryptochetum Rondani, 1875 (Diptera: Cryptochetidae) from Yunnan Province, China and an identification key to Chinese species
Figs 1–3. Cryptochetum euthyiproboscise sp. nov., holotype, ♂ (CR101). 1. Head, lateral view. 2. Head, dorsal view. 3. Wing, dorsal view. Scale bar = 0.1 mm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.