Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
51,140
datasets available to search
ShareScore release 0.7.1
Dataset results
51,140 results for “China”
Foundation Species in Forest Dynamics Plots of China 2004-2014
Foundation species structure forest communities and ecosystems but are difficult to identify without long-term observations or experiments. We used statistical criteria---outliers from size-frequency distributions and scale-dependent negative effects on alpha diversity and positive effects on beta diversity---to identify candidate foundation woody plant species in 12 large forest-dynamics plots spanning 26 degrees of latitude in China. We used these data to: [1] identify candidate foundation species in Chinese forests; [2] test the hypothesis---based on observations of a mid-latitude peak in functional trait diversity and high local species richness but few numerically dominant species in tropical forests---that foundation woody plant species are more frequent in temperate than tropical or boreal forests; and [3] compare these results with data from the Americas to suggest candidate foundation genera in Northern Hemisphere forests. Using the most stringent criteria, only two species of Acer, the canopy tree Acer ukurunduense and the shrubby treelet Acer barbinerve, were identified in temperate plots as candidate foundation species. Using more relaxed criteria, we identified four times more candidate foundation species in temperate plots (including species of Acer, Pinus, Juglans, Padus, Tilia, Fraxinus, Prunus, Taxus, Ulmus, and Corlyus) than in (sub)tropical plots (the shrubs or treelets Aporosa yunnanensis, Ficus hispida, Brassaiopsis glomerulata, and Orophea laui). Species diversity of co-occurring woody species was negatively associated with basal area of candidate foundation species more frequently at 5- and 10-m spatial grains (scale) than at a 20-m grain. Conversely, Bray-Curtis dissimilarity was positively associated with basal area of candidate foundation species more frequently at 5-m than at 10- or 20-m grains. Using either stringent or relaxed criteria supported the hypothesis that foundation species are more common in mid-latitude temperate forests. Com
Time series of carbon dioxide fluxes measured with eddy covariance for Danjiangkou Reservoir in Hubei Province, China during 2022-2024
This dataset contains half-hourly micrometeorological and eddy covariance flux measurements of carbon dioxide (CO₂) collected over the water surface of the Danjiangkou Reservoir in Hubei Province, China, from April 2022 to November 2024. The eddy covariance tower was installed at the deepest point of the reservoir, which serves as a critical water source for water supply and regional ecological functions in the middle reaches of the Yangtze River. Measurements were obtained using a LI-COR eddy covariance system (LI-COR Biosciences, Lincoln, NE, USA), and fluxes were calculated using EddyPro software (version 7.0.6). The dataset includes CO₂ and CH₄ fluxes as well as supporting micrometeorological, radiation, and water temperature measurements. All data were processed following established best practices for eddy covariance measurements, including comprehensive quality assurance and quality control procedures, which are fully documented and included with the dataset.
Increased inflammation and oxidative stress caused by accumulated metal particle exposure among metro station staff, Tianjin, China, 2023
Metro is a significant part of world transport, delivering over 58 billion passengers annually. The dilution effect of particulate matter (PM) from natural ventilation was limited in underground metro stations. What's worse, train operation processes generated PM rich in heavy metals. Though PM pollution in metro stations was reported widely, there is limited evidence of the adverse health effect of metro station PM. This dataset collected urinary samples from 74 metro station staff from three different metro stations in Tianjin, China, for inflammation and oxidative stress biomarker tests to better understand the potential health effects induced by metal particulates in metro stations. Also, an indoor air quality survey was conducted simultaneously in the metro stations.
ChinaHighNO₂: Daily Seamless 1 km Ground-Level NO₂ Dataset for China (2019–Present)
<p>ChinaHighNO<sub>2</sub> is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) ground-level NO<sub>2</sub> dataset for China <strong>from 2019 to the present</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.93, a root-mean-square error (RMSE) of 4.89 µg m<sup>-3</sup>, and a mean absolute error (MAE) of 3.48 µg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighNO<sub>2</sub> dataset in your scientific research, please cite the following references (Wei et al., EST, 2022; Wei et al., ACP, 2023):</p> <ul> <li> <p>Wei, J., Liu, S., Li, Z., Liu, C., Qin, K., Liu, X., Pinker, R., Dickerson, R., Lin, J., Boersma, K., Sun, L., Li, R., Xue, W., Cui, Y., Zhang, C., and Wang, J. <a href="https://weijing-rs.github.io/publications/Wei_et_al-EST-2022.pdf">Ground-level NO<sub>2</sub> surveillance from space across China for high resolution using interpretable spatiotemporally weighted artificial intelligence</a>. <em>Environmental Science & Technology</em>, 2022, 56(14), 9988–9998. https://doi.org/10.1021/acs.est.2c03834</p> </li> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>. <em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511–1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> </ul> <p><strong>Note that the ChinaHighNO<sub>2 </sub>dataset is also available for periods prior to 2019, but at a spatial resolution of 10 km:</strong></p> <p> all (including <strong>daily</strong>) data for the years <strong>2008–2018 </strong>is accessible at: <strong><a href="https://doi.org/10.5281/zenodo.4641542">https://doi.org/10.5281/zenodo.4641542</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
ChinaHighNO₂: Daily Seamless 10 km Ground-Level NO₂ Dataset for China (2008–2018)
<p>ChinaHighNO<sub>2</sub> is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 10 km (i.e., D10K, M10K, and Y10K) ground-level NO<sub>2</sub> dataset for China <strong>from 2008 to 2018</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.84, a root-mean-square error (RMSE) of 7.99 µg m<sup>-3</sup>, and a mean absolute error (MAE) of 5.34 µg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighNO<sub>2</sub> dataset in your scientific research, please cite the following references (Wei et al., ACP, 2023; Wei et al., EST, 2022):</p> <ul> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>. <em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511–1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> <li> <p>Wei, J., Liu, S., Li, Z., Liu, C., Qin, K., Liu, X., Pinker, R., Dickerson, R., Lin, J., Boersma, K., Sun, L., Li, R., Xue, W., Cui, Y., Zhang, C., and Wang, J. <a href="https://weijing-rs.github.io/publications/Wei_et_al-EST-2022.pdf">Ground-level NO<sub>2</sub> surveillance from space across China for high resolution using interpretable spatiotemporally weighted artificial intelligence</a>. <em>Environmental Science & Technology</em>, 2022, 56(14), 9988–9998. https://doi.org/10.1021/acs.est.2c03834</p> </li> </ul> <p><strong>Note that the ChinaHighNO<sub>2</sub> dataset was improved to a 1 km resolution after 2019:</strong></p> <p> all (including <strong>daily</strong>) data for the years after <strong>2019</strong><strong> </strong>are accessible at: <strong><a href="https://doi.org/10.5281/zenodo.4571660">https://doi.org/10.5281/zenodo.4571660</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
NECCPB-1: The first cropland parcel boundary dataset from meter-level imagery of Northeast China
<p>The Northeast China Plain is one of the world's three largest black soil regions, characterized by high organic matter content, rich nutrients, and strong water retention capabilities. Suitable climate conditions and abundant rainfall promote the growth of crops such as corn, soybeans, and rice, making it one of the main grain production bases in China, accounting for about one-fifth of the country's grain output. The grain production in the Northeast China black soil region is crucial for food security in China and globally. This area's farmland parcels are the basic units of agricultural production and the cornerstone of precision agriculture management, providing detailed information on cultivated land location, boundaries, shape, and area. Utilizing this parcel-scale information, governments and farm managers can devise more precise planting strategies and optimize management methods, thereby enhancing the quality and productivity of crops, ensuring a continuous food supply, and promoting sustainable agricultural development.</p> <p>The first cropland parcel boundary dataset from meter-level imagery of Northeast China (NECCPB-1) was developed based on deep learning models and a custom-designed automatic parcel merging strategy. A total of 10.22 TB of very-high-resolution (VHR) imagery was downloaded and uploaded, covering the entire region of Northeast China and an area of 1,240,000 km². After further removal of non-cropland regions based on phenological differences, 32,395,946 parcels were obtained.</p> <p> Rigorous validation using manually drawn reference parcels demonstrated that this dataset had high accuracy in parcel delineation (Extraction Precision, EP: 0.85) and high consistency with the reference parcels (|Completeness Deviation|, |CompD|: 0.02; Intersection over Union, IoU: 0.90). Further comparison with official Third Survey reports confirmed the high reliability of the NECCPB-1 dataset, which exhibited an average relative difference of -3.6% and an absolute relative difference of 9.8%.</p> <p>A series of cross-validations with seven widely used cropland datasets (ESA_GLC10, ESRI_GLC10, FROM_GLC10, CLD10, GLAD250, GFSAD30, and SinoLC-1). The recall, precision, and F1 scores of the NECCPB-1 were calculated as 0.91, 0.93, and 0.92, respectively, using publicly validated sample points of land cover. Moreover, NECCPB-1 performed best regarding cropland completeness, achieving an intersection ratio (IR) of 0.94, as calculated using the reference parcels.</p> <p>Due to the extensive size of the dataset and potential policy considerations, access to the data will be granted based on specific inquiries. Please contact us at zhengjia@iga.ac.cn or guotianhao@iga.ac.cn for further details. Please indicate your purpose and other details. Thank you!</p>
Japan-Educated Officials in China's Wartime Central Administration (1944)
<p>This dataset is made up of two files:</p> <ol> <li>"CNKI-20231014003254589" is the raw extraction of bibliographical references from CNKI on the Chinese students who stuided in Japan before 1945. It consists in the main academic outputs on the topic, including journal article,s M.A. thesis, and doctoral dissertations.</li> <li>"CNKIjp2" is the pre-processed and cleaned file that was used for statistical analysis and topic modeling.</li> </ol> <p>The markdown script presents the core of the methodological framework for my study of Chinese historiography on the Japan-educated students in the late imperial and republican period. I developed this script as part of the paper titled "Japan-Educated Officials in China’s Wartime Central Administration (1944)". In this paper, I present a review of the literature on the Chinese students who went to Japan to study between 1896 and 1945. While the literature in English and Japanese is small and allwo for close reading, the literature in Chinese is massive. To explore this literature and identify research trends, I chose to apply topic modeling to the dataet of references extracted from CNKI (知網). The documentary basis consists in the abstracts of the academic outputs.</p>
EU MarcoPolo project | SO2 emission inventory over China
<p>The aposteriori SO<sub>2</sub> emissions for year 2014, in the domain from 102°E to 132°E and from 15°N to 55°N, in a 0.25°x0.25° spatial resolution and monthly temporal resolution, have been provided to the MarcoPolo project and can be found at <a href="http://users.auth.gr/mariliza/MarcoPolo/SO2_EmissionInventory/">http://users.auth.gr/mariliza/MarcoPolo/SO2_EmissionInventory/</a>. For details on the creation of the inventory refer to <a href="http://users.auth.gr/mariliza/MarcoPolo/D3.4_SO2_emission_estimates.pdf">http://users.auth.gr/mariliza/MarcoPolo/D3.4_SO2_emission_estimates.pdf</a> and for the inclusion of the SO2 emission inventory to the MarcoPolo Emission Database refer to: <a href="http://users.auth.gr/mariliza/MarcoPolo/D4.2_DescriptionMarcoPoloInventory.pdf">http://users.auth.gr/mariliza/MarcoPolo/D4.2_DescriptionMarcoPoloInventory.pdf</a> as well as <a href="http://users.auth.gr/mariliza/MarcoPolo/D4.3_assessment_impact_updated_emission_inventories_v2.0.pdf">http://users.auth.gr/mariliza/MarcoPolo/D4.3_assessment_impact_updated_emission_inventories_v2.0.pdf</a> .</p> <p>The main reference to this dataset is found here:</p> <p>Koukouli, M. E., Theys, N., Ding, J., Zyrichidou, I., Mijling, B., Balis, D., and van der A, R. J.: Updated SO<sub>2</sub> emission estimates over China using OMI/Aura observations, Atmos. Meas. Tech., 11, 1817–1832, https://doi.org/10.5194/amt-11-1817-2018, 2018.</p> <p>The netcdf data files contain the following structure:</p> <ul> <li>Dimensions <ul> <li>lat = 129</li> <li>lon = 121</li> </ul> </li> <li>Attributes <ul> <li>author = "MariLiza Koukouli"</li> <li>contact information = "mariliza@auth.gr"</li> <li>institution = "Laboratory of Atmospheric Physics, Aristotle University of Thessaloniki"</li> <li>time frame = "2014"</li> <li>sector classification = "total emissions"</li> <li>emis_cat_name = "sulphur dioxide emissions"</li> <li>source_type_name = "sulphur dioxide emissions"</li> <li>pollutant_description = "updated sulphur dioxide emissions based on the CHIMERE model running the MEIC emissions and the OMI/Aura observations"</li> <li>unit_emissions = "Mg/month"</li> <li>nodata_value = "-9999.0"</li> </ul> </li> <li>Variables <ul> <li>float emissions(lon, lat)</li> </ul> </li> </ul>
30 m Normalized Difference Vegetation Index Maps of Pure Pixels over China for Estimation of Fractional Vegetation Cover (2014, 2018, 2022)
<p>Using multi-angle remote sensing data, we generated 30-m maps for the normalized difference vegetation index (NDVI) of fully-covered vegetation (<em>Vv</em>) and bare soils (<em>Vs</em>) across China in 2014, 2018 and 2022. These pixel-wise <em>Vv</em> and <em>Vs</em> maps can be integrated with the vegetation index (VI)-based model to facilitate the accurate and rapid estimation of fractional vegetation cover (FVC) across various spatial resolutions and large scales. The products were produced using a multi-angle algorithm (MultiVI), which effectively addressed the spatial variability inherent in <em>Vv</em> and <em>Vs</em> and enhanced the accuracy of FVC estimations in comparison to traditional statistical methods. The estimated FVC demonstrated a root mean square deviation (RMSD) of approximately 0.1 when evaluated against field-measured FVC across different experimental sites.</p>
The Biomass and Plant Functional Traits of Leymus chinensis Affected by Genotypic Diversity and Soil Nitrogen Addition through a Two-year Experiment, Tianjin, China, 2021-2023
In order to investigate the effects of soil nitrogen addition on the genotypic diversity of Leymus chinensis, 12 genotypes of Leymus chinensis were used as plant material and a two-factor experimental design was carried out in this study. Factor one was genotypic diversity of L. chinensis, including three levels: mono-genotype (G1), three genotypes (G3), and six genotypes (G6). Factor two was the soil nitrogen addition level, which included four levels: no nitrogen addition (N0), 2.5 g N/(m²·a) nitrogen application (N2.5), 5 g N/(m²·a) nitrogen application (N5), and 10 g N/(m²·a) nitrogen application (N10). Each treatment had 12 combinations as replicates, and 12 genotypes of L. chinensis were used. The frequency of each genotype was standardized across all treatment levels of genotypic diversity × soil nitrogen addition. The experiment commenced in September 2021 and soil nitrogen was applied every 2 months. Plants were cultivated in the experimental field at Nankai University, but were moved to a greenhouse for overwintering from November to February each year. During the experiment, there were no stresses or disturbances such as shading, drought, or insect feeding; weeds were regularly removed.
Working time, energy throughput and value added embodied in production, consumption and trade by subsectors for the US, the EU, China and rest of the world (2011)
<p>This repository contains the data needed to reproduce the results in:</p> <p>Pérez-Sánchez, L., Velasco-Fernández, R., Giampietro, M., The international division of labor and embodied working time in trade for the US, the EU and China, Ecological Economics. <a href="http://doi.org/10.1016/j.ecolecon.2020.106909">https://doi.org/10.1016/j.ecolecon.2020.1069097</a></p> <p>Sources of data are specified in the dataset (under tab "references")</p> <p> </p>
Relations in the Biographical Dictionary of Republican China - Standardized output
<p>This dataset contains the data on relations in the BDRC. It is based on the raw data output to be found in this collection. This file retained only the person-to-person relations. It served as a reference file to create the edge and node lists used for SNA under Cytoscape. All the corresponding networks are available as interactive networks in the <a href="http://public.ndexbio.org/#/group/4ea8024e-094c-11eb-948d-0ac135e8bacf?searchType=All&searchString=bdrc&searchTermExpansion=false">ENP-China Group</a> on the NDEx platform.</p>
ChinaHighPM2.5: VIIRS 6 km Ground-level PM2.5 Dataset for China
<p>ChinaHighPM<sub>2.5</sub> is one of the series of long-term, full-coverage, high-resolution, and high-quality datasets of ground-level air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from the big data (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence by considering the spatiotemporal heterogeneity of air pollution.</p> <p>This is the VIIRS derived yearly 6 km ground-level PM<sub>2.5</sub> dataset in China from 2013 to 2018, and this dataset yields a high quality with a cross-validation coefficient of determination (CV-R<sup>2</sup>) reaching 0.88 and a root-mean-square error (RMSE) of 16.52 µg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighPM<sub>2.5</sub> dataset for related scientific research, please cite the corresponding reference (Wei et al., TGRS, 2022):</p> <ul> <li> <p>Wei, J., Li, Z., Sun, L., Xue, X., Ma, Z., Liu, L., Fan, T., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-TGRS-2022.pdf">Extending the EOS long-term PM<sub>2.5</sub> data records since 2013 in China: application to the VIIRS Deep Blue aerosol products</a>. <em>IEEE Transactions on Geoscience and Remote Sensing</em>, 2022, 60, 4100412. https://doi.org/10.1109/TGRS.2021.3050999</p> </li> </ul> <p><strong>More CHAP datasets of different air pollutants can be found at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
ChinaHighCO: Daily Seamless 1 km Ground-Level CO Dataset for China (2019–Present)
<p>ChinaHighCO is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) ground-level CO dataset for China <strong>from 2019 to the present</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.80, a root-mean-square error (RMSE) of 0.29 mg m<sup>-3</sup>, and a mean absolute error (MAE) of 0.16 mg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighCO dataset in your scientific research, please cite the following reference (Wei et al., ACP, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>. <em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511–1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> </ul> <p><strong>Note that the ChinaHighCO<sub> </sub>dataset is also available for periods prior to 2019, but at a spatial resolution of 10 km:</strong></p> <p> all (including <strong>daily</strong>) data for the years <strong>2013–2018 </strong>are accessible at: <strong><a href="https://doi.org/10.5281/zenodo.4641530">https://doi.org/10.5281/zenodo.4641530</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
ChinaHighSO₂: Daily Seamless 1 km Ground-Level SO₂ Dataset for China (2019–Present)
<p>ChinaHighSO<sub>2</sub> is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) ground-level SO<sub>2</sub> dataset for China <strong>from 2019 to the present</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.84, a root-mean-square error (RMSE) of 10.07 µg m<sup>-3</sup>, and a mean absolute error (MAE) of 4.68 µg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighSO<sub>2</sub> dataset in your scientific research, please cite the following reference (Wei et al., ACP, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>. <em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511–1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> </ul> <p><strong>Note that the ChinaHighSO<sub>2 </sub>dataset is also available for periods prior to 2019, but at a spatial resolution of 10 km:</strong></p> <p> all (including <strong>daily</strong>) data for the years <strong>2013–2018 </strong>are accessible at: <strong><a href="https://doi.org/10.5281/zenodo.4641538">https://doi.org/10.5281/zenodo.4641538</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
Survey Data on Apple Farming in China: Agronomic Management, Advisory Channels, and Profitability
<p>The Survey results and original data are stored in a directory structured as the table:</p> <table style="width: 100%; height: 223.938px;"> <tbody> <tr style="height: 19.5938px;"> <td style="width: 18.8269%; height: 19.5938px;"><strong>Type</strong></td> <td style="width: 21.7597%; height: 19.5938px;"><strong>File Name</strong></td> <td style="width: 59.4134%; height: 19.5938px;"><strong>Description</strong></td> </tr> <tr style="height: 47.5938px;"> <td style="width: 18.8269%; height: 47.5938px;"> <p>Raw_Data_Spearate_Source</p> </td> <td style="width: 21.7597%; height: 47.5938px;">raw_data_english_telephone.xlsx</td> <td style="width: 59.4134%; height: 47.5938px;">Translated data in English corresponding to the Chinese telephone interview data</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 18.8269%; height: 19.5938px;"> </td> <td style="width: 21.7597%; height: 19.5938px;">raw_data_english_wechat.xlsx</td> <td style="width: 59.4134%; height: 19.5938px;">Translated data in English corresponding to the Chinese Wechat Mini Program data</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 18.8269%; height: 19.5938px;">Raw_Data_Total</td> <td style="width: 21.7597%; height: 19.5938px;">raw_data_english_total.xlsx</td> <td style="width: 59.4134%; height: 19.5938px;">Combined data from raw_data_english_telephone.xlsx and raw_data_english_wechat.xlsx</td> </tr> <tr style="height: 39.1875px;"> <td style="width: 18.8269%; height: 39.1875px;">Apple_Statistical_Data</td> <td style="width: 21.7597%; height: 39.1875px;">apple_2022_statistical_data.xlsx</td> <td style="width: 59.4134%; height: 39.1875px;">Contains data on apple planting area, production, and yield sourced from the China Statistics Bureau, along with the number of survey questionnaires collected from various provinces</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 18.8269%; height: 19.5938px;"> </td> <td style="width: 21.7597%; height: 19.5938px;">province_eng.xlsx</td> <td style="width: 59.4134%; height: 19.5938px;">Contains the English version of the provinces' names</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 18.8269%; height: 19.5938px;">Map_Boundary_line</td> <td style="width: 21.7597%; height: 19.5938px;">national_boundary_line.shp</td> <td style="width: 59.4134%; height: 19.5938px;">The country boundaires of China</td> </tr> <tr style="height: 19.5938px;"> <td style="width: 18.8269%; height: 19.5938px;"> </td> <td style="width: 21.7597%; height: 19.5938px;">province_boundary.shp</td> <td style="width: 59.4134%; height: 19.5938px;">The province boundaries of China</td> </tr> </tbody> </table> <p>For privacy reasons, personally identifiable information such as respondents’ names, telephone numbers, and specific addresses has been anonymized in the dataset. The file <em>raw_data_english_total.xlsx</em> contains 96 columns, each corresponding to a question in the questionnaire.</p>
Modern China Geospatial Database - Main Dataset
<p>MCGD_Data_V2.2 contains all the data that we have collected on locations in modern China, plus a number of locations outside of China that we encounter frequently in historical sources on China. All further updates will appear under the name "MCGD_Data" with a time stamp (e.g., MCGD_Data2023-06-21)</p> <p>You can also have access to this dataset and all the datasets that the ENP-China makes available on GitLab: https://gitlab.com/enpchina/IndexesEnp</p> <p>Altogether there are 464,970 entries. The data include seven variables:<br>- Name: Place names and their variants in Chinese, pinyin, and any recorded transliteration<br>- Prov_Zh: Chinese province names in Chinese characters (新疆, 江蘇, 河北, etc.)<br>- Prov_Py: Chinese province names in pinyin<br>- LAT: Latitude coordinates<br>- LONG: Longitude coordinates<br>- LocID: Location identifiers<br>- NameID: Location name identifiers</p> <p>The Name IDs all start with H followed by seven digits. This is the internal ID system of MCGD.</p> <p>Locations IDs that start with "D" are data points extracted from China Historical GIS (Harvard University); those that start with "E" are locations extracted from the data points in Geonames or data points we have added from various map sources.</p> <p>One of the main features of the MCGD Main Dataset is the systematic collection and compilation of place names from non-Chinese language historical sources. Locations were designated in transliteration systems that are hardly comprehensible today, which makes it very difficult to find the actual locations they correspond to. This dataset allows for the conversion from these obsolete transliterations to the current names and geocoordinates.</p> <p>From June 2021 onward, we have adopted a different file naming system to keep track of versions. From MCGD_Data_V1 we have moved to MCGD_Data_V2. In June 2022, we introduced time stamps, which result in the following naming convention: MCGD_Data_YYYY.MM.DD. </p> <p> </p> <p><strong>UPDATES</strong></p> <p><strong>MCGD_Data2025_08_06</strong> introduces a significant update with the addition of the <strong>‘Code’</strong> column. This column categorizes place names as follows:</p> <ul> <li> <p><strong>A</strong>: Canonical Chinese name</p> </li> <li> <p><strong>C</strong>: Alternative Chinese name</p> </li> <li> <p><strong>P</strong>: Romanized name in pinyin</p> </li> <li> <p><strong>W</strong>: Romanized name in another transliteration system</p> </li> </ul> <p>When the codes <strong>P</strong> or <strong>W</strong> are doubled (<strong>PP</strong>, <strong>WW</strong>), this indicates that the place name does not match any existing Chinese name in the dataset. These unmatched names will be reviewed and linked progressively, rather than through a systematic batch process, due to their high volume.The coding system is designed to facilitate name-matching operations between MCGD and place names extracted from historical sources using programming tools. It also enables filtering for more precise and efficient matching. The dataset contains a total of <strong>472,749 entries</strong>.</p> <p>MCGD_Data2025_02_28 includes a major change with the duplication of all the locations listed under Beijing, Shanghai, Tianjin, and Chongqing (北京, 上海, 天津, 重慶) and their listing under the name of the provinces to which they belonge origially before the creation of the four special municipalities after 1949. This is meant to facilitate the matching of data from historical sources. Each location has a unique NameID. Altogether there are 472,818 entries</p> <p>MCGD_Data2025_02_27 inclues an update on locations extracted from Minguo zhengfu ge yuanhui keyuan yishang zhiyuanlu 國民政府各院部會科員以上職員錄 (Directory of staff members and above in the ministries and committees of the National Government). Nanjing: Guomin zhengfu wenguanchu yinzhuju 國民政府文官處印鑄局國民政府文官處印鑄局, 1944). We also made corrections in the Prov_Py and Prov_Zh columns as there were some misalignments between the pinyin name and the name in Chines characters. The file now includes 465,128 entries.</p> <p>MCGD_Data2024_03_23 includes an update on locations in Taiwan from the Asia Directories. Altogether there are 465,603 entries (of which 187 place names without geocoordinates, labelled in the Lat Long columns as "Unknown").</p> <p>MCGD_Data2023.12.22 contains all the data that we have collected on locations in China, whatever the period. Altogether there are 465,603 entries (of which 187 place names without geocoordinates, labelled in the Lat Long columns as "Unknown"). The dataset also includes locations outside of China for the purpose of matching such locations to the place names extracted from historical sources. For example, one may need to locate individuals born outside of China. Rather than maintaining two separate files, we made the decision to incorporate all the place names found in historical sources in the gazetteer. Such place names can easily be removed by selecting all the entries where the 'Province' data is missing.</p>
ChinaHighO₃: Daily Seamless 1 km Ground-Level O₃ Dataset for China (2000–Present)
<p>ChinaHighO<sub>3</sub> is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) ground-level maximum daily 8-hour average (MDA8) O<sub>3</sub> dataset for China <strong>from 2000 to the present</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.89, a root-mean-square error (RMSE) of 15.77 µg m<sup>-3</sup>, and a mean absolute error (MAE) of 10.48 µg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighO<sub>3</sub> dataset in your scientific research, please cite the following references (Yang et al., RSE, 2025; Wei et al., RSE, 2022):</p> <ul> <li>Yang, Z., Li, Z., Cheng, F., Lv, Q., Li, K., Zhang, T., Zhou, Y., Zhao, B., Xue, W., and Wei, J. <a href="https://weijing-rs.github.io/publications/Yang_et_al-RSE-2025.pdf" target="_blank" rel="noopener">Two-decade surface ozone (O<sub>3</sub>) pollution in China: enhanced fine-scale estimations and environmental health implications</a>. <em>Remote Sensing of Environment</em>, 2025, 317, 114459. https://doi.org/10.1016/j.rse.2024.114459</li> </ul> <ul> <li> <p>Wei, J., Li, Z., Li, K., Dickerson, R., Pinker, R., Wang, J., Liu, X., Sun, L., Xue, W., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-RSE-2022.pdf">Full-coverage mapping and spatiotemporal variations of ground-level ozone (O<sub>3</sub>) pollution from 2013 to 2020 across China</a>. <em>Remote Sensing of Environment</em>, 2022, 270, 112775. https://doi.org/10.1016/j.rse.2021.112775</p> </li> </ul> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
ChinaHighCO: Daily Seamless 10 km Ground-Level CO Dataset for China (2013–2018)
<p>ChinaHighCO is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 10 km (i.e., D10K, M10K, and Y10K) ground-level CO dataset for China <strong>from 2013 to 2018</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.80, a root-mean-square error (RMSE) of 0.29 mg m<sup>-3</sup>, and a mean absolute error (MAE) of 0.16 mg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighCO dataset in your scientific research, please cite the following reference (Wei et al., ACP, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>. <em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511–1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> </ul> <p><strong>Note that the ChinaHighCO dataset was improved to a 1 km resolution after 2019:</strong></p> <p> all (including <strong>daily</strong>) data for the years after <strong>2019 </strong>are accessible at: <strong><a href="https://doi.org/10.5281/zenodo.10477022">https://doi.org/10.5281/zenodo.10477022</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
ChinaHighSO₂: Daily Seamless 10 km Ground-Level SO₂ Dataset for China (2013–2018)
<p>ChinaHighSO<sub>2</sub> is part of a series of long-term, seamless, high-resolution, and high-quality datasets of air pollutants for China (i.e., ChinaHighAirPollutants, CHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>Here is the big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 10 km (i.e., D10K, M10K, and Y10K) ground-level SO<sub>2</sub> dataset for China <strong>from 2013 to 2018</strong>. This dataset exhibits high quality, with a cross-validation coefficient of determination (CV-R<sup>2</sup>) of 0.84, a root-mean-square error (RMSE) of 10.07 µg m<sup>-3</sup>, and a mean absolute error (MAE) of 4.68 µg m<sup>-3</sup> on a daily basis.</p> <p>If you use the ChinaHighSO<sub>2</sub> dataset in your scientific research, please cite the following reference (Wei et al., ACP, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Wang, J., Li, C., Gupta, P., and Cribb, M. <a href="https://weijing-rs.github.io/publications/Wei_et_al-ACP-2023.pdf">Ground-level gaseous pollutants (NO<sub>2</sub>, SO<sub>2</sub>, and CO) in China: daily seamless mapping and spatiotemporal variations</a>. <em>Atmospheric Chemistry and Physics</em>, 2023, 23, 1511–1532. https://doi.org/10.5194/acp-23-1511-2023</p> </li> </ul> <p><strong>Note that the ChinaHighSO<sub>2</sub> dataset was improved to a 1 km resolution after 2019:</strong></p> <p> all (including <strong>daily</strong>) data for the years after <strong>2019 </strong>are accessible at: <strong><a href="https://doi.org/10.5281/zenodo.10476944">https://doi.org/10.5281/zenodo.10476944</a></strong></p> <p><strong>More CHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.