Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “Chinese local gazetter”
Mulberry Disasters in Chinese Local Gazetteers
<p>This dataset contains 404 mulberry disasters found in a digital collection of 4,000 Chinese local gazetteers (published in Erudition's Zhongguo Fangzhi Ku, or the Database of Chinese Local Gazetteers) that was curated during 2018 and 2019 by scholars at the Max Planck Institute for the History of Science to support their joint research paper entitled: “What Is Local Knowledge: Digital Humanities and Yuan Dynasty Disasters in Imperial China’s Local Gazetteers.” The paper appeared in the <em>Journal of Chinese History</em> in its 2020 spring issue. The authors use this dataset to ask what results the emerging methodology of analyzing data drawn from historical sources can produce for historical research, given that such data-driven analysis is inevitably quantitative and that many historians believe this would contradict the core value of studies in the humanities.</p> <p>This dataset is accompanied by a data paper to be published by Brill's Digital Concordances Platform. In the data paper, the authors give information about the full curation cycle of this dataset, including the research goal that motivated the curation, the sources that were used, and the curation/cleansing/preparation process. The authors also explain this dataset in detail, including the kind of information it contains and its overall temporal and geospatial distributions. In conclusion, they discuss the potential usages of the dataset. With this paper, we also hope to offer an exemplary data curation workflow for collecting data from historical sources that includes not only data curation (whether manual or semiautomatic), but also the practical steps of data cleansing and normalization that reflect the various historical considerations in the process.</p>
N-gram dataset of Chinese local gazetteers (中國地方誌)
<p>This dataset contains the N-grams (1-3) collected from 11083 Chinese local gazetteers (中國地方誌).</p> <p>The dataset comprises of the following resources:</p> <ul> <li><strong>local_gazetteer_1.7z</strong> Unigram dataset in tab separated format (one file per book, each row contains the N-gram and its count)</li> <li><strong>local_gazetteer</strong><strong>_2.7z</strong> Bigram dataset in tab separated format (one file per book, each row contains the N-gram and its count)</li> <li><strong>local_gazetteer</strong><strong>_3.7z</strong> Trigram dataset in tab separated format (one file per book, each row contains the N-gram and its count)</li> <li><strong>local_gazetteer</strong><strong>_metadata.xlsx</strong> Metadata of each book</li> </ul> <p> </p> <p>Dieses Datenset enthält die in 11083 chinesischen Lokalmonographien (中國地方誌) enthaltenen N-Gramme (1-3). </p> <p>Das Datenset besteht aus den folgenden Dateien:</p> <ul> <li><strong>local_gazetteer</strong><strong>_1.7z</strong><em> </em>Monogramm-Datenset im .txt Dateiformat mit Tabstopp als Trennzeichen (jede Datei enthält ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>local_gazetteer</strong><strong>_2.7z</strong><em> </em>Bigramm-Datenset im .txt Dateiformat mit Tabstopp als Trennzeichen (jede Datei enthält ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>local_gazetteer</strong><strong>_3.7z</strong><em> </em>Trigramm-Datenset im .txt Dateiformat mit Tabstopp als Trennzeichen (jede Datei enthält ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>local_gazetteer</strong><em><strong>_</strong></em><strong>metata.xlsx</strong> Metadaten der enthaltenen Bücher</li> </ul> <p> </p> <p>11083 中國地方誌n元語法統計資料 (N-gram Dataset)</p> <p>以下是檔案簡說:</p> <ul> <li><strong>local_gazetteer</strong><strong>_1.7z</strong><em> </em>中國地方誌一元分詞(Unigram)的統計資料 (每本書一個檔案, 以tab作欄區分, 每一行紀錄該N-gram在書中出現的次數)</li> <li><strong>local_gazetteer</strong><strong>_2.7z</strong><em> </em>中國地方誌二元分詞(Bigram)的統計資料 (每本書一個檔案, 以tab作欄區分, 每一行紀錄該N-gram在書中出現的次數)</li> <li><strong>local_gazetteer</strong><strong>_3.7z</strong><em> </em>中國地方誌三元分詞(Trigram)的統計資料 (每本書一個檔案, 以tab作欄區分, 每一行紀錄該N-gram在書中出現的次數)</li> <li><strong>local_gazetteer</strong><em><strong>_</strong></em><strong>metadata.xlsx</strong> 紀錄每本書的基本Metadata</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.