Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6 results for “N-Gram”

Learn how ShareScore rates datasets ↗
zenodo44/100

Text Generation using N-gram and GPT (metrics: ROUGE, BLEU and BERTScore)

<p>This publication presents a set of spreadsheets listing the user stories generated using N-gram and GPT models with metrics ROUGE, BLEU and BERTScore calculated. Each spreadsheet refers to one corpus of user stories processed.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Evaluation datasets and results of the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing"

<p>Event logs, process models, and results corresponding to the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing".</p> <p><em><strong>Inputs</strong></em>: preprocessed event logs and discovered process models (and their characteristics) used in the evaluation.</p> <ul> <li><em><strong>Real-life</strong></em>: preprocessed event logs (<em>xes</em> and <em>csv</em>) corresponding to the real-life processes used in the evaluation. Process models (<em>pnml</em>) discovered with the Inductive Miner infrequent for thresholds of 10%, 20%, and 50%. Characteristics (<em>txt</em>) of the event logs and process models. Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>).</li> <li><em><strong>Synthetic</strong></em>: simulated&nbsp;event logs (<em>csv</em>) corresponding to the synthetic processes used in the evaluation. Designed process models (<em>bpmn</em> and&nbsp;<em>pnml</em>). Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>). Ongoing cases with injected noise as described in the publication (under folders <em>noise_1</em>, <em>noise_2</em>, and <em>noise_3</em>).</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo40/100

N-gram dataset of Dao Zang Ji Yao (道藏輯要)

<p>This dataset contains the N-grams (1-3) collected from Dao Zang Ji Yao (道藏輯要).</p> <p>The dataset comprises of the following resources:</p> <ul> <li><strong>jiyao_1.7z</strong> Uni-gram dataset in tab seperated format (one file per book,&nbsp;each row contains the N-gram and its count)</li> <li><strong>jiyao_2.7z</strong> Bi-gram dataset in tab seperated format&nbsp;(one file per book, each row contains the N-gram and its count)</li> <li><strong>jiyao_3.7z</strong> Tri-gram dataset in tab seperated format&nbsp;(one file per book, each row contains the N-gram and its count)</li> <li><strong>jiyao_metadata.xlsx</strong> Metadata of each book</li> </ul> <p>&nbsp;</p> <p>Dieses Datenset enth&auml;lt die im Dao Zang Ji Yao (道藏輯要) enthaltenen&nbsp;N-Gramme (1-3).&nbsp;&nbsp;</p> <p>Das Datenset besteht aus den folgenden Dateien:</p> <ul> <li><strong>jiyao_1.7z</strong> Monogramm-Datenset im .txt Dateiformat mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>jiyao_2.7z</strong>&nbsp;Bigramm-Datenset im .txt Dateiformat mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>jiyao_3.7z</strong> Trigramm-Datenset im .txt Dateiformat mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>jiyao_metata.xlsx</strong> Metadaten der enthaltenen B&uuml;cher</li> </ul> <p>&nbsp;</p> <p>《道藏輯要》n元語法統計資料 (N-gram Dataset)</p> <p>以下是檔案簡說:</p> <ul> <li><strong>jiyao_1.7z</strong>&nbsp;《道藏輯要》一元分詞(Unigram)的統計資料 (每本書一個檔案, 以tab作欄區分, 每一行紀錄該 N-gram 在書中出現的次數)</li> <li><strong>jiyao_2.7z</strong>&nbsp;《道藏輯要》二元分詞(Bigram)的統計資料&nbsp;(每本書一個檔案, 以tab作欄區分, 每一行紀錄該 N-gram 在書中出現的次數)</li> <li><strong>jiyao_3.7z</strong> 《道藏輯要》三元分詞(Trigram)的統計資料&nbsp;(每本書一個檔案, 以tab作欄區分, 每一行紀錄該 N-gram 在書中出現的次數)</li> <li><strong>jiyao_metadata.xlsx</strong> 紀錄每本書的基本Metadata</li> </ul>

opencc-by-4.0Mar 2019View details →
zenodo40/100

N-gram dataset of Xu Xiu Si Ku Quan Shu (續修四庫全書)

<p>This dataset contains the N-grams (1-3) collected from Xu Xiu Si Ku Quan Shu (續修四庫全書).</p> <p>The dataset comprises of the following resources:</p> <ul> <li><strong>xuxiu<strong>_</strong>1.7z</strong> Unigram dataset in tab seperated format (one file per book, each row contains the N-gram and its count)</li> <li><strong>xuxiu</strong><strong>_2.7z</strong>&nbsp;Bigram dataset in tab seperated format (one file per book, each row contains the N-gram and its count)</li> <li><strong>xuxiu</strong><strong>_3.7z</strong> Trigram dataset in tab seperated format (one file per book,&nbsp;each row contains the N-gram and its count)</li> <li><strong>xuxiu_metadata.xlsx</strong>&nbsp;Metadata of each book</li> </ul> <p>&nbsp;</p> <p>Dieses Datenset enth&auml;lt die im Xu Xiu Si Ku Quan Shu (續修四庫全書) enthaltenen&nbsp;N-Gramme (1-3).&nbsp;&nbsp;</p> <p>Das Datenset besteht aus den folgenden Dateien:</p> <ul> <li><strong>xuxiu_1.7z</strong>&nbsp;Monogramm-Datenset im .txt Dateiformat mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>xuxiu_2.7z</strong>&nbsp;Bigramm-Datenset im .txt Dateiformat mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>xuxiu_3.7z</strong>&nbsp;Trigramm-Datenset im .txt Dateiformat mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>xuxiu_metata.xlsx</strong>&nbsp;Metadaten der enthaltenen B&uuml;cher</li> </ul> <p>&nbsp;</p> <p>《續修四庫全書》n元語法統計資料&nbsp;(N-gram Dataset)</p> <p>以下是檔案簡說:</p> <ul> <li><strong>xuxiu</strong><strong>_1.7z</strong>&nbsp;《續修四庫全書》一元分詞(Unigram)的統計資料 (每本書一個檔案,以tab作欄區分,每一行紀錄該 N-gram 在書中出現的次數)</li> <li><strong>xuxiu</strong><strong>_2.7z</strong>&nbsp;《續修四庫全書》二元分詞(Bigram)的統計資料 (每本書一個檔案,以tab作欄區分,每一行紀錄該 N-gram 在書中出現的次數)</li> <li><strong>xuxiu</strong><strong>_3.7z</strong>&nbsp;《續修四庫全書》三元分詞(Trigram)的統計資料 (每本書一個檔案,以tab作欄區分,每一行紀錄該 N-gram&nbsp;在書中出現的次數)</li> <li><strong>xuxiu</strong><strong>_metadata.xlsx</strong> 紀錄每本書的基本Metadata</li> </ul>

opencc-by-4.0Mar 2019View details →
zenodo40/100

N-gram dataset of Chinese local gazetteers (中國地方誌)

<p>This dataset contains the N-grams (1-3) collected from 11083 Chinese local gazetteers&nbsp; (中國地方誌).</p> <p>The dataset comprises of the following resources:</p> <ul> <li><strong>local_gazetteer_1.7z</strong> Unigram dataset in tab separated format (one file per book, each row contains the N-gram and its count)</li> <li><strong>local_gazetteer</strong><strong>_2.7z</strong> Bigram dataset in tab separated format&nbsp;(one file per book, each row contains the N-gram and its count)</li> <li><strong>local_gazetteer</strong><strong>_3.7z</strong> Trigram dataset in tab separated format&nbsp;(one file per book, each row contains the N-gram and its count)</li> <li><strong>local_gazetteer</strong><strong>_metadata.xlsx</strong> Metadata of each book</li> </ul> <p>&nbsp;</p> <p>Dieses Datenset enth&auml;lt die in 11083 chinesischen Lokalmonographien (中國地方誌) enthaltenen&nbsp;N-Gramme (1-3).&nbsp;&nbsp;</p> <p>Das Datenset besteht aus den folgenden Dateien:</p> <ul> <li><strong>local_gazetteer</strong><strong>_1.7z</strong><em> </em>Monogramm-Datenset im .txt Dateiformat&nbsp;mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>local_gazetteer</strong><strong>_2.7z</strong><em>&nbsp;</em>Bigramm-Datenset im .txt Dateiformat&nbsp;mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>local_gazetteer</strong><strong>_3.7z</strong><em> </em>Trigramm-Datenset im .txt Dateiformat&nbsp;mit Tabstopp als&nbsp;Trennzeichen (jede Datei enth&auml;lt ein Buch, jede Zeile ein N-Gramm mit der Anzahl der Vorkommnisse im Text)</li> <li><strong>local_gazetteer</strong><em><strong>_</strong></em><strong>metata.xlsx</strong> Metadaten der enthaltenen B&uuml;cher</li> </ul> <p>&nbsp;</p> <p>11083 中國地方誌n元語法統計資料 (N-gram Dataset)</p> <p>以下是檔案簡說:</p> <ul> <li><strong>local_gazetteer</strong><strong>_1.7z</strong><em>&nbsp;</em>中國地方誌一元分詞(Unigram)的統計資料 (每本書一個檔案, 以tab作欄區分, 每一行紀錄該N-gram在書中出現的次數)</li> <li><strong>local_gazetteer</strong><strong>_2.7z</strong><em> </em>中國地方誌二元分詞(Bigram)的統計資料&nbsp;(每本書一個檔案, 以tab作欄區分, 每一行紀錄該N-gram在書中出現的次數)</li> <li><strong>local_gazetteer</strong><strong>_3.7z</strong><em> </em>中國地方誌三元分詞(Trigram)的統計資料 (每本書一個檔案, 以tab作欄區分, 每一行紀錄該N-gram在書中出現的次數)</li> <li><strong>local_gazetteer</strong><em><strong>_</strong></em><strong>metadata.xlsx</strong> 紀錄每本書的基本Metadata</li> </ul>

opencc-by-4.0Mar 2019View details →
zenodo36/100

byteSteady: Fast Classification Using Byte-Level n-Gram Embeddings - Gene Classification Dataset

<p>The gene classification dataset used in the <a href="https://arxiv.org/abs/2106.13302">byteSteady paper</a>.</p>

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record