Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
99
datasets available to search
ShareScore release 0.9.0
Dataset results
99 results for “tokenization”
A Multi-Objective Optimization Framework for Code Generation Decoding Strategies Based on Code Token Trees
Open the record for dataset details and reuse information.
Pocket-based generation task data for Token-Mol 1.0 model finetuning
Open the record for dataset details and reuse information.
Data from: Classification of cryptocurrency coins and tokens by the dynamics of their market capitalisations
Open the record for dataset details and reuse information.
Tokenized and POS-Tagged Khmer Data of the Asian Language Treebank Project
<p>* Introduction</p> <p>This is the Khmer ALT of the Asian Language Treebank (ALT) Corpus. English texts sampled from English Wikinews were available under a Creative Commons Attribution 2.5 License.</p> <p>Please refer to<br> http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/index.html<br> for an introduction of the ALT project.</p> <p>Khmer ALT has been developed by NICT and NIPTICT. The license of Khmer ALT is</p> <p>Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License<br> https://creativecommons.org/licenses/by-nc-sa/4.0/</p> <p><br> * Contents</p> <p>- data_km.km-[tok|tag].nova : tokenized/POS-tagged Khmer sentences by the nova annotation system<br> # based on the following two guildelines<br> # http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/Khmer-annotation-guideline.pdf<br> # http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/Khmer-annotation-guideline-supplementary.pdf</p> <p><br> * Disclaimer</p> <p>[1] The content of the selected English Wikinews articles have been translated for this corpus. English texts sampled from English Wikinews were available under a Creative Commons Attribution 2.5 License. Users of the corpus are requested to take careful consideration when encountering any instances of defamation, discriminatory terms, or personal information that might be found within the corpus. Users of the corpus are advised to read Terms of Use in https://en.wikinews.org/wiki/Main_Page carefully to ensure proper usage.</p> <p>[2] NICT bears no responsibility for the contents of the corpus and the lexicon and assumes no liability for any direct or indirect damage or loss whatsoever that may be incurred as a result of using the corpus or the lexicon.</p> <p>[3] If any copyright infringement or other problems are found in the corpus or the lexicon, please contact us at alt-info[at]khn[dot]nict[dot]go[dot]jp. We will review the issue and undertake appropriate measures when needed.</p>
us-banks-10K-sec7-tokenized
<p>data_2436.csv contains tokenized words for 10-K sec7/7A text for US banks</p> <p>new_all_2436_mda_roa.csv contains all original sec7/7A</p>
Miner's token
Miners carried a number into the mine each day. As each tub was filled at the face by a miner, or more usually a pair of men working together, a token was hung on the tub so that when it was weighed at surface the amount could be ascribed to that particular number. When the men were paid, at the end of the week, or fortnight in earlier days, the amount per ton accrued to each number was paid out to that number. A pair of miners then divided the money between themselves. In the days of hand jumper drilling an experienced miner could expect to produce between 4 and 5 tons per shift for which rates could vary from one shilling to perhaps 1/10d per ton. 6 shifts producing, say, 27 tons, at 1/6d per ton produces £2 0s 6d (around £250 in today's money) from which would be deducted costs of gunpowder, sick club, union subs etc. In later and more recent years additional tokens were introduced to keep track of who was underground at any one time. Source: Objaverse 1.0 / Sketchfab
Beard Tax Token
A Coin That Stopped the Tsar's Police From Shaving You<br> "Money Taken" tokens, issued in 1699 during the reign of Peter the Great. All who wanted to keep their beard in St. Petersburg, had to pay tax, receiving such a token as proof that tax was paid. Heres a more colorful article about these tokens - https://www.atlasobscura.com/articles/what-was-the-beard-tax?utm_source=Atlas+Obscura+Daily+Newsletter&utm_campaign=43a8d52d3e-EMAIL_CAMPAIGN_2018_09_11_COPY_01&utm_medium=email&utm_term=0_f36db9c480-43a8d52d3e-66186853&ct=t(EMAIL_CAMPAIGN_9_11_2018_COPY_01)&mc_cid=43a8d52d3e&mc_eid=c396b12c99 Original Image from http://numistika.com/vkgm%20collection.html Source: Objaverse 1.0 / Sketchfab
Newport Bridge Token 20200620
Rhode Island Turnpike and Bridge Authority token. Date, Unknown. Composition, Brass. Weight, 7.6 g. Diameter, 28mm. Mint, Roger Williams Mint, Attleboro, Massachusetts, United States https://en.numista.com/catalogue/pieces127218.html From personal collection. Photogrammetric data created using focus stacking and manual turntable. Source: Objaverse 1.0 / Sketchfab
World of Warcraft Token
Model of WoW Token Source: Objaverse 1.0 / Sketchfab
Yggdrasil token
N/A Source: Objaverse 1.0 / Sketchfab
【outputs_search_best_number_tokens_large_scales_SS】only two 8B LLMs, only SS task, [1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files
<ul> <li>only two 8B LLMs, </li> <li>only SS task,</li> <li>[1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files</li> </ul>
【outputs_search_best_number_tokens_large_scales_Ir】only two 8B LLMs, only Ir task, [1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files
<ul> <li>only two 8B LLMs, </li> <li>only Ir task,</li> <li>[1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files</li> </ul>
【outputs_search_best_number_tokens_large_scales_HT】only two 8B LLMs, only HT task, [1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files
<ul> <li>only two 8B LLMs, </li> <li>only HT task,</li> <li>[1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files</li> </ul>
【outputs_search_best_number_tokens_large_scales_FB】only two 8B LLMs, only FB task, [1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files
<ul> <li>only two 8B LLMs, </li> <li>only FB task,</li> <li>[1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files</li> </ul>
【outputs_search_best_number_tokens_large_scales_SS 012】only two 8B LLMs, only SS 012 task, [1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files
<ul> <li>only two 8B LLMs, </li> <li>only SS 012 task,</li> <li>[1 (best # of tokens) + 1 (corresponding language materials)] ⨉ 6 (parts) = 12 files</li> </ul>
【outputs_best_number_tokens_to_tpm_large_scales】/best_number_tokens_log_part_12 (2), only two 8B LLMs, 30 (tpms) = 30 files
<ul> <li>Part 12 (2) </li> <li>only two 8B LLMs, </li> <li>only valid results (both MP p value and CI),</li> <li>all permutation controls,</li> <li>all seeds,</li> <li>all linguistic spans,</li> <li>all 5 tasks, </li> <li>30 (tpms) = 30 files</li> </ul>
【outputs_best_number_tokens_to_tpm_large_scales】/best_number_tokens_log_part_12 (1), only two 8B LLMs, 30 (tpms) = 30 files
<ul> <li>Part 12 (1)</li> <li>only two 8B LLMs, </li> <li>only valid results (both MP p value and CI),</li> <li>all permutation controls,</li> <li>all seeds,</li> <li>all linguistic spans,</li> <li>all 5 tasks, </li> <li>30 (tpms) = 30 files</li> </ul>
【outputs_best_number_tokens_to_tpm_large_scales】/best_number_tokens_log_part_18 (2), only two 8B LLMs, 5 (tpms) = 5 files
<ul> <li>Part 18 (2) only two 8B LLMs, </li> <li>only valid results (both MP p value and CI),</li> <li>all permutation controls,</li> <li>all seeds,</li> <li>all linguistic spans,</li> <li>all 5 tasks, </li> <li>5 (tpms) = 5 files</li> </ul>
【outputs_best_number_tokens_to_tpm_large_scales】/best_number_tokens_log_part_18 (1), only two 8B LLMs, 25 (tpms) = 25 files
<ul> <li>Part 18 (1) </li> <li>only two 8B LLMs, </li> <li>only valid results (both MP p value and CI),</li> <li>all permutation controls,</li> <li>all seeds,</li> <li>all linguistic spans,</li> <li>all 5 tasks, </li> <li>25 (tpms) = 25 files</li> </ul>
【outputs_best_number_tokens_to_tpm_large_scales】/best_number_tokens_log_part_18 (4), only two 8B LLMs, 8 (tpms) = 8 files
<ul> <li>Part 18 (4) </li> <li>only two 8B LLMs, </li> <li>only valid results (both MP p value and CI),</li> <li>all permutation controls,</li> <li>all seeds,</li> <li>all linguistic spans,</li> <li>all 5 tasks, </li> <li>8 (tpms) = 8 files</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.