Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
99
datasets available to search
ShareScore release 0.9.0
Dataset results
99 results for “tokenization”
TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia
<p><strong>Fixes in version 1.1 (= Zenodo's "version 2")</strong></p> <p>*In 20161101-revisions-part1-12-1728.csv, missing first data line is added.</p> <p>*In Current_content and Deleted_content files, some token values ('str' column) which contain regular quotes ('"') are fixed.</p> <p>*In Current_content and Deleted_content files, some wrong revision ID values for 'origin_rev_id', 'in' and 'out' columns are fixed.</p> <p> ------</p> <p><strong>This dataset contains every instance of all tokens (≈ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article revision it was originally created in, and (ii) lists with all the revisions in which the token was ever deleted and (potentially) re-added and re-deleted from its article, enabling a complete and straightforward tracking of its history.</strong></p> <p>This data would be exceedingly hard to create by an average potential user as it is (i) very expensive to compute and as (ii) accurately tracking the history of each token in revisioned documents is a non-trivial task. <br> Adapting a state-of-the-art algorithm, we have produced a dataset that allows for a range of analyses and metrics, already popular in research and going beyond, to be generated on complete-Wikipedia scale; ensuring quality and allowing researchers to forego expensive text-comparison computation, which so far has hindered scalable usage.</p> <p>This dataset, its creation process and use cases are described in a dedicated dataset paper of the same name, published at the ICWSM 2017 conference. In this paper, we show how this data enables, on token level, computation of provenance, measuring survival of content over time, very detailed conflict metrics, and fine-grained interactions of editors like partial reverts, re-additions and other metrics.</p> <p>Tokenization used: https://gist.github.com/faflo/3f5f30b1224c38b1836d63fa05d1ac94</p> <p>Toy example for how the token metadata is generated: <br> https://gist.github.com/faflo/8bd212e81e594676f8d002b175b79de8</p> <p><strong>Be sure to read the ReadMe.txt or - even more detailed - the supporting paper which is referenced under "related identifiers".</strong></p>
SPACCC_TOKEN
<p>[PlanTL/medicine/annotated corpus/guidelines/tokenization] First version of the tokenization annotations in the Spanish Clinical Case Corpus that have been carried out by means of the Spanish Clinical Case Corpus Part-of-Speech Tagger based on FreeLing3.1 (SPACCC_POS-TAGGER, <a href="https://github.com/PlanTL/SPACCC_POS-TAGGER">https://github.com/PlanTL/SPACCC_POS-TAGGER</a>).</p> <p>Copyright (c) 2018 Secretaría de Estado para el Avance Digital</p>
Initial Investors of the BB1 token
<p>In 2019 the Berlin based crypto company <a href="https://www.bitbondsto.com/">bitbond</a> issued the security token BB1 (ITIN: 8WR5-AKBG-X) in Germany’s first, completely regulated STO (security token offering). The BB1 token is structured like a bond of the issuer Bitbond Finance GmbH, with a fixed coupon of 4.00% p.a. (frequency: quarterly), and a floating coupon of 60% of the pre-tax profit of the issuer (frequency: annually).</p> <p>All BB1 related transactions can be observed at the <a href="https://stellar.expert/explorer/public/asset/BB1-GD5J6HLF5666X4AZLTFTXLY46J5SW7EXRKBLEYPJP33S33MXZGV6CWFN">Stellar</a> blockchain. We used this transpareny to extract a list of all investors who initially bought the token. We found 1021 initial investments in total, with a total investment volume of around 2,5 mio BB1 token. The list contains all investments, which have been claimed between July 2019 and June 2020.</p> <p>For each investment we extracted the following data:</p> <ul> <li><strong>Date. </strong>This is the date, at which the investor claimed the purchased tokens.</li> <li><strong>Recipient.</strong> Is the ID of the investros account in the Stellar network. All details about this account can be retrieved by: https://stellar.expert/explorer/public/account/[ID]</li> <li><strong>Symbol. </strong>The symbol of the BB1 token.</li> <li><strong>Volume.</strong> The number of BB1 tokens the investor has intially purchased.</li> </ul> <p>If you have any questions regarding this dataset please contact Lutz Maicher (lutz.maicher@uni-jena.de).</p> <p>There is already a piece at Medium where this data is used: <a href="https://medium.com/@Lutz.Maicher/how-versatile-is-the-crowd-of-sto-investors-the-case-of-the-bb1-token-53a3095925c8">"How versatile is the crowd of STO investors — the case of the BB1 token"</a></p>
Tokens of the noun 'person' in the West Polesian corpus
<p>This dataset contains all the tokens of the noun 'person' in free texts in the West Polesian corpus. Data were collected by Kristian Roncero in the Brest region (Belarus) between Jan 2016 and June 2017. Data are represented according to the IPA conventions, although punctuation marks are used and proper names have their first letter in upper case. I have tried to respect all the differences in the pronunciation, which means sometimes stems appear as palatalized (<em>ʧʲelovjek</em>-, <em>lʲud</em>-, as in Contemporary Standard Russian (CSR)); or most often unpalatalised (which is more in line with the general phonological rules of West Polesian) and the vocalism is not very consistent.</p> <p>The first letters of the code represent the village and speaker. The numbers after the speaker code indicate the file and the remaining the minute and second(s) where this sentence appears.</p>
Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"
<p>Dataset accompanying the paper "<em>Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens</em>"</p> <p>The zip file contains five folders:</p> <p>- <strong>Database:</strong> contains csv files for each speaker which contain the processed features</p> <p>- <strong>Recordings: </strong>the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing</p> <p>- <strong>Recordings_Normalised:</strong> same as recordings but after minimal audio preprocessing (min-max scaling)</p> <p>- <strong>Textgrids: </strong>contains the textgrids which are annotated on the word-level and on phoneme-level</p> <p>- <strong>TIMIT selection: </strong>contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found <a href="https://catalog.ldc.upenn.edu/LDC93s1">here.</a></p>
Influence of social networks as a distribution channel in volatile markets of non-fungible tokens..
<p>Dataset with tweets, prices (ETH & USD), volume (USD) and sentiment from BAYC, WOW and CoolCats NFTs. </p> <p>Obtained from Twitter and from the Ethereum blockchain using Dune Analytics (Dune.xyz)</p>
Deprecated - VAMDC extraction with query token = vald:1e937aca-42f7-403c-8ea3-5a97df96f628:get
<p>Deprecated because the file upload was corrupted</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://vald.astro.uu.se/atoms-12.07/tap/ by doing these queries: (query=select * where ( atomsymbol = 'h' ); ) . The corresponding dataset is versioned on 2018-04-04 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=1ce2c1b9-f4dd-493c-ae3a-a36a9a32cde9. Bibliographic references are: N.N. (2018). {Murphy}, P.~W. (1968). {Transition Probabilities in the Spectra of Ne I, Ar I, and Kr I}. Journal of the Optical Society of America (1917-1983). {Redfors}, A. (1991). {Oscillator strengths for Y III and Zr III in the IUE region}. aap.</p>
Deprecated - VAMDC extraction with query token = vald:bff5bd96-aefe-46b2-85ed-f18fd6f90161:get
<p>Deprecated because the file upload was corrupted.</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://vald.astro.uu.se/atoms-12.07/tap/ by doing these queries: (query=select * where ( IonCharge = '1' ) ) . The corresponding dataset is versioned on 2018-04-04 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=6609d3b9-0ecd-4cda-9693-8aed0f044d7c. Bibliographic references are: N.N. (2018). {Bridges}, J.~M. (1973). {Arc measurements of Fe II oscillator strengths}. {Ekberg}, J.~O. (1997). {Extended analysis of doubly ionized chromium, Cr III}. physscr. 10.1088/0031-8949/56/2/003 {Kurucz}, R.~L. (1975). None. {Murphy}, P.~W. (1968). {Transition Probabilities in the Spectra of Ne I, Ar I, and Kr I}. Journal of the Optical Society of America (1917-1983). {Redfors}, A. (1991). {Oscillator strengths for Y III and Zr III in the IUE region}. aap. {Smith}, G. (1988). {Oscillator strengths for neutral calcium lines of 2.9 eV excitation}. Journal of Physics B Atomic Molecular Physics. 10.1088/0953-4075/21/16/008 {Theodosiou}, C.~E. (1989). Accurate calculation of the 4p lifetimes of Ca^{ + }. Physical Review A. 10.1103/PhysRevA.39.4880</p>
Deprecated - VAMDC extraction with query token = topbase:9eff7c45-0683-4c41-a74c-15c1685464f8:head
<p>Deprecated because the file upload was corrupted.</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://topbase.obspm.fr/12.07/vamdc/tap/ by doing these queries: (query=select * where ( atomsymbol = 'li' ); ) . The corresponding dataset is versioned on 2017-01-30 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=5ab8c38a-bdad-4c38-be72-40cbcf015797. Bibliographic references are: N.N. (2018). Peach, G. and Saraph, H. E. and Seaton, M. J. (1988). Atomic data for opacity calculations. IX. The lithium isoelectronic sequence. Journal of Physics B Atomic Molecular Physics. Fernley, J. A. and Seaton, M. J. and Taylor, K. T. (1987). Atomic data for opacity calculations. VII - Energy levels, f values and photoionisation cross sections for He-like ions. Journal of Physics B Atomic Molecular Physics. Seaton, M. J. (1995). . unpublished.</p>
Deprecated - VAMDC extraction with query token = topbase:f5ca4588-1814-427c-b417-040f7fe09029:get
<p>Deprecated because the file upload was corrupted.</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://topbase.obspm.fr/12.07/vamdc/tap/ by doing these queries: (query=select * where ( IonCharge = '1' ) ) . The corresponding dataset is versioned on 2017-01-30 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=b540b211-ea04-4092-8375-d55fd068f8c7. Bibliographic references are: N.N. (2018). Fernley, J. A. and Hibbert, A. and Kingston, A. E. and Seaton, M. J. (1999). Atomic data for opacity calculations: XXIV. The boron-like sequence. Journal of Physics B Atomic Molecular Physics. Mendoza, C. and Eissner, W. and LeDourneuf, M. and Zeippen, C. J. (1995). Atomic data for opacity calculations. XXIII. The aluminium isoelectronic sequence. Journal of Physics B Atomic Molecular Physics. Hibbert, A. and Scott, M. P. (1994). Atomic data for opacity calculations. XXI. The neon sequence. Journal of Physics B Atomic Molecular Physics. Butler, K. and Mendoza, C. and Zeippen, C. J. (1993). Atomic data for opacity calculations. XIX. The magnesium isoelectronic sequence. Journal of Physics B Atomic Molecular Physics. Tully, John A. and Seaton, Michael J. and Berrington, Keith A. (1990). Atomic data for opacity calculations. XIV - The beryllium sequence. Journal of Physics B Atomic Molecular Physics. Luo, D. and Pradhan, A. K. (1989). Atomic data for opacity calculations. XI - The carbon isoelectronic sequence. Journal of Physics B Atomic Molecular Physics. Luo, D. and Pradhan, A. K. and Saraph, H. E. and Storey, P. J. and Yu, Yan (1989). Atomic data for opacity calculations. X - Oscillator strengths and photoionisation cross sections for O III. Journal of Physics B Atomic Molecular Physics. Peach, G. and Saraph, H. E. and Seaton, M. J. (1988). Atomic data for opacity calculations. IX. The lithium isoelectronic sequence. Journal of Physics B Atomic Molecular Physics. Fernley, J. A. and Seaton, M. J. and Taylor, K. T. (1987). Atomic data for opacity calculations. VII - Energy levels, f values and photoionisation cross sections for He-like ions. Journal of Physics B Atomic Molecular Physics. Seaton, M. J. (1995). . unpublished. Burke, V. M. and Lennon, D. J. (1995). . unpublished. Butler, K. and Zeippen, C.J. (1995). . unpublished. Butler, K. and Zeippen, C.J. (1995). . unpublished. Taylor, K.T. (1995). . unpublished. Butler, K. and Mendoza, C. and Zeippen, C.J. (1995). . unpublished. Storey, P. J. and Taylor, K. T. (1995). . unpublished. Saraph, H. E. and Storey, P. J. (1995). . unpublished.</p>
Deprecated - VAMDC extraction with query token = vald:8402ed6f-b887-4dfb-a5a5-c02c4e3c0ca0:get
<p>Deprecated because the file upload was corrupted.</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://vald.astro.uu.se/atoms-12.07/tap/ by doing these queries: (query=select * where ( atomsymbol = 'he' and radtranswavelength < 600 and radtranswavelength > 500 ); ) . The corresponding dataset is versioned on 2018-04-04 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=955c9268-d40a-4427-9bea-7b715f978769. Bibliographic references are: N.N. (2018). {Kurucz}, R.~L. (1975). None. {Smith}, G. (1988). {Oscillator strengths for neutral calcium lines of 2.9 eV excitation}. Journal of Physics B Atomic Molecular Physics. 10.1088/0953-4075/21/16/008</p>
Deprecated - VAMDC extraction with query token = starkb:64abe979-6ff1-47ef-bb80-fca200635ae8:head
<p>Deprecated because the file upload was corrupted.</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://stark-b.obspm.fr/12.07/vamdc/tap/ by doing these queries: (query=select * where ( atomsymbol = 'mg' and ioncharge = 0 and radtranswavelength <= 5200 and radtranswavelength >= 5100 ); ) . The corresponding dataset is versioned on 2017-06-23 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=60ccb3ad-4411-43d1-badd-5e09ee659f39. Bibliographic references are: N.N. (2018). Dimitrijević M.S. and Sahal-Bréchot S. (1994). Stark broadening parameters tables for Mg I lines of interest for Solar and Stellar spectra research. I. Bull. Astron. Belgrade. Dimitrijević M.S. and Sahal-Bréchot S. (1994). Stark broadening parameters tables for Mg I lines of interest for Solar and Stellar spectra research. II. Bull. Astron. Belgrade. Dimitrijević M.S. and Sahal-Bréchot S. (1996). Stark broadening of Mg I solar lines. A&AS. Dimitrijević M.S. and Sahal-Bréchot S. (1998). Electron-impact broadening of Mg II spectral lines for astrophysical and laboratory plasma research. Phys.Scr.. Dimitrijević M. S. and Sahal-Bréchot S. (1994). Stark Broadening Parameter Tables for MG I Lines of Interest for Solar and Stellar Spectra Research. I. Bull. Obs. Astron. Belgrade.</p>
Deprecated - VAMDC extraction with query token = gecasda:d9e14463-b0d6-43a7-98fc-633a5f9df75d:head
<p>Deprecated because the file upload was corrupted.</p> <p>This dataset comes from the VAMDC(Virtual Atomic and Molecular Data Center) node named http://vamdc.icb.cnrs.fr/gecasda/tap/ by doing these queries: (query=select * where ( inchikey in ['quzpnffhzprkjd-aklpvkdbsa-n', 'quzpnffhzprkjd-bjudxgsmsa-n', 'quzpnffhzprkjd-igmarmgpsa-n', 'quzpnffhzprkjd-oiobtwansa-n', 'quzpnffhzprkjd-oubtzvsysa-n'] ); ) . The corresponding dataset is versioned on 1 with the XSAMS 12.07 format. This data is associated with a unique identifier uuid=5c91e7c7-e0c9-474f-b76a-62bff7b38468. Bibliographic references are: N.N. (2018). Boudon,V. and Grigoryan,T. and Philipot,F and Richard,C and Kwabia Tchana,F and Manceron,L. and Rizopoulos,A and Vander Auwera,J and Encrenaz,T (2018). Line positions and intensities for the $ u_3$ band of 5 isotopologues of germane for planetary applications. Journal of Quantitative Spectroscopy and Radiative Transfer.</p>
Deprecated - VAMDC extraction with query token = mecasda:86283556-1a6a-4fd3-b513-de792c32edd4:get
<p>Deprecated because the file upload was corrupted.</p> <p>This is a dataset extracted from http://vamdc.icb.cnrs.fr/mecasda-12.07/tap/ VAMDC node.</p> <p>Query originating this dataset: query=select species;</p> <p>Data source version: 1</p> <p>Data format: XSAMS 12.07</p> <p>Query uuid in VAMDC query store: 9ca8e9a9-b6d7-4f5e-a510-13cb8bf4b759</p>
Deprecated - VAMDC extraction with query token = starkb:3dec15c4-fc5b-4679-aa6a-3244d77436eb:head
<p>Deprecated because the file upload was corrupted.</p> <p>This is a dataset extracted from http://stark-b.obspm.fr/12.07/vamdc/tap/ VAMDC node.</p> <p>Query originating this dataset: query=select * where ( atomsymbol = 'li' and ioncharge = 0 );</p> <p>Data source version: 2017-06-23</p> <p>Data format: XSAMS 12.07</p> <p>Query uuid in VAMDC query store: 17053a9a-e56e-451b-9bd2-8e0cddda0d5d</p>
[Artifact] Qubit Allocation as a Combination of Subgraph Isomorphism and Token Swapping
<p>Object-Oriented Programming, Systems, Languages & Applications (OOPSLA'19) artifact for the paper: "Qubit Allocation as a Combination of Subgraph Isomorphism and Token Swapping".</p>
Questionnaire on adnumerative forms in West Polesian and tokens from the corpus
<p>The following document contains a scratchy draft of the answers given by West Polesian speakers concerning numeral phrases and The data were gathered by Kristian Roncero in the region of Brest (Belarus) between 2016-2017 as a part of his PhD project.</p> <p>The first part of the questionnaire is inspired by Shevelov (1963). The second part contains several paradigms and the different answers multiple speakers have given for them. The first half is based on answers given to visual stimuli; whilst the second one are based on direct elicitation for primarily masculine human (virile) nouns, as I suspected they had a special distribution of the adnumerative.The last part contains some hits of numeral phrases with lower numerals in West Polesian from free texts (note that many data are missing here).</p>
Data from: More than a token photo: humanising scientists enhances student engagement
Open the record for dataset details and reuse information.
LatinISE test data for SemEval 2020 task 1 with additional token versions of the corpora
<p>This data collection contains the Latin test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>: </p> <ul> <li>a Latin text corpus pair (`corpus1/lemma`, `corpus2/lemma`)</li> <li>40 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)</li> <li>the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</li> </ul> <p>The corpus data have been automatically lemmatized and part-of-speech tagged, and have been partially corrected by hand. For homonyms, the lemmas are followed by the '\#' symbol and the number of the homonym according to the Lewis-Short dictionary of Latin when this number is greater than 1. For example, the lemma 'dico' corresponds to the first homonym in the Lewis-Short dictionary and 'dico\#2' corresponds to the second homonym, cf. Lewis-Short dictionary.</p> <p>__Corpus 1__</p> <ul> <li>based on: <a href="http://hdl.handle.net/11372/LRT-3170">LatinISE</a> (McGillivray and Kilgarriff 2013), <a href="https://app.sketchengine.eu/#dashboard?corpname=preloaded/latinise_4">version on Sketch Engine</a></li> <li>language: Latin</li> <li>time covered: from the beginning of the second century before Christ (BC) to the end of the first century BC</li> <li>size: ~1.7 million tokens</li> <li>format: lemmatized, sentence length >= 2, no punctuation, sentences randomly shuffled</li> <li>encoding: UTF-8</li> </ul> <p>__Corpus 2__</p> <ul> <li>based on: <a href="http://hdl.handle.net/11372/LRT-3170">LatinISE</a> (McGillivray and Kilgarriff 2013) , <a href="https://app.sketchengine.eu/#dashboard?corpname=preloaded/latinise_4">version on Sketch Engine</a></li> <li>language: Latin</li> <li>time covered: from the beginning of the first century after Christ (AD) to the end of the twenty-first century AD</li> <li>size: ~9.4 million tokens</li> <li>format: lemmatized, sentence length >= 2, no punctuation, sentences randomly shuffled</li> <li>encoding: UTF-8</li> </ul> <p>Find more information on the data in the papers referenced below.</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (<code>corpus1/token/</code>, <code>corpus2/token/</code>). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p>The creation of the data was supported by the CRETA center and the CLARIN-D grant funded by the German Ministry for Education and Research (BMBF).</p> <p><strong>References</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p> <p>McGillivray, B. and Kilgarriff, A. (2013). <a href="https://www.sketchengine.co.uk/wp-content/uploads/2015/05/Latin_historical_corpus_2013.pdf">Tools for historical corpus research, and a corpus of Latin</a>. In Paul Bennett, Martin Durrell, Silke Scheible, Richard J. Whitt (eds.), New Methods in Historical Corpus Linguistics, Tübingen: Narr.<br> </p>
Token-based data sets for the analysis of the academic language of literary studies and linguistics
<p>These are the token-based data sets used for my PhD thesis ("Potentiale syntaktischer Annotationen für die datengeleitete Sprachbeschreibung am Beispiel der Wissenschaftssprachen der Germanistik", publication in progress).</p> <p>For python scripts and further data see https://github.com/melandresen/dissertation.</p> <p>Due to copyright law, the annotated texts of the corpus could only be published without the token layer. The files provided here include the token-based frequency data that have been derived from the origial texts and can be used as input to the analysis scripts in the GitHub-Repository.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.