Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

57

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

57 results for “bilingualism”

Learn how ShareScore rates datasets ↗
zenodo40/100

Bilingual English-German word embedding models for scientific text

<p>This data set contains three word embedding models, constructed from the same training corpus of English and German parallel scientific texts (abstracts and research project descriptions). All text was pre-processed by language-specific stemming with the Porter stemming algorithm, removing numbers, and lower-casing.</p> <p>The first model is a 1000-dimensional Latent Semantic Analysis model, constructed from concatenating the English and German texts. The input data was a&nbsp;m&times;n (297,852&times;923,864) document-term matrix of tf-idf weights. This was processed with truncated SVD. There are two files, the word vectors in file lsa_1000_Vmat.csv (the V* term by latent factors matrix of right singular values) and the dimension weights in lsa_1000_d_weights.csv (the 1000 values of the diagonal of the <span class="math-tex">\(\Sigma\)</span> matrix.</p> <p>lsa_1000_Vmat.csv has two fields, the term and its vector representation in LSA space, separated by a &quot;|&quot; character. The structure looks like this:</p> <p>tarifplural|{5.00599733151825e-08,-1.43071379136936e-08,8.32862290483082e-08,-6.08010721687266e-08,1.15831140150142e-07,-2.46470313387358e-08,3.43215595753282e-07,6.24301666802575e-07,-2.62907158945831e-07,-1.04120313981517e-07,4.5864574355164e-07,-2.31799632277312e-07,8.37354377858843e-07,8.22507467711628e-07,4.07585381069368e-07,-4.26358988941922e-08,-8.38652991154651e-07,1.98091851171759e-07,-3.94768548759816e-08,-4.28802181962385e-07, ...}</p> <p>The other two models are a basic Random Indexing and a Reflective Random Indexing model, contained in same file, RI_training.csv. Both models have 1000 dimensions. The data structure is as follows.</p> <ul> <li>language: either &quot;en&quot; (English) or &quot;de&quot; (German), the language of the term</li> <li>term: the term as a character string</li> <li>term_collection_count: integer, number of times the term occurred in the training data</li> <li>c_vector: vector of 1000 reals, RI context vector of the term. formatted like this: &quot;{0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.12309149,0,0,-0.12309149,0,0,0,0,0,0,0,0,0,0,0,0,0,0, ...}&quot;</li> <li>n_docs: integer, number of different documents which contained the term</li> <li>c_vector_o2: vector of 1000 reals, RRI context vector of the term, formatted like c_vector above</li> </ul> <p>1,034,860 rows.</p> <p>All files are aggressively compressed with GNU gzip and will require much more disk space when uncompressed.&nbsp; Note the special formatting of the vector numeric variables, which are different for the two models.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Documentation of a bilingual Pyu inscription (PYU011) held at the the Pagan museum

<p>This data set includes photographs (.jpg), RTIs (.ptm) and related files documenting a bilingual Chinese and Pyu inscription (inventory number PYU011) held at the Pagan museum. The photographer was James Miles or Archeovision, working on behalf of the Pyu epigraphy sub-project (PI, Nathan W. Hill of SOAS University of London) of the ERC synergy grant "Beyond Boundaries: Religion, Region, Language and the State" (Identifier: ASIA 609823) in collaboration with the project "From Vijayapuri to Sriksetra? The Beginnings of Buddhist Exchange across the Bay of Bengal as Witnessed by Inscriptions from Andhra Pradesh and Myanmar" (PI Arlo Giffiths of the EFEO) funded by The Robert H. N. Ho Family Foundation.</p>

opencc-by-4.0Nov 2016View details →
zenodo40/100

Documentation of a Sanskrit-Pyu bilingual inscription (PYU012) around the base of a Buddha statue held by the Śrī Kṣetra museum

<p>This data set includes includes photographs (.jpg), RTIs (.ptm) and related files documenting a Sanskrit-Pyu bilingual inscription (inventory number PYU012)  around the base of a headless Buddha statue held by the Śrī Kṣetra museum. The photographer was James Miles or Archeovision, working on behalf of the Pyu epigraphy sub-project (PI, Nathan W. Hill of SOAS University of London) of the ERC synergy grant "Beyond Boundaries: Religion, Region, Language and the State" (Identifier: ASIA 609823) in collaboration with the project "From Vijayapuri to Sriksetra? The Beginnings of Buddhist Exchange across the Bay of Bengal as Witnessed by Inscriptions from Andhra Pradesh and Myanmar" (PI Arlo Giffiths of the EFEO) funded by The Robert H. N. Ho Family Foundation.</p>

opencc-by-4.0Nov 2016View details →
zenodo40/100

Bilingual Dataset of Multinational Brand Posts and Consumer Replies from Platform X (Turkish and English)

<div> <div> <div> <div>&nbsp;</div> </div> </div> </div> <div> <p>This dataset consists of posts and consumer replies from a curated selection of Fortune 500 companies from 2021 that actively maintained Platform X accounts in both English and Turkish. The sample includes 33 multinational brands spanning various industries: fast-moving consumer goods (FMCG), fast food, technology, automotive, apparel, retail, finance, and logistics.</p> <p>All Platform X messages posted by these brands, along with corresponding consumer responses, were collected over a five-year period, from June 2016 to June 2021. The dataset initially contained a total of 8,101,034 English and 735,459 Turkish messages. Among these, the English dataset includes 2,708,874 brand posts and 5,392,160 consumer replies, while the Turkish dataset comprises 245,422 brand posts and 490,037 consumer replies.</p> <p>All files are in GZIP compressed parquet format.</p> </div>

restrictedcc-by-4.0Nov 2024View details →
zenodo40/100

Machine-readable Northern Karelian Proper-Livvi bilingual translation dictionary

<p>This machine readable bilingual translation dictionary of Northern Karelian Proper (ISO-639: krl) to Livvi aka Olonets-Karelian (ISO-639 olo) was created by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd&#39; with translation suggestions generated by Khalid Alnajjar and Mika H&auml;m&auml;l&auml;inen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Machine-readable Finnish-Livvi bilingual translation dictionary

<p>This machine readable bilingual translation dictionary of Finnish to Livvi aka Olonets-Karelian (ISO-396: olo) was proofread and extended by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd&#39; with translation suggestions generated by Khalid Alnajjar and Mika H&auml;m&auml;l&auml;inen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Machine-readable Finnish-Karelian bilingual translation dictionary

<p>This machine readable bilingual translation dictionary of Finnish to Northern Karelian Proper (ISO-396: krl) was proofread and extended by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd&#39; with translation suggestions generated by Khalid Alnajjar and Mika H&auml;m&auml;l&auml;inen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Main Study Data)

<p>Datasets from Main&nbsp;study of&nbsp;<strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong>&nbsp;(project linked here:&nbsp;https://doi.org/10.17605/OSF.IO/S6VDN).</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Pilot Study Data)

<p>Datasets from Pilot&nbsp;study of&nbsp;<strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong>&nbsp;(project linked here:&nbsp;https://doi.org/10.17605/OSF.IO/S6VDN).</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Word-related parameters and discriminative power of NWRT for bilingual children (Italian-German) with/without DLD

<p>Dataset with word-related parameters, adult monolingual ratings&nbsp;and mean repetition scores by bilingual children with/without (risk of) DLD. Word-related parameters include length, number of syllables, number of clusters,&nbsp; monolingual adults&#39; direct (in-person) repetition rates and direct ratings of Pronounceability (PR), Word-likeness (WLike); online ratings of Pronounceability (avPR), Specificity (avSP, corresponding to&nbsp;percent target language assignment), and children&#39;s repetition rates (averages from children with/without DLD).&nbsp;IT, Italian, GER, German.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Bilingual Dataset for Information Retrieval and Question Answering over the Spanish Workers Statute

<p>A bilingual dataset of questions and answers over a key document in Spanish labor law legislation is presented. The document contains 150 questions and their respective answers in the form of one part number from the 130 parts in which the Workers Statute is divided (articles and other provisions), and with the most relevant excerpt of information for the answer.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Theories and Principles of Bilingual Education

<p>This annotated bibliography discusses chapters from three books (<em>Foundations of Bilingual Education and Bilingualism,&nbsp;Immersion Education: International Perspectives,&nbsp;</em>and&nbsp;<em>Literacy and Bilingualism: A Handbook for All Teachers)</em>&nbsp;and an article entitled&nbsp;<em>Biliteracy, Empowerment, and Transformative Pedagogy.</em> The&nbsp;academic works discussed are mainly&nbsp;about theories and principles of bilingualism, bilingual education, and immersion education&nbsp;which relate&nbsp;to second/foreign language teaching in different contexts.</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Vikidia En/Fr bilingual dataset for Automatic Readability Assessment

<p>Vikidia.org is a children&#39;s encyclopedia, with content targeting 8-13 year old children, in several European languages. Our dataset contains 24660 texts distributed across 6165 articles in 2 reading levels, for English and French respectively i.e., each text in the corpus has four versions: en, en-simple, fr and fr-simple, and there are 6165 slugs in total.&nbsp;The uniqueness of the current dataset is that these are parallel, document level aligned texts in four versions - en, en-simple, fr, fr-simple. While we did not create paragraph/sentence level alignments on the corpus, we hope that this will be a useful dataset for future English and French research on ARA and Automatic Text Simplification. This is the first such dataset in ARA, and perhaps the first readily available French readability dataset.</p> <p>This dataset is used in the paper &quot;A neural pairwise ranking model for automatic readability assessment&quot; by Justin Lee and Sowmya Vajjala, to appear in Findings of ACL 2022.&nbsp;</p>

opencc-by-3.0Mar 2022View details →
zenodo36/100

Annotated Dataset for Bilingual Code-Mixed English-Malay Sentiment Analysis and Sarcasm Detection in Public Security Domain

<p>Tweets from X, and post with comment from TikTok was acquired <span>from 11 September until 21 September 2022</span>. Data from both platforms was merged and selected. Three annotators manually labelling the selected data for sentiment and sarcasm. Sentiment labels are &lsquo;positive&rsquo;, &lsquo;negative&rsquo;, and &lsquo;neutral&rsquo;. Sarcasm label is &lsquo;sarcastic&rsquo; and &lsquo;not sarcastic&rsquo;. Majority voting is considered for each label. Language identification label produced for each data.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

juliacarbajal/bilingual_assimilation: First release of data & analysis scripts for bilingual assimilation project.

<p>This release contains all the app data and vocabulary questionnaires collected for Carbajal et al. (in preparation)&#39;s project on bilingual assimilation (OSF project <a href="https://osf.io/52z9g/">https://osf.io/52z9g/</a>). It also includes analysis scripts as of the 22nd of March, 2018.</p>

opencc-by-nc-sa-4.0Dec 2017View details →
zenodo36/100

NeuroVox: Bilingual Brain-to-Speech Translation and Neural Activity Detection

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
ClinicalTrials.gov36/100

Auditory Processing in Spanish-English Bilinguals: Is Performance Better When Tested in Spanish or English?

ClinicalTrials.gov study NCT05452486. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad36/100

Is mere exposure enough? the effects of bilingual environments on infant cognitive development

Open the record for dataset details and reuse information.

publicFeb 2020View details →
zenodo32/100

Bilingual inscription, Tomb FA5, Mleiha, Sharjah

Bilingual South-Arabian / Aramaic funerary inscription discovered inside the burial chamber of tomb FA-5, Mleiha, Sharjah, UAE. (Overlaet 2018:14 Figure 9, Catalog No. 1). Central panel in south Arabian script, text along the rim in Aramaic. Text reads "Memorial and tomb of Amud son of Gurr son of Ali, inspector of the king of Oman, which built over him his son Amud son of Amud son if Gurr, inspector of the king of Oman." Dated to either 222/221 or 215/214 BCE. First record evidence of an Oman Kingdom (Overlaet 2018:13). Model of the original after stabilization. 358 photos. Completely processed (aligned, scaled, modeled, cleaned, simplified, unwrapped, textured, meshed) in Reality Capture. B. Overlaet 2018. Mleiha, An Arab Kingdom on the Caravan Trails (Brussels 30.10-30.12.2018). Sharjah: Sharjah Archaeology Authority. Source: Objaverse 1.0 / Sketchfab

opencc-by-nc-1.0Mar 2019View details →
zenodo32/100

Figs. 1–5. Hologymnetis reyesi, male. 1 in Description of the Female ofHologymnetis reyesiGasca and Deloya (Coleoptera: Scarabaeidae: Cetoniinae: Gymnetini), with New State Records for Mexico and a Bilingual Key to the Species ofHologymnetisMartínez

Figs. 1–5. Hologymnetis reyesi, male. 1) Dorsal habitus; 2) Ventral habitus; Parameres: 3) Dorsal view; 4) Ventral view; 5) Lateral view. Photographs by Jiri Zidek.

opennotspecifiedMar 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record