Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

296

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

296 results for “language models”

Learn how ShareScore rates datasets ↗
zenodo40/100

Figure 9. A model that helps the diagnosis prediction-Modern Tools in Patient-Centred Speech Therapy for Romanian Language

<p>We have already implemented many modules of Logo-DM, such as: data cleaning module, data transformation module, feature extraction module, data clustering module and a classification module for diagnosis prediction. &nbsp;Figure 9 shows the model achieved using a decision tree built on complex examination data that aims to predict the patient&rsquo;s diagnosis. Currently, we are testing the built models on new cases in order to estimate their quality.</p>

opencc-by-4.0Jan 2016View details →
zenodo40/100

Figure 12: Language faculty as two-way determination through a universal-Brain Functors: A mathematical model of intentional perception and action

<p>A brain functor, broadly put, is any universal mechanism of determination that can factor determination either way through a universal&ndash;rather than an adjunction that factors one way determination through two (receiving and sending) universals. In some contexts in the life sciences, determination is strictly one way so one might expect to find a semiadjunction but not a two-way system like a brain functor. An application of the scheme for a brain functor in the cognitive sciences is to model the language faculty where there is two way determination between vocal stimuli and internal representations. The previous semiadjunctions for language understanding and language action can be merged to arrive at the brain-like function of the language faculty.</p>

opencc-by-4.0Jan 2016View details →
zenodo40/100

Figure 8: Language production through a sending universal-Brain Functors: A mathematical model of intentional perception and action

<p>The dual to &quot;language understanding&quot; is language production or linguistic action (e.g., &quot;speech acts&quot;). The role of the specific het is played by some auditory output such as utterances (Humboldt&rsquo;s &quot;vocal stimulus&quot;). But the corresponding internal specific hom is the speech act (i.e., internal speech with intentionality) that through the language faculty produces the same outputs but as intentional speech.</p>

opencc-by-4.0Jan 2016View details →
zenodo40/100

BRAIN Journal-Redesigning a Flexible Material Master Data Application with Language Dependency-Figure 3. The logical data model of tables and views

<p>The five tables, named TAB, TABT, AREA, ARET and FLD, are combined within three views (TABV, AREV and FLDV) which build a cluster view, TAFC (figure 3).&nbsp;</p>

opencc-by-4.0Jun 2016View details →
zenodo40/100

Language modeling data for Swahili

<p>The Swahili dataset developed specifically for language modeling task. The dataset contains 28,000 unique words with 6.84M, 970k, and 2M words for the train, valid and test partitions respectively which represent the ratio 80:10:10. The entire dataset is lowercased, has no punctuation marks and, the start and end of sentence markers have been incorporated to facilitate easy tokenization during language modeling. The train partition is the largest in order to support unsupervised learning of word representations while the hyper-parameters are adjusted based on the performance on the valid partition before evaluating the language model on the test partition.</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

DYNA: Disease-Specific Language Model for Variant Pathogenicity

<p>For coding variant effect predictions (VEPs), our approach centers on clinical variant sets specifically related to inherited cardiomyopathies (CM) and arrhythmias (ARM). We utilize a pre-compiled dataset comprised of rare missense pathogenic and benign variants, categorized using a cohort-based approach for diseases such as cardiomyopathy and arrhythmias, as detailed in the previous report by Zhang et al. ClinVar CM and ARM datasets include all missense variants in CM and ARM, respectively, are extracted from ClinVar (Landrum et al.). In the realm of non-coding VEPs, our focus shifts to splicing-related variants, utilizing a dataset from the multiplexed assay for exon recognition by Chong et al., which highlights the significant impact of rare genetic variants on splicing disruptions. Similarly, the ClinVar Splicing dataset, compiled from ClinVar, encompasses all benign sequences and pathogenic variants pertinent to splicing.</p> <p>&nbsp;</p> <p>For the ClinVar CM and ARM datasets, we translate the DNA sequences into protein sequences using the human genome assembly hg38 from https://www.ncbi.nlm.nih.gov/grc/human. We employed the GFF file, MANE.GRCh38.v1.1.ensembl\_genomic.gff.gz from https://www.ncbi.nlm.nih.gov/refseq/MANE, to annotate coding versus non-coding regions for each gene, as only coding DNA sequences are translated into proteins. Additionally, protein domains, cataloged in the Pfam database (Finn et al.), are essential for the functional characterization of proteins. These domains are identified by aligning the translated sequences to known domain structures, thereby facilitating deeper insights into protein function.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Large Language Model-Based Classification of Flash Flood Impacts Across the United States

<p>This repository contains the data sets used for the publication of the journal article titled&nbsp;<em>Large Language Model-Based Classification of Flash Flood Impacts Across the United States</em>.</p> <p>This is the first release of the data with a Zenodo DOI attached to the README.md file.&nbsp;</p> <p>Further information about the data can be found in the GitHub repository's README.md file.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Replication Package for "PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems"

<p>This package contains the traceback data, pre-trained models, and static word embeddings used in the paper, PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Supplemental material for: Software System Testing assisted by Large Language Models: An Exploratory Study

<p>This is the supplemental material of the paper titled as &ldquo;Software System Testing Assisted by Large Language Models: An Exploratory Study&rdquo; presented at the 36th International Conference on Testing Software and Systems.</p> <p>It contains the raw execution data generated by both models, GPT-4o and GPT-4omini, during the exploratory study. The supplementary material includes the following files:</p> <ul> <li><em>GPT-4ominiRQ1-2ExecutionData.zip</em>: contains the JSON outputs from the OpenAI API for the GPT-4o mini model. Each output is labeled according to the research question number and the corresponding timestamp (for RQ1) or the requested test case (for RQ2), all provided in plain text format.</li> <li><em>GPT-4oRQ1-2ExecutionData.zip</em>: contains the JSON outputs from the OpenAI API for the GPT-4o model. Like the previous file, each output is named in plain text format based on the research question number and timestamp (for RQ1) or the requested test case (for RQ2).</li> </ul> <p>To cite this work:&nbsp;</p> <p>C. Augusto, J. Mor&aacute;n, A. Bertolino, C. de la Riva and J. Tuya, &ldquo;S<em>oftware System Testing assisted by Large Language Models: An Exploratory Study</em>&rdquo;, in <em>Testing Software and Systems</em> (pp. 239&ndash;255). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-80889-0_17</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Dataset for "GPTCast: a weather language model for precipitation nowcasting"

<p>Dataset for "<em><strong>GPTCast: a weather language model for precipitation nowcasting</strong></em>"</p> <ul> <li>Preprint:&nbsp;<a href="https://arxiv.org/abs/2407.02089">https://arxiv.org/abs/2407.02089</a></li> <li>Code:&nbsp;<a href="https://github.com/DSIP-FBK/GPTCast">https://github.com/DSIP-FBK/GPTCast</a></li> <li>Pretrained models:&nbsp;<a href="https://doi.org/10.5281/zenodo.13594332">https://doi.org/10.5281/zenodo.13594332</a></li> </ul> <p>Version 2 of this dataset contains also the "Forecaster Test Set" (fts.tar) which includes all generated forecasts for GPTCast8x8, GPTCast16x16, and Linda, to ensure reproducibility of the results.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Transfer learning and DNA language models enhance transcription factor binding predictions

<p>This is the dataset for replicating the results of the paper called "Transfer learning and DNA language models enhance transcription factor binding predictions" by Ekin Deniz Aksu and Martin Vingron.</p> <p>See https://github.com/ekinda/tfbs_prediction_paper</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Protein language model embeddings and predictions of the human proteome

<p>Residue and sequence embeddings of the human proteome (SwissProt for organism Human, downloaded on&nbsp;2021.06.09)&nbsp;computed using bio_embeddings (bioembeddings.com) using the ProtT5 embedder at full precision (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3).</p> <p>Additionally:</p> <p>- Sequence-level&nbsp;predictions of subcellular localization in 10 classes using LA (https://www.biorxiv.org/content/10.1101/2021.04.25.441334v1)</p> <p>- Residue-level three state secondary structure prediction (alpha, sheet or other) using models reported&nbsp;in the ProtTrans paper (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3)</p> <p>&nbsp;</p> <p>Files included:</p> <p>- human.fasta --&gt; FASTA-formatted sequences of human from SwissProt</p> <p>-&nbsp;DSSP3_human_ProtT5Sec.fasta --&gt; Secondary structure predictions in three states for each residue of each protein&nbsp;in human.fasta. &quot;H&quot; stands for Helix; &quot;E&quot; stands for Sheet; &quot;C&quot; stands for Other.</p> <p>-&nbsp;subcell_human_LA_ProtT5.csv --&gt; Subcellular location (10 states) and memrane-boundness (2 states)&nbsp;for each protein in human.fasta</p> <p>-&nbsp;embeddings_file.h5 --&gt; per-residue embeddings of sequences in human.fasta. Each dataset&nbsp;in the .h5 file represents a protein sequence and contains a matrix of length Lx1024, with L being the length of the protein sequence. Datasets are indexed using integers. The original sequence identifier (from the FASTA header) can be accessed through the &quot;original_id&quot; attribute. See&nbsp;https://docs.bioembeddings.com/v0.2.0/notebooks/open_embedding_file.html for information on how to open the file</p> <p>-&nbsp;reduced_embeddings_file.h5 --&gt; per-sequence embeddings of sequences in human.fasta (obtained by mean-pooling the residue-embeddings along the length dimension of the protein sequence). Each dataset&nbsp;in the .h5 file represents a protein sequence and contains a vector of size 1024 (meaning, each sequence has the same dimension).</p>

openafl-3.0Jun 2021View details →
zenodo40/100

TLMD: Tigrinya Language Modeling Dataset

<p>A monolingual dataset built for Tigrinya language modeling. To the best of our knowledge, this is the largest dataset for Tigrinya of its kind. The data was collected from various sources across the web&nbsp;including news, blogs, and books. The largest portion of the data, ~75%, comes from over 2150 issues of the <em>Haddas Ertra</em> newspaper and other magazines published by <a href="https://www.shabait.com">www.shabait.com</a>.</p> <p>Data Statistics:</p> <ul> <li>Total size: ~0.5GB</li> <li>Around 40 million tokens</li> <li>Over 2 million lines</li> <li>367 unique characters</li> <li>Train split: 98%, 1.97 million lines</li> <li>Validation split: 2%, 43k lines</li> </ul> <p>We have done a light-weight cleanup of the data:<br> &nbsp;- Removal of Tigrinya text with legacy and non-standard encoding systems<br> &nbsp;- Normalization of punctuation and special characters<br> &nbsp;- Removal of redundant white spaces and empty lines<br> &nbsp;- Rejoining or fixing broken sentences when possible<br> &nbsp;- Removal of foreign words</p> <p>We avoid applying any form of tokenization, extensive cleanup, and preprocessing operations in order not to take away potentially useful information, those&nbsp;decisions are left to the use-case researchers or developers.</p> <p>This dataset is shared solely&nbsp;to advance&nbsp;research on natural language processing for Tigrinya.&nbsp;While the dataset authors do not claim any copyright on the content, some of the original sources&nbsp;may do. To use the content for commercial purposes or other forms of redistribution of the data, permission&nbsp;shall be acquired from the original owners, mainly shabait.com.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Replication Package: Model-Driven Engineering for the Interoperability of Simulation Modeling Languages: a Case Study in the Space Industry

<p>Replication package &quot;Architectural Support for Software Performance in Continuous Software Engineering: a Systematic Mapping Study&quot;.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

OntoLAMA: LAnguage Model Analysis for Ontology Subsumption Inference

<h3><strong>About</strong></h3> <p>OntoLAMA is a set of language model (LM) probing datasets for ontology subsumption inference. The work follows the "LMs-as-KBs" literature but focuses on conceptualised knowledge extracted from formalised KBs such as the OWL ontologies. Specifically, the subsumption inference (SI) task is introduced and formulated in the Natural Language Inference (NLI) style, where the sub-concept and the super-concept involved in a subsumption axiom are verbalised and fitted into a template to form the premise and hypothesis, respectively. The sampled axioms are verified through ontology reasoning. The SI task is further divided into Atomic SI and Complex SI where the former involves only atomic named concepts and the latter involves both atomic and complex concepts. Real-world ontologies of different scales and domains are used for constructing OntoLAMA and in total there are four Atomic SI datasets and two Complex SI datasets.</p> <p>&nbsp;</p> <table> <tbody> <tr> <th>Dataset Source</th> <th>#Concepts</th> <th>#EquivAxioms</th> <th>#Datasets(Train/Dev/Test)</th> </tr> </tbody> <tbody> <tr> <td>Schema.org</td> <td>894</td> <td>N/A</td> <td> <p>Atomic SI: 808/404/2, 830</p> </td> </tr> <tr> <td>DOID</td> <td>11,157</td> <td>N/A</td> <td> <p>Atomic SI: 90,500/11,312/11,314</p> </td> </tr> <tr> <td>FoodOn</td> <td>30,995</td> <td>2,383</td> <td> <p>Atomic SI: 768,486/96,060/96,062</p> <p>Complex SI: 3,754/1,850/13,080</p> </td> </tr> <tr> <td>GO</td> <td>43,303</td> <td>11,456</td> <td> <p>Atomic SI: 772,870/96,608/96,610</p> <p>Complex SI: 72,318/9,040/9,040</p> </td> </tr> <tr> <td>MNLI</td> <td>N/A</td> <td>N/A</td> <td> <p>biMNLI: 235,622/26,180/12,906</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <h3><strong>Citation</strong></h3> <p>The relevant paper has been accepted at Findings of ACL 2023: <a href="https://aclanthology.org/2023.findings-acl.213/">https://aclanthology.org/2023.findings-acl.213/</a>.</p> <pre>```<br>@inproceedings{he2023language, title={Language Model Analysis for Ontology Subsumption Inference}, author={He, Yuan and Chen, Jiaoyan and Jimenez-Ruiz, Ernesto and Dong, Hang and Horrocks, Ian}, booktitle={Findings of the Association for Computational Linguistics: ACL 2023}, pages={3439--3453}, year={2023} }<br>```</pre> <h3><strong>Links</strong></h3> <ul> <li>See instructions at:&nbsp;<a href="https://krr-oxford.github.io/DeepOnto/ontolama/">https://krr-oxford.github.io/DeepOnto/ontolama/</a></li> <li>We have made available a convenient access of these datasets through Huggingface:&nbsp;<a href="https://huggingface.co/datasets/krr-oxford/OntoLAMA">https://huggingface.co/datasets/krr-oxford/OntoLAMA</a></li> <li>The arxiv version is available at:&nbsp;<a href="https://arxiv.org/abs/2302.06761">https://arxiv.org/abs/2302.06761</a></li> </ul> <h3><strong>Contact</strong></h3> <p>Yuan He (<code>yuan.he(at)cs.ox.ac.uk</code>)</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

ThoughtSource: A central hub for large language model reasoning data (code snapshot)

<p><strong>ThoughtSource is a meta-dataset and software library for chain-of-thought reasoning in large language models (LLMs). This repository contains a snapshot of the associated GitHub repository.</strong></p>

openmit-licenseJul 2023View details →
zenodo40/100

Results and log of LLM-KG-Bench runs described in article "Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering", Meyer et al. 2023

<p>Results and logs of <a href="https://github.com/AKSW/LLM-KG-Bench">LLM-KG-Bench</a> runs described in article &quot;Developing a Scalable Benchmark for Assessing Large Language Models in Knowledge Graph Engineering&quot;, Meyer et al., to appear in <a href="https://2023-eu.semantics.cc/page/accepted_posters">SEMANTICS 2023 poster track</a> proceedings.</p>

opencc-by-4.0Aug 2023View details →
dryad40/100

Data and code on the Moral Machine experiment on large language models (LLMs)

<p>As large language models (LLMs) have become more deeply integrated into various sectors, understanding how they make moral judgments has become crucial, particularly in the realm of autonomous driving. This study utilized the Moral Machine framework to investigate the ethical decision-making tendencies of prominent LLMs, including GPT-3.5, GPT-4, PaLM 2, and Llama 2, to compare their responses to human preferences. While LLMs' and humans' preferences such as prioritizing humans over pets and favoring saving more lives are broadly aligned, PaLM 2 and Llama 2, especially, evidence distinct deviations. Additionally, despite the qualitative similarities between the LLM and human preferences, there are significant quantitative disparities, suggesting that LLMs might lean toward more uncompromising decisions, compared to the milder inclinations of humans. These insights elucidate the ethical frameworks of LLMs and their potential implications for autonomous driving.</p>

opencc-zeroSep 2023View details →
zenodo40/100

Ancient Greek language models

<p>In this repository, we release a series of vector space models of Ancient Greek, trained following different architectures and with different hyperparameter values.&nbsp;</p> <p>Below is a breakdown of all the models released, with an indication of the training method and hyperparameters. The models are split into &lsquo;<strong>Diachronica&rsquo; </strong>and &lsquo;<strong>ALP&rsquo; </strong>models, according to the published paper they are associated with.</p> <blockquote> <p>[<strong>Diachronica</strong>:] Stopponi, Silvia, Nilo Pedrazzini, Saskia Peels-Matthey, Barbara McGillivray &amp; Malvina Nissim. Forthcoming. Natural Language Processing for Ancient Greek: Design, Advantages, and Challenges of Language Models, <em>Diachronica</em>.</p> <p>[<strong>ALP</strong>:] Stopponi, Silvia, Nilo Pedrazzini, Saskia Peels-Matthey, Barbara McGillivray &amp; Malvina Nissim. 2023. Evaluation of Distributional Semantic Models of Ancient Greek: Preliminary Results and a Road Map for Future Work. <em>Proceedings of the Ancient Language Processing Workshop associated with the 14th International Conference on Recent Advances in Natural Language Processing (RANLP 2023)</em>. 49-58. Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-087-8.2023_006</p> </blockquote> <h1><em>Diachronica</em> models</h1> <h2>Training data</h2> <p>Diorisis corpus (Vatri &amp; McGillivray 2018). Separate models were trained for:</p> <ol> <li>Classical subcorpus</li> <li>Hellenistic subcorpus</li> <li>Whole corpus</li> </ol> <p>Models are named according to the (sub)corpus they are trained on (i.e. <code>hel_</code> or <code>hellenestic</code> is appended to the name of the models trained on the Hellenestic subcorpus, <code>clas_</code> or <code>classical</code> for the Classical subcorpus, <code>full_</code> for the whole corpus).</p> <h2>Models</h2> <h3><strong><em>Count-based</em></strong></h3> <blockquote> <p>Software used: LSCDetection (Kaiser et al. 2021;&nbsp;<a href="https://github.com/Garrafao/LSCDetection">https://github.com/Garrafao/LSCDetection</a>)</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; With Positive Pointwise Mutual Information applied (folder PPMI spaces). For each model, a version trained on each subcorpus after removing stopwords is also included (<code>_stopfilt</code> is appended to the model names). Hyperparameter values: <code>window=5</code>, <code>k=1</code>, <code>alpha=0.75</code>.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; With both Positive Pointwise Mutual Information <em>and</em> dimensionality reduction with Singular Value Decomposition applied (folder PPMI+SVD spaces). For each model, a version trained on each subcorpus after removing stopwords is also included (<code>_stopfilt</code> is appended to the model names). Hyperparameter values: <code>window=5</code>, <code>dimensions=300</code>, <code>gamma=0.0</code>.</p> <h3><strong><em>Word2Vec</em></strong></h3> <blockquote> <p>Software used: CADE (Bianchi et al. 2020;&nbsp;<a href="https://github.com/vinid/cade">https://github.com/vinid/cade</a>).</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; Continuous-bag-of-words (CBOW). Hyperparameter values: <code>size=30</code>, <code>siter=5</code>, <code>diter=5</code>, <code>workers=4</code>, <code>sg=0</code>, <code>ns=20</code>.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; Skipgram with Negative Sampling (SGNS). Hyperparameter values: <code>size=30</code>, <code>siter=5</code>, <code>diter=5</code>, <code>workers=4</code>, <code>sg=1</code>, <code>ns=20</code>.</p> <h3><strong><em>Syntactic word embeddings</em></strong></h3> <p>Syntactic word embeddings were also trained on the Ancient Greek subcorpus of the PROIEL treebank (Haug &amp; J&oslash;hndal 2008), the Gorman treebank (Gorman 2020), the PapyGreek treebank (Vierros &amp; Henriksson 2021), the Pedalion treebank (Keersmaekers et al. 2019), and the Ancient Greek Dependency Treebank (Bamman &amp; Crane 2011) largely following the SuperGraph method described in Al-Ghezi &amp; Kurimo (2020) and the Node2Vec architecture (Grover &amp; Leskovec 2016) (see <a href="https://github.com/npedrazzini/ancientgreek-syntactic-embeddings#graph-based-syntactic-word-embeddings">https://github.com/npedrazzini/ancientgreek-syntactic-embeddings</a> for more details). Hyperparameter values: window=1, min_count=1.</p> <h1><em>ALP</em> models</h1> <h2>Training data</h2> <p>Archaic, Classical, and Hellenistic portions of the Diorisis corpus (Vatri &amp; McGillivray 2018) merged, stopwords removed according to the list&nbsp;made by Alessandro Vatri, available at https://figshare.com/articles/dataset/Ancient_Greek_stop_words/9724613.</p> <h2>Models</h2> <h3><strong><em>Count-based</em></strong></h3> <blockquote> <p>Software used: LSCDetection (Kaiser et al. 2021;&nbsp;<a href="https://github.com/Garrafao/LSCDetection">https://github.com/Garrafao/LSCDetection</a>)&nbsp;</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; With Positive Pointwise Mutual Information applied (folder ppmi_alp).&nbsp; Hyperparameter values: <code>window=5</code>, <code>k=1</code>, <code>alpha=0.75</code>. Stopwords were removed from the training set.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; With both Positive Pointwise Mutual Information <em>and</em> dimensionality reduction with Singular Value Decomposition applied (folder ppmi_svd_alp). Hyperparameter values: <code>window=5</code>, <code>dimensions=300</code>, <code>gamma=0.0</code>. Stopwords were removed from the training set.</p> <h3><strong><em>Word2Vec</em></strong></h3> <blockquote> <p>Software used: Gensim library (Řehůřek and Sojka, 2010)</p> </blockquote> <p>a.&nbsp;&nbsp;&nbsp;&nbsp; Continuous-bag-of-words (CBOW). Hyperparameter values: <code>size=30</code>, <code>window=5</code>, <code>min_count=5</code>, <code>negative=20</code>, <code>sg=0</code>. Stopwords were removed from the training set.</p> <p>b.&nbsp;&nbsp;&nbsp;&nbsp; Skipgram with Negative Sampling (SGNS). Hyperparameter values: <code>size=30</code>, <code>window=5</code>, <code>min_count=5</code>, <code>negative=20</code>, <code>sg=1</code>. Stopwords were removed from the training set.</p> <h1><strong>References</strong></h1> <p>Al-Ghezi, Ragheb &amp; Mikko Kurimo. 2020. Graph-based syntactic word embeddings. In Ustalov, Dmitry, Swapna Somasundaran, Alexander Panchenko, Fragkiskos D. Malliaros, Ioana Hulpuș, Peter Jansen &amp; Abhik Jana (eds.), <em>Proceedings of the Graph-based Methods for Natural Language Processing (TextGraphs)</em>, 72-78.</p> <p>Bamman, D. &amp; Gregory Crane. 2011. The Ancient Greek and Latin dependency treebanks. In Sporleder, Caroline, Antal van den Bosch &amp; Kalliopi Zervanou (eds.), <em>Language Technology for Cultural Heritage. Selected Papers from the LaTeCH [Language Technology for Cultural Heritage] Workshop Series. Theory and Applications of Natural Language Processing</em>, 79-98. Berlin, Heidelberg: Springer.</p> <p>Gorman, Vanessa B. 2020. Dependency treebanks of Ancient Greek prose. <em>Journal of Open Humanities Data</em> 6(1).</p> <p>Grover, Aditya &amp; Jure Leskovec. 2016. Node2vec: scalable feature learning for networks. In <em>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD &lsquo;16)</em>, 855-864.</p> <p>Haug, Dag T. T. &amp; Marius L. J&oslash;hndal. 2008. Creating a parallel treebank of the Old Indo-European Bible translations. In <em>Proceedings of the Second Workshop on Language Technology for Cultural Heritage Data (LaTeCH)</em>, 27&ndash;34.</p> <p>Keersmaekers, Alek, Wouter Mercelis, Colin Swaelens &amp; Toon Van Hal. 2019. Creating, enriching and valorizing treebanks of Ancient Greek. In Candito, Marie, Kilian Evang, Stephan Oepen &amp; Djam&eacute; Seddah (eds.), <em>Proceedings of the 18th International Workshop on Treebanks and Linguistic Theories</em> <em>(TLT, SyntaxFest 2019)</em>, 109-117.</p> <p>Kaiser, Jens, Sinan Kurtyigit, Serge Kotchourko &amp; Dominik Schlechtweg. 2021. Effects of Pre- and Post-Processing on type-based Embeddings in Lexical Semantic Change Detection. In <em>Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics</em>.</p> <p>Schlechtweg, Dominik, Anna H&auml;tty, Marco del Tredici &amp; Sabine Schulte im Walde. 2019. A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In <em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</em>, 732-746, Florence, Italy. ACL.</p> <p>Vatri, Alessandro &amp; Barbara McGillivray. 2018. The Diorisis Ancient Greek Corpus: Linguistics and Literature.&nbsp;<em>Research Data Journal for the Humanities and Social Sciences</em>&nbsp;3, 1, 55-65, Available From: Brill&nbsp;<a href="https://doi.org/10.1163/24523666-01000013" target="_blank" rel="noopener">https://doi.org/10.1163/24523666-01000013</a></p> <p>Vierros, Marja &amp; Erik Henriksson. 2021. PapyGreek treebanks: a dataset of linguistically annotated Greek documentary papyri. <em>Journal of Open Humanities Data</em> 7.</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Robustness of large language models in moral judgments

Open the record for dataset details and reuse information.

publicFeb 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record