Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
95
datasets available to search
ShareScore release 0.9.0
Dataset results
95 results for “Sentence”
Figure 5. The mental spaces set up by the second interpretation of the sentence in the film, John is riding a unicorn.-Representing Mental Spaces and Dynamics of Natural Language Semantics
<p>Second interpretation: The proper name John exists only in film space without having a<br> counterpart in base space. Therefore, the second interpretation of the sentence is: John is a<br> film character who is riding a unicorn in the film (Figure 5).</p>
Figure 4. The mental spaces set up by the first interpretation of the sentence in the film, John is riding a unicorn-Representing Mental Spaces and Dynamics of Natural Language Semantics
<p>First interpretation: The reality space (let’s call it R) contains an element a associated with<br> the proper name John. The noun phrase a unicorn introduces an element b ׳ to the film space<br> (call it F). I is the connector linking a in the space B to a ׳ in the space F (Figure 4). Since the<br> elements of both mental spaces are co-referential, this connector is an identity connector.<br> The rectangles represent the internal structure of the spaces next to them. The dashed line<br> indicates that the space F is set up in relation to R and that it is subordinate to R in<br> discourse.</p>
Figure 2. Semantic representation of the sentence Peter finished the discussion in MultiNet after Helbig [8, p. 447].-Representing Mental Spaces and Dynamics of Natural Language Semantics
<p>In Figure 3, semantic frame of the concept Finish realized in the form of the verb finish<br> requires two C-roles: An agent represented by the relation AGT, and an affected entity represented<br> by the relation AFF. Here agent is Peter and the affected entity is an abstract object ([SORT = ad]<br> means the concept is a dynamic abstraction).</p>
Sentence Compression Using Constituency Analysis of Sentence Structure
<p>Simply stated, producing a shorter format of a given sentence is the task of sentence compression. The challenging part of this process is preserving the most important information as well as grammaticality in the compressed version. There are different ways of achieving this purpose, among which, we try to come up with a rule-based extractive method for the Persian language. Our approach involves identifying removable constituents within a sentence and eliminating them in the order of insignificance until the desired compression rate is achieved. To develop this method, we created a compression corpus of 600 sentences from two available treebanks in the Persian language. 300 sentences are used for extracting deletion rules as training data and the remaining 300 are used for testing the system. Its applicability even using a limited training corpus and user’s authority over the input compression rate are the benefits of the presented method in this paper. Additionally, our compression system is open-ended, enabling the addition or removal of deletion rules to refine its performance. The results suggest applying rule-based methods for language processing tasks can be quite efficient in syntactically rich languages such as Persian, yielding desirable outcomes and offering distinct advantages.</p>
Molecular Biology Open Access Pubmed Word and Sentence Representations
<p><strong>Natural Language Embeddings about Molecular Biology</strong></p> <p>This dataset is concerned with developing a tailored training data set for word and sentence embedding based on biomedical text that has some component associated with molecular work (as opposed to the other range of work indexed in PubMed like non molecular clinical work, studies of human behavior, etc). </p> <p><strong>Raw Data</strong></p> <p>In order to develop natural language embeddings (for words and sentences), we queried PMC and MEDLINE for molecular papers only by using high-level MeSH terms to restrict interest to papers with a molecular focus. We used the following MeSH terms:</p> <ul> <li>Cells [A11]</li> <li>Multiprotein Complexes [D05.500]</li> <li>Protein Aggregates [D05.875]</li> <li>Hormones [D06]</li> <li>Enzymes and Coenzymes [D08]</li> <li>Carbohydrates [D08]</li> <li>Lipids [D10]</li> <li>Amino Acids, Peptides and Proteins [D12]</li> <li>Nucleic Acids, Nucleotides and Nucleosides [D13]</li> <li>Biological Factors [D23]</li> <li>Pharmaceutical Preparations [D26]</li> <li>Metabolism [G03]</li> <li>Genetic Phenomena [G06]</li> </ul> <p>Queries for these terms use the following string:</p> <blockquote> <p>"cells"[MeSH Terms] OR "Multiprotein Complexes"[mh] OR "Protein Aggregates"[mh] OR "Hormones, Hormone Substitutes, and Hormone Antagonists"[mh] OR "Enzymes and Coenzymes"[mh] OR "Carbohydrates"[mh] OR "Lipids"[mh] OR "Amino Acids, Peptides, and Proteins"[mh] OR "Nucleic Acids, Nucleotides, and Nucleosides"[mh] OR "Biological Factors"[mh] OR "Pharmaceutical Preparations"[mh] OR "Metabolism"[mh] OR "Cell Physiological Phenomena"[mh] OR "Genetic Phenomena"[mh]</p> </blockquote> <p>PubMed returns 11,447,521 abstracts. PMC returns, 1,720,266 documents, 509,722 of these are open access. We downloaded, parsed and concatenated 403,825 PMC open access documents into a single file `molecular_oa_pmc.tsv`. This is a 33GB TSV file with the following columns:</p> <ul> <li>File:Paragraph - a unique identifier for each paragraph</li> <li>SentenceId - the local number of the sentence in the document</li> <li>Sentence Text - tokenized text of the sentence (based on <a href="https://github.com/ClearTK/cleartk/blob/master/cleartk-token/src/main/java/org/cleartk/token/tokenizer/TokenAnnotator.java">ClearTk's TokenAnnotator.java</a>)</li> <li>Codes - <code>exLink</code> for the presence of a citation, <code>inLink</code> for the presence of link to a Figure</li> <li>Figures - Figure codes</li> <li>Headings - High level section of the paper</li> <li>Offset_Begin - offset of the start of the sentence within the paper</li> <li>Offset_End - offset of the start of the sentence within the paper</li> </ul> <p>We repeated the same process for PubMed abstracts to generate a 3.6G file (`molecular_oa_medline.tsv`) with three columns:</p> <ol> <li>Pubmed ID</li> <li>A Boolean value indicating whether the article is a review</li> <li>Text</li> </ol> <p>We concatenated the text columns of these two files into a single 30GB file (`molecular_oa.txt`) where each line is a single sentence and the text is fully tokenized. </p> <p>These three files are archived in `molecular_oa_raw_text.tar.gz`.</p> <p><strong>Fasttext Embedding</strong></p> <p>We trained a fasttext model on the raw training data (https://fasttext.cc/) using the standard `skipgram` parameter. A gzipped copy of the word embeddings is included in `fasttext.model.vec.gz` </p> <p><strong>GloVe Embedding</strong></p> <p>We trained GloVe models on the raw training data (https://nlp.stanford.edu/projects/glove/). A gzipped copy of the best performing word embeddings is included in `bio_GloVe_300.tar.gz` </p>
Simple Italian sentences ranked by readability
<p>The dataset contains 500,000 sentences extracted from the Paisà corpus (https://www.corpusitaliano.it/) which have been selected for being easy to read according to four parameters: token number, average word length, depth of the parse tree and verb "arity". The sentences are ranked by readability.</p>
Dataset for training classifiers of comparative sentences
<p><br> As there was no large publicly available cross-domain dataset for comparative argument mining, we create one composed of sentences, potentially annotated with BETTER / WORSE markers (the first object is better / worse than the second object) or NONE (the sentence does not contain a comparison of the target objects). The BETTER sentences stand for a pro-argument in favor of the first compared object and WORSE-sentences represent a con-argument and favor the second object. </p> <p>We aim for minimizing dataset domain-specific biases in order to capture the nature of comparison and not the nature of the particular domains, thus decided to control the specificity of domains by the selection of comparison targets. We hypothesized and could confirm in preliminary experiments that comparison targets usually have a common hypernym (i.e., are instances of the same class), which we utilized for selection of the compared objects pairs. </p> <p>The most specific domain we choose, is computer science with comparison targets like programming languages, database products and technology standards such as Bluetooth or Ethernet. Many computer science concepts can be compared objectively (e.g., on transmission speed or suitability for certain applications). The objects for this domain were manually extracted from List of-articles at Wikipedia. In the annotation process, annotators were asked to only label sentences from this domain if they had some basic knowledge in computer science. </p> <p>The second, broader domain is brands. It contains objects of different types (e.g., cars, electronics, and food). As brands are present in everyday life, anyone should be able to label the majority of sentences containing well-known brands such as Coca-Cola or Mercedes. Again, targets for this domain were manually extracted from `List of''-articles at Wikipedia.</p> <p>The third domain is not restricted to any topic: random. For each of 24~randomly selected seed words 10 similar words were collected based on the distributional similarity API of JoBimText (http://www.jobimtext.org). Seed words created using randomlists.com: book, car, carpenter, cellphone, Christmas, coffee, cork, Florida, hamster, hiking, Hoover, Metallica, NBC, Netflix, ninja, pencil, salad, soccer, Starbucks, sword, Tolkien, wine, wood, XBox, Yale.</p> <p>Especially for brands and computer science, the resulting object lists were large (4493 in brands and 1339 in computer science). In a manual inspection, low-frequency and ambiguous objects were removed from all object lists (e.g., RAID (a hardware concept) and Unity (a game engine) are also regularly used nouns). The remaining objects were combined to pairs. For each object type (seed Wikipedia list page or the seed word), all possible combinations were created. These pairs were then used to find sentences containing both objects. The aforementioned approaches to selecting compared objects pairs tend minimize inclusion of the domain specific data, but do not solve the problem fully though. We keep open a question of extending dataset with diverse object pairs including abstract concepts for future work. </p> <p>As for the sentence mining, we used the publicly available index of dependency-parsed sentences from the Common Crawl corpus containing over 14 billion English sentences filtered for duplicates. This index was queried for sentences containing both objects of each pair. For 90% of the pairs, we also added comparative cue words (better, easier, faster, nicer, wiser, cooler, decent, safer, superior, solid, terrific, worse, harder, slower, poorly, uglier, poorer, lousy, nastier, inferior, mediocre) to the query in order to bias the selection towards comparisons but at the same time admit comparisons that do not contain any of the anticipated cues. This was necessary as a random sampling would have resulted in only a very tiny fraction of comparisons. Note that even sentences containing a cue word do not necessarily express a comparison between the desired targets (dog vs. cat: He's the best pet that you can get, better than a dog or cat.). It is thus especially crucial to enable a classifier to learn not to rely on the existence of clue words only (very likely in a random sample of sentences with very few comparisons). For our corpus, we keep pairs with at least 100 retrieved sentences.</p> <p>From all sentences of those pairs, 2500 for each category were randomly sampled as candidates for a crowdsourced annotation that we conducted on figure-eight.com in several small batches. Each sentence was annotated by at least five trusted workers. We ranked annotations by confidence, which is the figure-eight internal measure of combining annotator trust and voting, and discarded annotations with a confidence below 50%. Of all annotated items, 71% received unanimous votes and for over 85% at least 4 out of 5 workers agreed -- rendering the collection procedure aimed at ease of annotation successful.</p> <p>The final dataset contains 7199 sentences with 271 distinct object pairs. The majority of sentences (over 72%) are non-comparative despite biasing the selection with cue words; in 70% of the comparative sentences, the favored target is named first.</p> <p>You can browse though the data here: https://docs.google.com/spreadsheets/d/1U8i6EU9GUKmHdPnfwXEuBxi0h3aiRCLPRC-3c9ROiOE/edit?usp=sharing </p> <p>Full description of the dataset is available in the <a href="https://arxiv.org/abs/1809.06152">workshop paper at ACL 2019 conference</a>. Please cite this paper if you use the data: Franzek, Mirco, Alexander Panchenko, and Chris Biemann. "Categorization of Comparative Sentences for Argument Mining." <em>arXiv preprint arXiv:1809.06152</em> (2018).</p> <pre>@inproceedings{franzek2018categorization, title={Categorization of Comparative Sentences for Argument Mining}, author={Panchenko, Alexander and Bondarenko, and Franzek, Mirco and Hagen, Matthias and Biemann, Chris}, booktitle={Proceedings of the 6th Workshop on Argument Mining at ACL'2019}, year={2019}, address={Florence, Italy} }</pre> <p> </p> <p> </p>
Short sentences on R analyses in a health informatics subject
<p>The dataset contains the list of sentences written by students, with a unique ID, its type (if either given for the hypothesis or normality test), its degree in a range from 0 to 1, and its fail/pass result, flanked with (i) the gold standard, (ii) an alternative gold standard, and (iii) the negated gold standard.</p>
Sentences with negative actors: negative strength quantified
<p>Files: data1.xml,data3.xml,data3.xml (3 annotators) - XML validiert</p> <p>- 439 sentences <br> - target: a negative cause (an actor etc.) represented by the Lemma<br> - id: sentence number<br> - string: the plain sentence<br> - strength: negativity strength of the target<br> - labels 0-3<br> - 0 no negative entity found (or parsing error)<br> - 1 slightly negative, 2 negative, 3 stronly negative<br> <br> - 115 out of 439 sentences with tag 0: i.e. sentences do not contain a negative actor<br> - different reasons (see the paper below): modal, future tense etc. but also parsing errors</p> <p><br> Data source: Facebook posts of the AfD, a German right-wing party</p> <p>Examples:</p> <p>no actor here: passive voice<br> <sent><id>1</id><target>Junge</target><strength>0</strength><string>"Verletzt wurde auch ein 11-jähriger Junge . "</string></sent><br> stronly negative:<br> <sent><id>411</id><target>Euro</target><strength>3</strength><string>"Der Euro ruiniert Europa . "</string></sent><br> negative:<br> <sent><id>214</id><target>Merkel</target><strength>2</strength><string>"Merkel verantwortet zusätzliche 50 Milliarden Sozialkosten bis 2018 . "</string></sent><br> slightly negative:<br> <sent><id>154</id><target>Meuthen</target><strength>1</strength><string>"Meuthen schadet der Partei . "</string></sent></p> <p>References:</p> <p>@inproceedings{nodalida,<br> month = {Juni},<br> author = {Manfred Klenner and Anne G{\"o}hring and Sophia Conrad},<br> booktitle = {Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa)},<br> address = {Reykjavik, Iceland},<br> title = {Getting Hold of Villains and other Rogues},<br> publisher = {Virtual Event},<br> pages = {435--439},<br> year = {2021},<br> language = {english},<br> url = {https://doi.org/10.5167/uzh-204265},<br> abstract = {In this paper, we introduce the first corpus specifying negative entities within sentences. We discuss indicators for their presence, namely particular verbs, but also the linguistic conditions when their prediction should be suppressed. We further show that a fine-tuned Bert-based baseline model outperforms an over-generating rule-based approach which is not aware of these further restrictions. If a perfect filter were applied, both would be on par.}<br> }<br> </p>
NNS-500 Acceptability judgment task dataset based on the sentences written by non-native English speakers
<p><strong>Acceptability judgment task (AJT):</strong> AJT is a common method in empirical linguistics to gather information about the internal grammar of speakers of a language, which is considered a promising area to evaluate neural language models' linguistic knowledge. There is a Corpus of Linguistic Acceptability (CoLA) whose creators think Boolean judgments sufficient; similarly, some non-English resources cast acceptability as a binary classification task.</p> <p><strong>Dataset:</strong> NNS-500 dataset based on the sentences written by non-native speakers (which is important from the point of view of the source of unacceptable sentences) and labelled by a university English teacher is intended for testing the pre-trained neural networks. It has 350 acceptable and 150 unacceptable sentences, which is 70% of acceptability (this compares to 69.2% in the CoLA out-of-domain set). The dataset markup includes standard data: id number, sentence, indication of acceptability – 1, indication of unacceptability – 0, type of error (morphology, syntax, semantics), and detailed information about the source (group number with the year of admission to the university, number according to list of the group members, and gender of the student). For the use of the assessment of EFL learners' linguistic competence, the first 100 sentences of the dataset (id 1‒100) include the ones written by the study group with a high level of academic performance (Group A) and another 100 sentences of the dataset (id 101‒200) are taken from the writing assignments of students with a poor academic performance (Group B). From each group of students, 5 people were selected (a total 10 participants); 20 sentences were randomly selected from each student's written work (14 ‒ unacceptable, 6 ‒ acceptable); there are more sentences with errors, since they are very important for the error analysis. The rest of the dataset consists of 290 acceptable and 10 unacceptable sentences (id 201‒500) taken from the works of students of different study groups of an intermediate level.</p>
Spanish Workers' Statute Sentences Dataset
<p>The Spanish Workers' Statute contains Spain's fundamental rules of labor law. It is divided into three main sections marked as "Titles". The initial title encompasses individual labor relationships, while the subsequent title covers the entitlements related to collective representation and workers' assemblies within companies. Lastly, the third title refers to collective bargaining and the agreements reached through such negotiations. In total, the three sections gather 92 articles.</p> <p>The dataset was obtained after automatically separating the raw Statute text into sentences and manually cleaning the partial result. This process leads to a dataset of 1235 sentences.</p>
Sentence-aligned student translations of Crito (Ancient Greek, English, German, Persian)
<p>This dataset is a corpus of five student's translation of Plato's Crito aligned at sentence-level with the original Ancient Greek text, one German translation, and two English translations. The Ancient Greek text is the Burnet edition, made available by Perseus Digital Library. For more information, see: <br> https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0169%3Atext%3DCrito%3Asection%3D43a<br> </p> <p>The details of the eight translations (five Persian, two English, and one German translations) are as follows:</p> <ul> <li>German Translation: The German translation of Schleiermmacher available on Project Gutenberg has been aligned at the sentence level.<br> For more information, see: Plato, F. Schleiermacher, Platons Werke, In der Realschulbuchhandlung, 1809<br> Link to the text on Project Gutenberg:<br> https://www.projekt-gutenberg.org/platon/platowr1/kriton.html<br> </li> <li>English Translations: Two different English translations of "Crito", one by Benjamin Jowett and the other by Harold North Fowler are included in the dataset.<br> For more information on Jowett's translation, see:<br> Plato, H. N. Fowler, W. Lamb, Plato in Twelve Volumes, Vol. 1 translated by Harold North<br> Fowler; Introduction by W.R.M. Lamb, volume 1, Harvard University Press and Wiliam<br> Heinemann Ltd., Cambridge, MA and London, 1966.<br> Fowler's translation on Perseus Digital Library:<br> https://www.perseus.tufts.edu/hopper/text?doc=plat.+crito+43a<br> For more information on Jowett's translation, see:<br> Plato, B. Jowett, Crito, The Internet Classics Archive, Massachusetts Institute of Technology,<br> http://classics.mit.edu/Plato/crito.html.<br> </li> <li>Persian Translations: The dataset consists of five Persian translations by students who have already completed a 30-hour Homeric Greek course. Each translator has translated the text into Persian using treebanks, commentaries, lexicon entries, and English and German translations. The translators themselves aligned the Persian translations to the Greek text at word-level using Ugarit. The alignments are available in their Ugarit profile:<br> Shouresh Assimi: https://ugarit.ialigner.com/userProfile.php?userid=50956<br> Aylar Mahmoudzadeh Sarabi: https://ugarit.ialigner.com/userProfile.php?userid=63464&tgid=9576<br> Nima Mohammadi: https://ugarit.ialigner.com/userProfile.php?userid=52434&tgid=9362<br> Kimia Nikpour: https://ugarit.ialigner.com/userProfile.php?userid=52378<br> Farshid Rahimi: https://ugarit.ialigner.com/userProfile.php?userid=50932&tgid=9727</li> </ul> <p>The group's initial goal was to produce one finalized translation of Crito to Persian, but due to the intriguing variations in the translations and the text's intricacy, it was decided to provide three finalized translations rather than one. The finalized translations will be available in Beyond Translation as part of the Perseus Digital Library under a Creative Commons license. For more information on our final versions of Crito, see: http://beyond-translation.perseus.org</p>
Video recordings for the female French Matrix Sentence Test
<p>This dataset provides the video recordings for the female French Matrix Sentence Test, a speech intelligibility test. The audio recordings are not provided in the dataset, but research licenses are available from Hörzentrum Oldenburg gGmbH. Please refer to Hörzentrum Oldenburg gGmbH for the audio material (<a href="https://www.hz-ol.de/en/matrix.html">https://www.hz-ol.de/en/matrix.html</a>).</p> <p>For testing audiovisual speech perception, the video recordings should be played together with the original speech mentioned above. Half a second of silence at the beginning of each audio file should be added for synchronous playback. An example with video and original audio is provided here (exampleWithAudio_00200_JeanLuc_ramasse_quinze_classeurs_jaunes.mp4). The research software that runs the test is available for collaborative work with Hörzentrum Oldenburg gGmbH. Please contact <a href="mailto:sales@hz-ol.de.">sales@hz-ol.de</a></p> <p>The scripts (Matlab) and guidelines to create the video recordings can be found in the public repository <a href="https://github.com/gerardllorach/audiovisualdubbedMST">https://github.com/gerardllorach/audiovisualdubbedMST</a>. The videos for the German version can be found here: <a href="https://doi.org/10.5281/zenodo.3673062">https://doi.org/10.5281/zenodo.3673062</a></p> <p>Main authors: Loïc Le Ruhn, Gerard Llorach Tó</p> <p>Corresponding author: Tanguy Delmas (tanguy.delmas (at) pasteur.fr)</p> <p> </p> <p>References:</p> <p>Le Rhun, L., Llorach, G., Delmas, T., Suied, C., Arnal, L., & Lazard, D. (2023). A standardized test to evaluate audio-visual speech intelligibility in French. <em>medRxiv</em>, 2023-01. https://doi.org/10.1101/2023.01.18.23284110</p> <p>Jansen, S., Luts, H., Wagener, K. C., Kollmeier, B., Del Rio, M., Dauman, R., James, C., Fraysse, B., Vormès, E., Frachet, B., Wouters, J., & van Wieringen, A. (2012). Comparison of three types of French speech-in-noise tests: A multi-center study. International Journal of Audiology, 51(3), Art. 3. https://doi.org/10.3109/14992027.2011.633568 </p> <p>Llorach, G., Kirschner, F., Grimm, G., Zokoll, M. A., Wagener, K. C., & Hohmann, V. (2022). Development and evaluation of video recordings for the OLSA matrix sentence test. International Journal of Audiology, 61(4), 311‑321. https://doi.org/10.1080/14992027.2021.1930205 </p>
Evoking the N400 Event-Related Potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)
Open the record for dataset details and reuse information.
[Stimulus Set] Evoking the N400 event-related potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)
Open the record for dataset details and reuse information.
Video recordings for the female German Matrix Sentence Test (OLSA)
<p>This dataset provides the video recordings for the female German Matrix Sentence Test (OLSA), a speech intelligibility test. The audio recordings are not provided in the dataset, but research licenses are available from Hörzentrum Oldenburg gGmbH. Please refer to Hörzentrum Oldenburg gGmbH for the audio material (<a href="https://www.hz-ol.de/en/matrix.html">https://www.hz-ol.de/en/matrix.html</a>). An example can be seen in "exampleWithAudio.mp4". For testing audiovisual speech perception, the video recordings should be played together with the original speech OLSAf mentioned above. One second of silence at the beginning of each audio file should be added for synchronous playback.</p> <p>Additionally, a web demo is provided to see how the test functions without sound (try it here with Chrome: <a href="http://www.staff.uni-oldenburg.de/gerard.llorach.to/AVOLSA/">http://www.staff.uni-oldenburg.de/gerard.llorach.to/AVOLSA/</a>). In order to make the web demo work, the folders "Video.zip" and "webData.zip" should be unzipped and placed where "A_demo.html" is. The web demo permits to test visual-only, audiovisual, and audio-only modalities. For testing with audio (if acquired), please drag and drop the audio files ("01248.wav", "02064.wav"...) in the web interface. For further information about the web demo, please read the introduction inside the file "A_demo.html". For further information about the material and how it was recorded please refer to the references.</p> <p>Special thanks to Jutta Birkigt, the talker of the audio and video recordings of the female German Matrix Sentence Test.</p> <p>Reference:</p> <p>Llorach, G., Kirschner, F., Grimm, G., Zokoll, M.A., Wagener, K.C. and Hohmann, V., 2021. Development and evaluation of video recordings for the OLSA matrix sentence test. <em>International Journal of Audiology</em>, pp.1-11.</p>
Crossing hands behind your back reduces recall of manual action sentences and alters brain dynamics
<p>The experiment involved a 2 Verb (manual, attentional) x 2 Hand posture (front, behind) within-participant design. Half of the participants were assigned randomly to one of the two sets of 120 experimental sentences. The sentences were divided into 12 experimental blocks of 11 sentences each: 5 manual and 5 attentional sentences as experimental materials, and 1 filler sentence that was added at the beginning of each block to discard recall primacy effects and excluded from the analysis. An additional block of 5 sentences was used as practice. Each experimental block consisted of: a) learning phase, including the 10 experimental sentences in random order, plus the initial filler sentence, b) one-minute distractive task, c) recall phase.</p> <p>For more information about the data, please refer to the Corresponding Author</p>
An xml collection of example sentences for Lhasa Tibetan verbs
<p>This is an xml collection of example sentences for Lhasa Tibetan verbs</p>
Sentence representations generated by Inner Attention model (arxiv: 1707.03103)
<p>600-dimensional sentence vector representations created by the model described in the paper "Refining Raw Sentence Representations for Textual Entailment Recognition via Attention".</p> <p>The dataset is in tab-delimited format: ID\tSENTENCE_TYPE\tVECTOR, where ID is the id corresponding to the sentence pair as specified in the Repeval 2017 test dataset for both matched and mismatched evaluations, available in https://inclass.kaggle.com/c/multinli-matched-evaluation/download/multinli_0.9_test_matched_unlabeled.jsonl and https://inclass.kaggle.com/c/multinli-mismatched-evaluation/download/multinli_0.9_test_mismatched_unlabeled.jsonl (you will probably have to create an account to download them).</p> <p>SENTENCE_TYPE can either be p, meaning the sentence is the premise or h, meaning it is the hypothesis.</p> <p>VECTOR is a space-delimited 600-dim vector.</p>
Average sentence length in novels dataset (Avg-Sent-Len)
<p>This dataset includes values on average sentence length per novel, as well as very basic metadata (author name and year of firstpublication), for several sets of 100 novels in multiple languages (English, French, German and Hungarian, for the period 1840-1920) derived from the European Literary Text Collection (ELTeC) as well as for one larger set of 1150 English-language novels from Project Gutenberg (for the period 1800-1960). </p><p>This dataset is suitable for learning data visualization (scatterplots and regression plots) as well as data analytics (linear / polynomial regression). </p><p>Source of the data: <a href="https://github.com/christofs/sentlens">https://github.com/christofs/sentlens</a>. </p><p>Use of the data: <a href="https://dragonfly.hypotheses.org/1152">https://dragonfly.hypotheses.org/1152</a>. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.