Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
117
datasets available to search
ShareScore release 0.7.1
Dataset results
117 results for “text analysis”
Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text
<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>
Data and analysis script for "The (non)effect of personalization in climate texts on credibility of climate scientists: A case study on sustainable travel"
<p>Dataset and analysis script for the article "<strong>The (non)effect of personalization in climate texts on credibility of climate scientists</strong><strong>: A case study on sustainable travel</strong>", under review at Geoscience Communication (https://doi.org/10.5194/egusphere-2024-543)</p>
Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 5
<p>Dataset containing four .xlsx and .csv files for the exercises in Episode 5 of the <a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a> lesson of the <a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a> project. The original data was collected from <a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>
Aššur and His Friends: A Statistical Analysis of Neo-Assyrian Texts
<p>This is the data used for and generated during our research for the article "Aššur and His Friends: A Statistical Analysis of Neo-Assyrian Texts", published in <em>Journal of Cuneiform Studies </em>71 (2019).</p>
Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 2
<p>Dataset containing three subgenre-specific .xlsx files for the exercises in Episode 2 of the <a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a> lesson of the <a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a> project. The original data was collected from <a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>
Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text
<p>Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text covering the Olympic legacy of Rio 2016 and London 2012. Data was searched via Google search engine. It is composed of sentiment labels assigned to 1271 news articles in total.</p> <p><strong>News outlets:</strong></p> <ul> <li>BBC</li> <li>Daily Mail</li> <li>The Telegraph</li> <li>The Guardian</li> <li>Globo</li> <li>Estadao</li> <li>Folha de S. Paulo</li> </ul> <p><strong>Events covered by the articles:</strong></p> <ul> <li>London 2012 Olympic legacy</li> <li>Rio 2016 Olympic legacy</li> </ul> <p>All classifiers were used in texts in English. Text originally published in Portuguese by the Brazilian media were automatically translated.</p> <p><strong>Sentiment classifiers used:</strong></p> <ul> <li>Vader</li> <li>BERT (Trained on Amazon data)</li> <li>BERT (Trained on twitter data - 140)</li> </ul> <p>Each document (spreadsheet - xlsx) refers to one outlet and one event (London 2012 or Rio 2016).</p> <p><strong>How were labels assigned to the texts?</strong></p> <p>These labels are a combination of the three sentiment classifiers listed above. If two of them agree with the same label, then this label would be considered as right. Otherwise, the label ‘other’ was assigned.</p> <p>For news article body text: the proportion of sentences of each sentiment type was used to assign labels to the whole article instead of averaging the sentence scores. For example, if the proportion of sentences with negative labels is greater than 50%, then the article is assigned a negative label.</p> <p><strong>The documents are composed of the following columns:</strong></p> <ul> <li>Rank: the position of the article on Google search ranking</li> <li>Date: date of article's publication (DD/MM/YYYY)</li> <li>Link: article's link</li> <li>Title: article's title</li> <li>Sentiment_Title: final sentiment for article headline</li> <li>Sentiment_Text: final sentiment for article's body text</li> </ul> <p><em>PS: Documents do not include articles' body text. </em></p> <p><strong>Sentiment is presented in labels as follows:</strong></p> <ul> <li>Pos: Positive</li> <li>Neg: Negative</li> <li>Neutral: Neutral</li> <li>other: inconclusive - if each of the 3 classifiers assigned a different label to the article, the label 'other' was used. Therefore, 'other' identifies contradictory results.</li> </ul> <p> </p>
Text-fig. 5. Vegetation zones in P. R. China (Editorial Committee of Vegetation Map of China, The Chinese Academy of Sciences 2007), and assumed location of extant reference vegetation type of Wiesa fossil assemblage (rectangle), as revealed from qualitative floristic analysis. Extant reference vegetation type present in southern belt of zone of subtropical evergreen broadleaved forest, with minor overlap into zone of tropical forest. in Assessment Of Phytogeographic Reference Regions For Cenozoic Vegetation: A Case Study On The Miocene Flora Of Wiesa (Germany)
Text-fig. 5. Vegetation zones in P. R. China (Editorial Committee of Vegetation Map of China, The Chinese Academy of Sciences 2007), and assumed location of extant reference vegetation type of Wiesa fossil assemblage (rectangle), as revealed from qualitative floristic analysis. Extant reference vegetation type present in southern belt of zone of subtropical evergreen broadleaved forest, with minor overlap into zone of tropical forest.
Text-fig. 4. Graphical visualization of Phytogeographic Reference Regions Assessment (PRRA) of nearest living relative genera of fossil-taxa from late Early Miocene Wiesa assemblage in eastern Germany. Analysis yields only NLRs which have modern distribution area (partly) in E and SE Asia. For relationships of fossil-taxa to nearest living relatives or ecological equivalents, see Tab. 6; taxa used for analysis marked with asterisks. Three geographic resolutions conducted: a – grid with 1.5° latitude/longitude resolution, b – grid with 2°, c – grid with 3°; similarity column indicates cooccurrences of genera of nearest living relatives in single grid box. Maximum value in our analysis: grid box marked with arrow in map a, located in western Yunnan Province, P. R. China and southern Kachin Province, NE Myanmar (east of Myitkyina city), area with 97.371 7–98.874 2° longitude and 24.586 7–25.837 5° latitude, yields 23 co-occurring species of 13 genera (Tab. 7). in Assessment Of Phytogeographic Reference Regions For Cenozoic Vegetation: A Case Study On The Miocene Flora Of Wiesa (Germany)
Text-fig. 4. Graphical visualization of Phytogeographic Reference Regions Assessment (PRRA) of nearest living relative genera of fossil-taxa from late Early Miocene Wiesa assemblage in eastern Germany. Analysis yields only NLRs which have modern distribution area (partly) in E and SE Asia. For relationships of fossil-taxa to nearest living relatives or ecological equivalents, see Tab. 6; taxa used for analysis marked with asterisks. Three geographic resolutions conducted: a – grid with 1.5° latitude/longitude resolution, b – grid with 2°, c – grid with 3°; similarity column indicates cooccurrences of genera of nearest living relatives in single grid box. Maximum value in our analysis: grid box marked with arrow in map a, located in western Yunnan Province, P. R. China and southern Kachin Province, NE Myanmar (east of Myitkyina city), area with 97.371 7–98.874 2° longitude and 24.586 7–25.837 5° latitude, yields 23 co-occurring species of 13 genera (Tab. 7).
Рис. 4. ГенитаΛьные структуры самки Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 4. Female genitalia structures of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text) in Morphometric analysis of genitalia of Ctenoceratoda tancrei (Graeser, 1892) (Lepidoptera, Noctuidae)
Рис. 4. ГенитаΛьные структуры самки Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 4. Female genitalia structures of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text)
Рис. 3. ЭΑеагус Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 3. Aedeagus of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text) in Morphometric analysis of genitalia of Ctenoceratoda tancrei (Graeser, 1892) (Lepidoptera, Noctuidae)
Рис. 3. ЭΑеагус Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 3. Aedeagus of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text)
A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text Spatializations
<p>Result Files for the Paper "A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text Spatializations" to be published at IEEE Vis 2024</p>
Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.
<p>This dataset is a subset of 596 documents from the <em>Registre d'Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Històric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called diplomatic criteria. Additionally, transcripts were tagged with <br> extra enriching/complementary information (e.g. expansion of the abbreviations, hyphen marks, etc.). Along with the transcripts the layout of the document is detected and recorded. Pages have been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d'Història Rural</em></a> and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights
<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook </a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - images
<p>In this notebook we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains high quality versions of the images used in the analysis.</p>
Text-fig. 5. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MPwarm. For legend see Text-fig. 3. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)
Text-fig. 5. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MPwarm. For legend see Text-fig. 3.
Text-fig. 4. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAP. For legend see Text-fig. 3. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)
Text-fig. 4. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAP. For legend see Text-fig. 3.
Text-fig. 1. Location of the Tossignano and Monte Tondo sites in the Romagna Apennines (modified after Lugli et al. 2010). 1 – Tossignano quarry, 2 – Monte Tondo quarry. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)
Text-fig. 1. Location of the Tossignano and Monte Tondo sites in the Romagna Apennines (modified after Lugli et al. 2010). 1 – Tossignano quarry, 2 – Monte Tondo quarry.
Text-fig. 3. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAT. Right-hand positoned large bold figures and shaded areas in each case indicate the Coexistence Interval, with the number of overlapping taxa being at a maximum. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)
Text-fig. 3. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAT. Right-hand positoned large bold figures and shaded areas in each case indicate the Coexistence Interval, with the number of overlapping taxa being at a maximum.
Text-fig. 2. Stratigraphy of western part of the Romagna Apennines (after Roveri et al. 2006). Symbol "arrow" – stratigraphical position of the studied floras of Tossignano and Monte Tondo. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)
Text-fig. 2. Stratigraphy of western part of the Romagna Apennines (after Roveri et al. 2006). Symbol "arrow" – stratigraphical position of the studied floras of Tossignano and Monte Tondo.
Text-fig. 4. Electrophoresis after amplification: Electrophoretical analysis of mitochondrial DNA. mtDNA sequences were amplified by primers F15.412 and R16.169 (450 bp), R16.269 (550 bp), R16.519 (800 bp). Lane 1 are primers F15.412 + R16.169, lane 2 primers F15.412 + R16.269, lane 3 primers F15.412 + R16.519, NC – negative control – water, L – 100 bp DNA ladder (band size from 100 bp to 1500 bp). in Genetic Analysis Of Possibly The Oldest Greyhound Remains Within The Territory Of The Czech Republic As Proof Of A Local Elite Presence At Chotěbuz-Podobora Hillfort In The 8 -9 Century Ad
Text-fig. 4. Electrophoresis after amplification: Electrophoretical analysis of mitochondrial DNA. mtDNA sequences were amplified by primers F15.412 and R16.169 (450 bp), R16.269 (550 bp), R16.519 (800 bp). Lane 1 are primers F15.412 + R16.169, lane 2 primers F15.412 + R16.269, lane 3 primers F15.412 + R16.519, NC – negative control – water, L – 100 bp DNA ladder (band size from 100 bp to 1500 bp).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.