Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

117

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

117 results for “text analysis”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text

<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Data and analysis script for "The (non)effect of personalization in climate texts on credibility of climate scientists: A case study on sustainable travel"

<p>Dataset and analysis script for the article "<strong>The (non)effect of personalization in climate texts on credibility of climate scientists</strong><strong>: A case study on sustainable travel</strong>", under review at Geoscience Communication (https://doi.org/10.5194/egusphere-2024-543)</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 5

<p>Dataset containing four .xlsx and .csv files for the exercises in Episode 5 of the&nbsp;<a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a>&nbsp;lesson of the&nbsp;<a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a>&nbsp;project. The original data was collected from&nbsp;<a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Aššur and His Friends: A Statistical Analysis of Neo-Assyrian Texts

<p>This is the data used for and generated during our research for the article &quot;A&scaron;&scaron;ur and His Friends: A Statistical Analysis of Neo-Assyrian Texts&quot;, published in <em>Journal of Cuneiform Studies </em>71 (2019).</p>

opencc-by-sa-3.0Mar 2019View details →
zenodo40/100

Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 2

<p>Dataset containing three subgenre-specific .xlsx files for the exercises in Episode 2 of the <a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a> lesson of the <a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a> project. The original data was collected from <a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text

<p>Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text covering the Olympic legacy of Rio 2016 and London 2012. Data was searched via Google search engine. It is composed of sentiment labels assigned to 1271 news articles in total.</p> <p><strong>News outlets:</strong></p> <ul> <li>BBC</li> <li>Daily Mail</li> <li>The Telegraph</li> <li>The Guardian</li> <li>Globo</li> <li>Estadao</li> <li>Folha de S. Paulo</li> </ul> <p><strong>Events covered by the articles:</strong></p> <ul> <li>London 2012 Olympic legacy</li> <li>Rio 2016 Olympic legacy</li> </ul> <p>All classifiers were used in texts in English. Text originally published in Portuguese by the Brazilian media were automatically translated.</p> <p><strong>Sentiment classifiers used:</strong></p> <ul> <li>Vader</li> <li>BERT (Trained on Amazon data)</li> <li>BERT (Trained on twitter data - 140)</li> </ul> <p>Each document (spreadsheet - xlsx) refers to one outlet and one event (London 2012 or Rio 2016).</p> <p><strong>How were labels assigned to the texts?</strong></p> <p>These labels are a combination of the three sentiment classifiers listed above. If two of them agree with the same label, then this label would be considered as right. Otherwise, the label &lsquo;other&rsquo; was assigned.</p> <p>For news article body text: the proportion of sentences of each sentiment type was used to assign labels to the whole article instead of averaging the sentence scores. For example, if the proportion of sentences with negative labels is greater than 50%, then the article is assigned a negative label.</p> <p><strong>The documents are composed of the following columns:</strong></p> <ul> <li>Rank: the position of the article on Google search ranking</li> <li>Date: date of article&#39;s publication (DD/MM/YYYY)</li> <li>Link: article&#39;s link</li> <li>Title: article&#39;s title</li> <li>Sentiment_Title: final sentiment for article headline</li> <li>Sentiment_Text: final sentiment for article&#39;s body text</li> </ul> <p><em>PS: Documents do not include articles&#39; body text. </em></p> <p><strong>Sentiment is presented in labels as follows:</strong></p> <ul> <li>Pos: Positive</li> <li>Neg: Negative</li> <li>Neutral: Neutral</li> <li>other: inconclusive - if each of the 3 classifiers assigned a different label to the article, the label &#39;other&#39; was used. Therefore, &#39;other&#39; identifies contradictory results.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Text-fig. 5. Vegetation zones in P. R. China (Editorial Committee of Vegetation Map of China, The Chinese Academy of Sciences 2007), and assumed location of extant reference vegetation type of Wiesa fossil assemblage (rectangle), as revealed from qualitative floristic analysis. Extant reference vegetation type present in southern belt of zone of subtropical evergreen broadleaved forest, with minor overlap into zone of tropical forest. in Assessment Of Phytogeographic Reference Regions For Cenozoic Vegetation: A Case Study On The Miocene Flora Of Wiesa (Germany)

Text-fig. 5. Vegetation zones in P. R. China (Editorial Committee of Vegetation Map of China, The Chinese Academy of Sciences 2007), and assumed location of extant reference vegetation type of Wiesa fossil assemblage (rectangle), as revealed from qualitative floristic analysis. Extant reference vegetation type present in southern belt of zone of subtropical evergreen broadleaved forest, with minor overlap into zone of tropical forest.

opencc-by-4.0Aug 2022View details →
zenodo40/100

Text-fig. 4. Graphical visualization of Phytogeographic Reference Regions Assessment (PRRA) of nearest living relative genera of fossil-taxa from late Early Miocene Wiesa assemblage in eastern Germany. Analysis yields only NLRs which have modern distribution area (partly) in E and SE Asia. For relationships of fossil-taxa to nearest living relatives or ecological equivalents, see Tab. 6; taxa used for analysis marked with asterisks. Three geographic resolutions conducted: a – grid with 1.5° latitude/longitude resolution, b – grid with 2°, c – grid with 3°; similarity column indicates cooccurrences of genera of nearest living relatives in single grid box. Maximum value in our analysis: grid box marked with arrow in map a, located in western Yunnan Province, P. R. China and southern Kachin Province, NE Myanmar (east of Myitkyina city), area with 97.371 7–98.874 2° longitude and 24.586 7–25.837 5° latitude, yields 23 co-occurring species of 13 genera (Tab. 7). in Assessment Of Phytogeographic Reference Regions For Cenozoic Vegetation: A Case Study On The Miocene Flora Of Wiesa (Germany)

Text-fig. 4. Graphical visualization of Phytogeographic Reference Regions Assessment (PRRA) of nearest living relative genera of fossil-taxa from late Early Miocene Wiesa assemblage in eastern Germany. Analysis yields only NLRs which have modern distribution area (partly) in E and SE Asia. For relationships of fossil-taxa to nearest living relatives or ecological equivalents, see Tab. 6; taxa used for analysis marked with asterisks. Three geographic resolutions conducted: a – grid with 1.5° latitude/longitude resolution, b – grid with 2°, c – grid with 3°; similarity column indicates cooccurrences of genera of nearest living relatives in single grid box. Maximum value in our analysis: grid box marked with arrow in map a, located in western Yunnan Province, P. R. China and southern Kachin Province, NE Myanmar (east of Myitkyina city), area with 97.371 7–98.874 2° longitude and 24.586 7–25.837 5° latitude, yields 23 co-occurring species of 13 genera (Tab. 7).

opencc-by-4.0Aug 2022View details →
zenodo40/100

Рис. 4. ГенитаΛьные структуры самки Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 4. Female genitalia structures of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text) in Morphometric analysis of genitalia of Ctenoceratoda tancrei (Graeser, 1892) (Lepidoptera, Noctuidae)

Рис. 4. ГенитаΛьные структуры самки Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 4. Female genitalia structures of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text)

opencc-by-4.0Dec 2022View details →
zenodo40/100

Рис. 3. ЭΑеагус Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 3. Aedeagus of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text) in Morphometric analysis of genitalia of Ctenoceratoda tancrei (Graeser, 1892) (Lepidoptera, Noctuidae)

Рис. 3. ЭΑеагус Ctenoceratoda tancrei с усΛовными обозначениями морфометрических характеристик (см. текст) Fig. 3. Aedeagus of Ctenoceratoda tancrei with morphometric characteristics abbreviations (see text)

opencc-by-4.0Dec 2022View details →
zenodo40/100

A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text Spatializations

<p>Result Files for the Paper "A Large-Scale Sensitivity Analysis on Latent Embeddings and Dimensionality Reductions for Text Spatializations" to be published at IEEE Vis 2024</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.

<p>This dataset is a subset of 596 documents from the&nbsp;<em>Registre d&#39;Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Hist&ograve;ric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary&nbsp;typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines&nbsp;written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called&nbsp;diplomatic criteria. Additionally, transcripts were tagged with&nbsp;<br> extra enriching/complementary information (e.g. expansion of the&nbsp;abbreviations, hyphen marks, etc.). Along with the transcripts &nbsp;the layout of the document is detected and recorded. Pages have&nbsp;been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d&#39;Hist&ograve;ria Rural</em></a>&nbsp;and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>

opencc-by-nc-4.0Jul 2018View details →
zenodo40/100

Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights

<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook&nbsp;</a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - images

<p>In this notebook we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains high quality versions of the images used in the analysis.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Text-fig. 5. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MPwarm. For legend see Text-fig. 3. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)

Text-fig. 5. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MPwarm. For legend see Text-fig. 3.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 4. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAP. For legend see Text-fig. 3. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)

Text-fig. 4. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAP. For legend see Text-fig. 3.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 1. Location of the Tossignano and Monte Tondo sites in the Romagna Apennines (modified after Lugli et al. 2010). 1 – Tossignano quarry, 2 – Monte Tondo quarry. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)

Text-fig. 1. Location of the Tossignano and Monte Tondo sites in the Romagna Apennines (modified after Lugli et al. 2010). 1 – Tossignano quarry, 2 – Monte Tondo quarry.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 3. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAT. Right-hand positoned large bold figures and shaded areas in each case indicate the Coexistence Interval, with the number of overlapping taxa being at a maximum. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)

Text-fig. 3. CA climate charts for the Monte Tondo and Tossignano floras, showing climatic ranges of the Nearest Living Relatives of the fossil taxa with respect to MAT. Right-hand positoned large bold figures and shaded areas in each case indicate the Coexistence Interval, with the number of overlapping taxa being at a maximum.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 2. Stratigraphy of western part of the Romagna Apennines (after Roveri et al. 2006). Symbol "arrow" – stratigraphical position of the studied floras of Tossignano and Monte Tondo. in Palaeoenvironmental Analysis Of The Messinian Macrofossil Floras Of Tossignano And Monte Tondo (Vena Del Gesso Basin, Romagna Apennines, Northern Italy)

Text-fig. 2. Stratigraphy of western part of the Romagna Apennines (after Roveri et al. 2006). Symbol "arrow" – stratigraphical position of the studied floras of Tossignano and Monte Tondo.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 4. Electrophoresis after amplification: Electrophoretical analysis of mitochondrial DNA. mtDNA sequences were amplified by primers F15.412 and R16.169 (450 bp), R16.269 (550 bp), R16.519 (800 bp). Lane 1 are primers F15.412 + R16.169, lane 2 primers F15.412 + R16.269, lane 3 primers F15.412 + R16.519, NC – negative control – water, L – 100 bp DNA ladder (band size from 100 bp to 1500 bp). in Genetic Analysis Of Possibly The Oldest Greyhound Remains Within The Territory Of The Czech Republic As Proof Of A Local Elite Presence At Chotěbuz-Podobora Hillfort In The 8 -9 Century Ad

Text-fig. 4. Electrophoresis after amplification: Electrophoretical analysis of mitochondrial DNA. mtDNA sequences were amplified by primers F15.412 and R16.169 (450 bp), R16.269 (550 bp), R16.519 (800 bp). Lane 1 are primers F15.412 + R16.169, lane 2 primers F15.412 + R16.269, lane 3 primers F15.412 + R16.519, NC – negative control – water, L – 100 bp DNA ladder (band size from 100 bp to 1500 bp).

opencc-by-4.0Oct 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record