Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “Brazilian portuguese”

Learn how ShareScore rates datasets ↗
zenodo44/100

A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition (BRESSAY)

<p>The BRESSAY dataset comprises images of handwritten essays in Brazilian Portuguese, which present a series of challenges to optical recognition models. These images were sourced from multiple online platforms, limiting our ability to standardize the capture process. Due to these varied sources and the lack of a uniform collection method, the dataset provides a realistic reflection of real-world conditions. Each essay is unique, contributed by different writers, and addresses a specific content topic. Furthermore, the constraints placed on the writers often lead to various handwriting scenarios, including hard-to-read words, connected words, noise, overwriting, and struck-through texts.</p> <h3><strong>Technical Details</strong></h3> <p>The BRESSAY dataset represents a comprehensive collection of handwritten essays in Brazilian Portuguese, offering detailed insights into various handwriting scenarios. It covers a total of 1,000 pages, each contributed by a unique writer, resulting in 1,000 distinct handwriting styles. This aspect of the dataset adds a layer of diversity, which is further emphasized by the total of 4,214 paragraphs, 30,090 lines, and 416,826 words. Regarding unique tokens, we have 41,318 unique words, and 107 unique characters.</p> <h3><strong>Data Structure</strong></h3> <p>The dataset is organized as follows:</p> <ul> <li>data/: Main folder containing segmented essay images <ul> <li>lines/: Images of individual lines <ul> <li>&nbsp; PNG files: Line images</li> <li>&nbsp; TXT files: Transcriptions of lines</li> </ul> </li> <li>pages/: Full page essay images <ul> <li>&nbsp; PNG files: Page images</li> <li>&nbsp; TXT files: Transcriptions of pages</li> </ul> </li> <li>paragraphs/: Images of paragraphs <ul> <li>&nbsp; PNG files: Paragraph images</li> <li>&nbsp; TXT files: Transcriptions of paragraphs</li> </ul> </li> <li>words/: Images of individual words <ul> <li>&nbsp; PNG files: Word images</li> <li>&nbsp; TXT files: Transcriptions of words</li> </ul> </li> </ul> </li> <li>sets/: Contains partition files <ul> <li>test.txt: Names of images in the test set</li> <li>validation.txt: Names of images in the validation set</li> <li>training.txt: Names of images in the training set</li> </ul> </li> </ul> <h3><strong>Dataset Usage and Annotations</strong></h3> <p>Each name in test.txt, validation.txt and training.txt represents the name of the page and all its content (words, lines, paragraphs) must be in the respective partition.</p> <p>Annotations used in the dataset:</p> <ul> <li>&nbsp; <code>##@@???@@##</code>: Superscript text that has become unidentifiable and unreadable.</li> <li>&nbsp; <code>$$@@???@@$$</code>: Subscript text that has become unidentifiable and unreadable.</li> <li>&nbsp; <code>@@???@@</code>: Text that cannot be read or identified due to its illegibility.</li> <li>&nbsp; <code>##--xxx--##</code>: Text that has been added as a superscript and subsequently crossed out, rendering it illegible.</li> <li>&nbsp; <code>$$--xxx--$$</code>: Text that has been added as a subscript and subsequently crossed out, rendering it illegible.</li> <li>&nbsp; <code>--xxx--</code>: Text that has been crossed out in a way that makes it unreadable.</li> <li>&nbsp; <code>##--text--##</code>: Text that has been added as a superscript and subsequently crossed out, but remains legible.</li> <li>&nbsp; <code>$$--text--$$</code>: Text that has been added as a subscript and subsequently crossed out, but remains legible.</li> <li>&nbsp; <code>##text##</code>: Text added as a superscript in the line, typically as a correction or additional note.</li> <li>&nbsp; <code>$$text$$</code>: Text added as a subscript in the line, typically as a correction or additional note.</li> <li>&nbsp; <code>--text--</code>: Text that has been crossed out but remains readable.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo44/100

Brazilian Portuguese COVID-19 Tweets

<p><strong>Brazilian Portuguese&nbsp;symptoms about COVID-19:</strong></p> <ul> <li><strong>Source</strong>: Twitter</li> <li><strong>Start</strong>: 2019-01-01 (January 1st)</li> <li><strong>End</strong>: 2021-09-30 (September 30th)</li> <li><strong>Tweets</strong>: 13,859,059 <ul> <li>Year 2019 [full year]: 4,043,958 obs. of 26 variables&nbsp;(Brazil_Portuguese_COVID19_Tweets2019.csv)</li> <li>Year 2020 [full year]: 6,155,844 obs. of 26 variables (Brazil_Portuguese_COVID19_Tweets2020.csv)</li> <li>Year 2021 [Q1 - Q3]: 3,659,257&nbsp;obs. of 26 variables (Brazil_Portuguese_COVID19_Tweets2021.csv)</li> </ul> </li> </ul> <p><strong>Search terms (56 symptoms keywords about COVID-19):</strong></p> <p><strong>(1)</strong> adinamia, <strong>(2)</strong> ageusia, <strong>(3)</strong> anosmia, <strong>(4)</strong> boca azulada, <strong>(5)</strong> calafrio, <strong>(6)</strong> cansa&ccedil;o, <strong>(7) </strong>cefaleia,&nbsp;<strong>(8)</strong> cianose,&nbsp;<strong>(9)</strong> colora&ccedil;&atilde;o azulada no rosto,&nbsp;<strong>(10)</strong> congest&atilde;o nasal,&nbsp;<strong>(11)</strong> conjuntivite,&nbsp;<strong>(12) </strong>coriza,&nbsp;<strong>(13)</strong> desconforto respirat&oacute;rio,&nbsp;<strong>(14)</strong> diarreia,&nbsp;<strong>(15)</strong> dificuldade para respirar,&nbsp;<strong>(16)</strong> diminui&ccedil;&atilde;o do apetite,&nbsp;<strong>(17)</strong> dispneia,&nbsp;<strong>(18)</strong> dist&uacute;rbio gustativo,&nbsp;<strong>(19)</strong> dist&uacute;rbio olfativo,&nbsp;<strong>(20)</strong> dor abdominal,&nbsp;<strong>(21)</strong> dor de cabe&ccedil;a,&nbsp;<strong>(22)</strong> dor de garganta,&nbsp;<strong>(23)</strong> dor no corpo,&nbsp;<strong>(24)</strong> dor no peito,&nbsp;<strong>(25) </strong>dor persistente no t&oacute;rax,&nbsp;<strong>(26) </strong>erup&ccedil;&atilde;o cut&acirc;nea na pele,&nbsp;<strong>(27)</strong> fadiga,&nbsp;<strong>(28)</strong> falta de ar,&nbsp;<strong>(29)</strong> febre,&nbsp;<strong>(30)</strong> gripe,&nbsp;<strong>(31)</strong> hiporexia,&nbsp;<strong>(32)</strong> inapet&ecirc;ncia,&nbsp;<strong>(33)</strong> infec&ccedil;&atilde;o respirat&oacute;ria,&nbsp;<strong>(34)</strong> l&aacute;bio azulado,&nbsp;<strong>(35)</strong> mialgia,&nbsp;<strong>(36)</strong> nariz entupido,&nbsp;<strong>(37) </strong>n&aacute;usea,&nbsp;<strong>(38)</strong> obstru&ccedil;&atilde;o nasal,&nbsp;<strong>(39)</strong> perda de apetite,&nbsp;<strong>(40)</strong> perda do olfato,&nbsp;<strong>(41)</strong> perda do paladar,&nbsp;<strong>(42)</strong> pneumonia,&nbsp;<strong>(43)</strong> press&atilde;o no peito,&nbsp;<strong>(44)</strong> press&atilde;o no t&oacute;rax,&nbsp;<strong>(45)</strong> prostra&ccedil;&atilde;o,&nbsp;<strong>(46)</strong> quadro gripal,&nbsp;<strong>(47)</strong> quadro respirat&oacute;rio,&nbsp;<strong>(48)</strong> queda da satura&ccedil;&atilde;o,&nbsp;<strong>(49)</strong> resfriado,&nbsp;<strong>(50)</strong> rosto azulado,&nbsp;<strong>(51)</strong> satura&ccedil;&atilde;o baixa,&nbsp;<strong>(52)</strong> satura&ccedil;&atilde;o de o2 menor que 95%,&nbsp;<strong>(53)</strong> s&iacute;ndrome respirat&oacute;ria aguda grave,&nbsp;<strong>(54) </strong>srag,&nbsp;<strong>(55)</strong> tosse,&nbsp;<strong>(56)</strong> v&ocirc;mito.</p> <p><strong>Variables:</strong></p> <pre><code>Variable str Description ---------------------------------------------------------------------------------- id (integer64) - Tweet identifier conversation_id (integer64) - Tweet conversation identifier date (POSIXct) - Tweet created date (format: YYYY-MM-DD hh:mm:ss) tweet (chr) - Symptoms mention about COVID-19 language (chr) - Tweet language: Portuguese hashtags (chr) - Sign (#) used to identify specific topic user_id (integer64) - User identifier username (chr) - Twitter user name link (chr) - Tweet url urls (chr) - External urls from tweet photos (chr) - Photos posted in message (link) video (int) - Video posted in message (1=True;0=False) thumbnail (chr) - Thumbnail posted in message retweet (logi) - Message reposted by another user nlikes (int) - Number of tweet likes nreplies (in) - Number of tweet replies nretweets (int) - Number of tweet retweets Near (logi) - Near a certain City (Example: London) geo (logi) - Geo coordinates (lat,lon,km/mi.) user_rt_id (logi) - User retweet identifier user_rt (logi) - Retweet user retweet_id (logi) - Retweet identifier reply_to (chr) - Answer to someone retweet_date (logi) - Retweet created date (format: YYYY-MM-DD hh:mm:ss) symptoms (chr) - Symptoms mentioned nsymptoms (int) - Number of symptons mentioned </code></pre> <p><em>str: Compactly Display the Structure of an Arbitrary R Object</em></p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Dataset of suicidal ideation texts in Brazilian Portuguese - Boamente System

<p>We obtained non-clinical texts from tweets (user posts of the online social network Twitter). To find suicide-related tweets, we used the Twitter API to download tweets in a personalized way based on search terms associated with suicide. After different experiments to retrieve relevant texts, 5699 tweets were collected in May 2021. Each downloaded tweet had user-specific information (for example, user ID, timestamp, language, location, number of likes, etc.). Still, we kept only the post content (suicide-related texts) and discarded the additional data. Therefore, all texts were anonymized.&nbsp;</p><p>After data collection, three psychologists were invited to perform the data annotation, in which they individually labeled each tweet. To avoid bias in the annotation process, we selected psychologists with different psychological approaches, namely cognitive behavioral theory, psychoanalytic theory, and humanistic theory. Professionals had to classify each tweet as negative for suicidal ideation (annotated as 0), or positive for suicidal ideation (annotated as 1).&nbsp;</p><p>All tweets with at least one divergence between psychologists (n = 1513) were excluded, resulting in a dataset with 4186 instances. 398 duplicate tweets were excluded. The final dataset consists of 2691 instances labeled negative and 1097 labeled positive.</p><p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

A Dataset of Polarities and Emotions from Brazilian Portuguese Play Store Reviews

<p>User reviews play a crucial role in shaping consumer perceptions and guiding decision-making processes in the digital marketplace. With the rise of mobile applications, platforms like the Google Play Store serve as hubs for users to express their opinions and experiences with various apps and services. Understanding the polarities and emotions conveyed in these reviews provides valuable insights for developers, marketers, and researchers alike.</p> <p>The dataset consists of user reviews collected from the "Trending" section of the Google Play Store in May 2023. A total of 300 reviews were gathered for each of the top 10 most downloaded applications during this period. Each review in the dataset has been meticulously labeled for polarity, categorizing sentiments as positive, negative, or neutral, and emotion, encompassing a range of emotional responses such as happiness, sadness, surprise, fear, disgust and anger.</p> <p>Additionally, it's worth noting that this dataset underwent a rigorous annotation process. Three annotators independently classified the reviews for polarity and emotion. Afterward, they reconciled any discrepancies through discussion and arrived at a consensus for the final annotations. This ensures a high level of accuracy and reliability in the labeling process, providing researchers and practitioners with trustworthy data for analysis and decision-making.</p> <p>It's important to highlight that all reviews in this dataset are in Brazilian Portuguese, reflecting the specific linguistic and cultural nuances of the Brazilian market. By leveraging this dataset, stakeholders gain access to a robust resource for exploring user sentiment and emotion within the context of popular mobile applications in Brazil.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Refinement dataset SPIRA

<p>This dataset was collected over the internet and in hospital wards with the goal of detecting respiratory insufficiency (typically caused by COVID-19). This data collection is part of the SPIRA Project, whose goal is developing a system for recognizing respiratory insufficiency through speech analysis. The datasets presented here were used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger.</p> <p>The spira_trimmed_data file contains the original ~1 hour dataset collected over the internet (control) and in hospital wards (patients) by the SPIRA Project. This is as described in the paper: Deep learning against COVID-19: Respiratory insufficiency detection in Brazilian Portuguese Speech. We include it here for completeness.</p> <p>The spira_control_full_mp3 file contains the complete ~18 hours control data collected over the internet by the SPIRA project. While not useful for respiratory insufficiency detection, the dataset may be used for identifying age and gender as we mention in our paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Pretraining Datasets raw audios from CORAA

<p>This repository contains all the pretraining datasets used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger. These datasets are part of a collection of datasets from the TaRSila project (see https://sites.google.com/view/tarsila-c4ai). The audios published here were in part also published with annotations and transcriptions as the CORAA dataset (see https://github.com/nilc-nlp/CORAA). Here we publish the original raw audios from the following datasets (without transcriptions) - ALIP, C-Oral, SP2010, NURC-Recife, NURC-S&atilde;o Paulo and Programa Certas Palavras. In total, the datasets contain about 800 hours of Brazilian Portuguese Speech.</p> <p>The audios have been converted to mp3 to facilitate the upload. ALIP, C-Oral and SP2010 are integrally contained in one file each. Programa Certas Palavras and NURC-Recife are split in 3 parts each, while NURC-SP is split in 7 parts of roughly equal size. More information on the datasets can be found in the paper Acoustic models of Brazilian Portuguese Speech based on Neural Transformers as well as on the original references which created these datasets.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Data for "Is Stack Overflow in Portuguese attractive for Brazilian Users?"

<p>Data for Botto-Tobar et al.&nbsp;Is Stack Overflow in Portuguese attractive for Brazilian Users?.&nbsp;ICGSE 2018.</p> <p>This data&nbsp;was built based on data dump from Stack Exchange (https://stackexchange.com) website.&nbsp;It contains two separate databases (Stack Overflow in English&nbsp;and Stack Overflow in Portuguese):</p> <ul> <li>Users</li> <li>Posts, decomposed by Answers and Questions</li> <li>Tags</li> <li>PostTags</li> <li>GenderUser</li> <li>UserLocation</li> </ul> <p>For more information, please visit&nbsp;http://www.win.tue.nl/~mbottoto/files/papers/conference_papers/sopt_icgse2018.pdf or write to&nbsp;<em>m.a.botto.tobar@tue.nl</em></p> <p>&nbsp;</p>

opencc-by-4.0May 2018View details →
zenodo40/100

COVID19.BR: A Dataset of Misinformation about COVID-19 in Brazilian Portuguese WhatsApp Messages

<p>COVID19.BR is provided in a csv&nbsp;file where the columns are date, hour, phone number, international phone code, if the user is Brazilian its state, the text content of the message, word count, character count, and if the message contained media (audio, image, or video). Each row represents a WhatsApp message.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

THLS - An open source dataset for Brazilian Portuguese speech processing

<p>THLS Open Source Brazilian Portuguese Speech Dataset with 1000 sentences balanced phonetically.</p> <p>&nbsp;</p> <p>Authors:</p> <ul> <li>Luiz Felipe Vecchietti</li> <li>Thalles Melo Batista Pieroni (Voice)</li> </ul> <p>&nbsp;</p> <p>https://gitlab.com/lfelipesv/1000-sentences-thls-dataset</p>

opencc-by-4.0Dec 2023View details →
ClinicalTrials.gov32/100

Adaptation of the One-session CBT Protocol for Prevention of Mental Illness for Brazilian Portuguese

ClinicalTrials.gov study NCT03518411. IPD Sharing: UNDECIDED. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo12/100

Brazilian Portuguese CHV

<p>Part of Brazilian Portuguese CHV&nbsp;</p>

restrictedOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record