Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
201
datasets available to search
ShareScore release 0.9.0
Dataset results
201 results for “portuguese”
Mesiodistal distance of portuguese population
<p>Mesiodistal distance of portuguese population</p>
CATCH-EyoU: Educating Critical Citizens? Portuguese Teachers and Students' Visions about Critical Thinking in School.
<p>The dataset integrates the perspectives on critical thinking in school of Portuguese teachers (in total, 16 interviews) and students (in total, 8 focus groups).</p>
VeLePor: European Portuguese Verbal Paradigms in Phonemic Notation
<p>This is a collection of European Portuguese verbal paradigms, in phonemic notation. They are suited for both computational and manual analysis.</p>
Portuguese DBnary archive in original Lemon format
<p>The DBnary dataset is an extract of Wiktionary data from many language editions in RDF Format. Until July 1st 2017, the lexical data extracted from Wiktionary was modeled using the lemon vocabulary.</p> <p>This dataset contains the full archive of all DBnary dumps in Lemon format containing lexical information from Portuguese language edition, ranging from 26th August 2012 to 1st July 2017.</p> <p>After July 2017, DBnary data has been modeled using the ontolex model and will be available in another Zenodo entry.<br> </p>
emoUERJ: an emotional speech database in Portuguese
<p><strong>GOAL</strong></p> <p>Since language is a key issue in speech emotion recognition (SER) and there are few databases in Portuguese, this database was developed at the State University of Rio de Janeiro aiming to the development of specific SER models for this language.</p> <p> </p> <p><strong>DATABASE DESIGN</strong></p> <p>Ten sentences were made available to eight actors, equally divided between genders, and they were free to choose the phrases for record audios in four emotions target: happiness, anger, sadness or neutral. The following phrases were used:</p> <ul> <li>Não importa quem está certo. (It doesn't matter who is right.)</li> <li>Você perde tempo demais com a Internet. (You waste too much time on the Internet.)</li> <li>A garrafa está na geladeria. (The bottle is in the fridge.)</li> <li>Eu estou me sentindo doente hoje. (I'm feeling sick today)</li> <li>Eu estou um pouco atrasado. (I'm a little late)</li> <li>Nos fins de semana, eu sempre ia para a casa dele(a). (On weekends, I always used to go to his/her house)</li> <li>De quem são essas malas que estão debaixo da mesa? (Whose bags are under the table?)</li> <li>Ele volta na quarta-feira. (He comes back on wednesday)</li> <li>Já chega! Eu vou tomar um banho e ir para a cama. (Enough! I'm going to take a shower and go to bed)</li> <li>Você poderia arrumar a mesa, por favor? (Could you set the table, please?)</li> </ul> <p>The result of this process was 377 audios distributed as follows</p> <ul> <li>happiness: 91</li> <li>anger: 94</li> <li>sadness: 100</li> <li>neutral: 92</li> </ul> <p><strong>FILE IDENTIFICATION</strong></p> <p>Each database file corresponds to a phrase recorded by an actor expressing one of the four emotions and was named as follows:</p> <ul> <li>Position 1: actor's gender ('m' for man or 'w' for woman)</li> <li>Positions 2 and 3: actor's id (from 01 to 04)</li> <li>Position 4: emotion (h: happiness, a: anger, s: sadness, n: neutral)</li> <li>Positions 5 and 6: recording identification</li> </ul> <p>For example, the file 'w04a11' was the eleventh audio recorded by actress 04 interpreting the anger emotion.</p>
Interview guide from the paper "Act Against the Aggression A Preliminary Model of a Digital Platform Aimed at Cyberaggressions in the Portuguese University Environment"
<p>Interview guide used in the study "Act Against the Aggression A Preliminary Model of a Digital Platform Aimed at Cyberaggressions in the Portuguese University Environment".</p>
ESTER-Pt: An Evaluation Suite for TExt Recognition in Portuguese
<p><em><strong>disclaimer</strong></em>: Version accepted as full paper in ICDAR 2023.</p> <p>Optical Character Recognition (OCR) is a technology that enables machines to read and interpret printed or handwritten texts from scanned images or photographs. However, the accuracy of OCR systems can vary depending on several factors, such as the quality of the input image, the font used, and the language of the document. As a general tendency, OCR algorithms perform better in resource-rich languages as they have more annotated data to train the recognition process. We propose ESTER-Pt, an Evaluation Suite for TExt Recognition in Portuguese in this work. Despite being one of the largest languages in terms of speakers, OCR in Portuguese remains largely unexplored. Our evaluation suite comprises four types of resources: synthetic text-based documents, synthetic image-based documents, real scanned documents, and a hybrid set with real image-based documents that were synthetically degraded.</p>
[Free] Portuguese Mandolin #02
New mandolin from my 3d workshop. Made few months ago but finally textured....I reused some of the parts from #01 and made the rest of the body based on portuguese style mandolins. This one comes with gold hardware, walnut body, neck, head and spruce sound board. All ornaments are custom made for this pupouse. -Modeled and unwrapped in 3dsmax, textured in Substance Painter and Photoshop. Source: Objaverse 1.0 / Sketchfab
ICF portuguese codebook for Atlas.ti
<p>This ICF (International Classification of Functioning, Disability and Health) codebook was created by the author for the qualitative analysis of health documents and can be used as a basis for content analysis and ICF linking rules for measurement instruments used in clinical health routine. It was structured to be used in the Atlas.ti qualitative analysis software.</p>
Ontolex-lemon and TIAD versions of Apertium Portuguese-Catalan dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2019-2020, Hèctor Alòs i Font 2016, Bror Hultberg 2008-2016, Francis M. Tyers 2012, Anthony J. Bentley 2009, Pasquale Minervini 2009, Mireia Ginestí Rosell 2008, Jim O'Regan 2008, Carme Armentano-Oller 2008, Mikel L. Forcada
Ontolex-lemon and TIAD versions of Apertium Portuguese-Galician dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2007--2008, imaxin|software (http://www.imaxin.com) (c) 2005--2008, Universitat d'Alacant (Transducens group)
Ontolex-lemon and TIAD versions of Apertium Spanish-Portuguese dictionary
<p>OntoLex-lemon and TSV conversion of Apertium Bidix. For more details, see <a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2007-2013, Universitat d'Alacant (Grup Transducens) 2007-2013, Francis M. Tyers 2007-2013, Mikel L. Forcada 2007-2013, Sergio Ortiz 2007-2014, Gema Ramírez Sánchez 2007, Oscar Senra Gómez 2008-2011, Jim O'Regan 2008, Mireia Ginestí Rosell 2009, Pminervini 2010, Carlos Ramisch 2012-2014, Kevin Brubeck Unhammer 2012, Anthony J. Bentley 2013-2014, Miriam Antunes Gonçalves 2013, Filip Petkovski 2014, Alicia Nieva
Portuguese Comparative Sentences: A Collection of Labeled Sentences on Twitter and Buscapé
<p>More and more customers demand online reviews of products and comments on the Web to make decisions about buying a product over another. In this context, sentiment analysis techniques constitute the traditional way to summarize a user’s opinions that criticizes or highlights the positive aspects of a product. Sentiment analysis of reviews usually relies on extracting positive and negative aspects of products, neglecting comparative opinions. Such opinions do not directly express a positive or negative view but contrast aspects of products from different competitors. </p> <p>Here, we present the first effort to study comparative opinions in Portuguese, creating two new Portuguese datasets with comparative sentences marked by three humans. This repository consists of three important files: (1) lexicon that contains words frequently used to make a comparison in Portuguese; (2) Twitter dataset with labeled comparative sentences; and (3) Buscapé dataset with labeled comparative sentences.</p> <p>The lexicon is a set of 176 words frequently used to express a comparative opinion in the Portuguese language. In these contexts, the lexicon is aggregated in a filter and used to build two sets of data with comparative sentences from two important contexts: (1) Social Network Online; and (2) Product reviews.</p> <p>For Twitter, we collected all Portuguese tweets published in Brazil on 2018/01/10 and filtered all tweets that contained at least one keyword present in the lexicon, obtaining 130,459 tweets. Our work is based on the sentence level. Thus, all sentences were extracted and a sample with 2,053 sentences was created, which was labeled for three human manuals, reaching an 83.2% agreement with Fleiss' Kappa coefficient. For Buscapé, a Brazilian website (https://www.buscape.com.br/) used to compare product prices on the web, the same methodology was conducted by creating a set of 2,754 labeled sentences, obtained from comments made in 2013. This dataset was labeled by three humans, reaching an agreement of 83.46% with the Fleiss Kappa coefficient.</p> <p>The Twitter dataset has 2,053 labeled sentences, of which 918 are comparative. The Buscapé dataset has 2,754 labeled sentences, of which 1,282 are comparative.</p> <p><strong>The datasets contain these labeled properties:</strong></p> <ul> <li> <p><em>text</em>: the sentence extracted from the review comment.</p> </li> <li> <p><em>entity_s1: </em>the first entity compared in the sentence<em>.</em></p> </li> <li> <p><em>entity_s2: </em>the second entity compared in the sentence.</p> </li> <li> <p><em>keyword: </em>the comparative keyword used in the sentence to express comparison.</p> </li> <li> <p><em>preferred_entity: </em>the preferred entity.</p> </li> <li> <p><em>id_start: </em>the keyword's initial position in the sentence.</p> </li> <li> <p><em>id_end: </em>the keyword's final position in the sentence.</p> </li> <li> <p><em>type: </em>the sentence label, which specifies whether the phrase is a comparison<em>.</em></p> </li> </ul> <p><strong>Additional Information:</strong></p> <p><em><strong>1 </strong>- </em>The sentences were separated using a sentence tokenizer.</p> <p><em><strong>2 </strong>- </em>If the compared entity is not specified, the field will receive a value: "__".</p> <p><em><strong>3</strong> - </em>The property <em>"type"</em> can contain five values, they are:</p> <ul> <li> <p><em>0: Non-comparative (</em>Não Comparativa<em>)</em>.</p> </li> <li> <p><em>1: Non-Equal-Gradable (</em>Gradativa com Predileção<em>)</em>.</p> </li> <li> <p><em>2: Equative</em> <em>(</em>Equitativa<em>).</em></p> </li> <li> <p><em>3: Superlative (</em>Superlativa<em>).</em></p> </li> <li> <p><em>4: Non-Equal-Gradable (</em>Não Gradativa<em>).</em></p> </li> </ul> <p> </p> <p>If you use this data, please cite our paper as follows: </p> <p><em>"Daniel Kansaon, Michele A. Brandão, Julio C. S. Reis, Matheus Barbosa,Breno Matos, and Fabrício Benevenuto. 2020. Mining Portuguese Comparative Sentences in Online Reviews. In Brazilian Symposium on Multimedia and the Web (WebMedia ’20), November 30-December 4, 2020, São Luís, Brazil. ACM, New York, NY, USA, 8 pages. https://doi.org/10.1145/3428658.3431081"</em></p> <p>--------------</p> <p><strong>Plus Information:</strong></p> <p>We make the raw sentences available in the dataset to allow future work to test different pre-processing steps. Then, if you want to obtain the exact sentences used in the paper above, you must reproduce the pre-processing step described in the paper (<em>Figure 2</em>). </p> <p>For each sentence with more than one keyword in the dataset: </p> <ul> <li>You need to extract three words before and three words after the comparative keyword, creating a new sentence that will receive the existing value in the “<em>type</em>” field as a label;</li> <li>The original sentence will be divided into <em>n</em> new sentences. (<em>n</em>) is the number of keywords in the sentence;</li> <li>The stopwords should not be accounted for as part of this range (<em>3 words</em>);</li> </ul> <p>Note that: the final processed sentence can have more than six words because the stopwords are not counted as part of the range.</p>
16th C Portuguese Coat of Arms, Montemor-o-Novo
This piece belongs to the collection of the Museum of São Domingos, in Montemor-o-Novo, carved in limestone, where in the center are the arms of King D. Manuel, between two armillary spheres. At the bottom is the following inscription: ESTAS ARMAS MADOU AQUI POR FRCO FARZAO, JUIZ DE FORA EM ESTA VILA NO PRIMEIRO DIA DE JUNHO ERA DE MIL E D E XI ANOS. Source: Objaverse 1.0 / Sketchfab
FIGURE 2 in On the diversity of the genus Pisione (Polychaeta, Pisionidae) along the Portuguese continental shelf, with a key to European species
FIGURE 2. Classification (A) and Principal Coordinates Analysis (B) based on morphological descriptors of the Pisione species occurring in the Portuguese continental shelf. Descriptors are represented as vectors. W10—width at chaetiger 10; CP2/ CP3—ratio between the length of the dorsal cirri of parapodia 2 (CP2) and parapodia 3 (CP3); nrT—number of teeth of the supra-acicular chaetae; P1—protruding length of the notoaciculae; IA—presence/absence of infra-acicular simple chaetae.
FIGURE 1 in On the diversity of the genus Pisione (Polychaeta, Pisionidae) along the Portuguese continental shelf, with a key to European species
FIGURE 1. Distribution and relative abundance of Pisione species along the Portuguese continental shelf (northeastern Atlantic).
FIGURE 6 in Lumbrineridae (Polychaeta) from the Portuguese continental shelf (NE Atlantic) with the description of four new species
FIGURE 6. Ordination analysis based on morphological descriptors of specimens of Abyssoninoe, Lumbrineris, Gallardoneris, Lumbrinerides, Lumbrineriopsis and Ninoe species (A) and of Lumbrineris luciliae sp. nov., L. lusitanica sp. nov. and L. pinaster sp. nov. (B). The most correlated variables (rho>0.8) are shown as dashed vectors. Legend: A.P.L.—postchaetal lobe shape in anterior parapodia; CMHH—composite multidentate hooded hook; SMHH—simple multidentate hooded hook; SBHH—simple bidentate hooded hook; MI attach. lam.—MI attachment lamellae; MIII unid. + knob – MIII unidentate followed by a knob; prominent proj. MIII—prominent projection in the basal part of MIII; MIV unid. + dev. plate—MIV unidentate with a developed plate; MIV unid. + pointed tooth—MIV unidentate with a pointed tooth; W10—width at chaetiger 10 excluding parapodia.
FIGURE 5 in Lumbrineridae (Polychaeta) from the Portuguese continental shelf (NE Atlantic) with the description of four new species
FIGURE 5. Lumbrineris pinaster sp. nov. Paratype (ECOSUR0131) A, anterior end, dorsal view; B, parapodium 3, frontal view; C, parapodium 13, frontal view; D, parapodium 153, frontal view; E, composite multidentate hooded hook, from parapodium 3; F, simple multidentate hooded hook with long hood, from parapodium 13; G, preacicular simple multidentate hooded hook with short hood, from parapodium 79; H, postacicular simple multidentate hooded hook with short hood, from parapodium 79; I, maxillae III and IV, dorsal view. Scale bars: A, 0.4 mm; B, C, 0.5 mm; D, I, 0.025 mm; E, F, 0.012 mm.
FIGURE 1 in Lumbrineridae (Polychaeta) from the Portuguese continental shelf (NE Atlantic) with the description of four new species
FIGURE 1. Study area: the Portuguese continental shelf. Grey dots correspond to sampling sites where Lumbrineridae specimens were found.
FIGURE 3 in Lumbrineridae (Polychaeta) from the Portuguese continental shelf (NE Atlantic) with the description of four new species
FIGURE 3. Lumbrineris luciliae sp. nov. Paratype (ECOSUR0129) A, anterior end, dorsal view; B, parapodium 3, frontal view; C, parapodium 13, frontal view; D, parapodium 77, frontal view; E, composite multidentate hooded hook, from parapodium 3; F, simple multidentate hooded hook from parapodium 77; G, acicula from parapodium 86; H, maxillae III and IV, dorsal view. Scale bars: A, 1.0 mm; B, C, D, H 0.1 mm; E, F, 0.025 mm; G, 0.01 mm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.