Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
201
datasets available to search
ShareScore release 0.9.0
Dataset results
201 results for “portuguese”
Data for "Is Stack Overflow in Portuguese attractive for Brazilian Users?"
<p>Data for Botto-Tobar et al. Is Stack Overflow in Portuguese attractive for Brazilian Users?. ICGSE 2018.</p> <p>This data was built based on data dump from Stack Exchange (https://stackexchange.com) website. It contains two separate databases (Stack Overflow in English and Stack Overflow in Portuguese):</p> <ul> <li>Users</li> <li>Posts, decomposed by Answers and Questions</li> <li>Tags</li> <li>PostTags</li> <li>GenderUser</li> <li>UserLocation</li> </ul> <p>For more information, please visit http://www.win.tue.nl/~mbottoto/files/papers/conference_papers/sopt_icgse2018.pdf or write to <em>m.a.botto.tobar@tue.nl</em></p> <p> </p>
Text-fig. 1 Simplified geological map of NW Portugal (adapted from Portuguese Geological Survey 1:500.000 lithological map). in Overview Of The Stratigraphy And Initial Quantitative Biogeographical Results From The Devonian Of The Albergaria-A-Velha Unit (Ossa-Morena Zone, W Portugal)
Text-fig. 1 Simplified geological map of NW Portugal (adapted from Portuguese Geological Survey 1:500.000 lithological map).
Linked collectors and determiners for: Portuguese Decapoda Crustacea - Museu Nacional de História Natural e da Ciência collections.
Natural history specimen data linked to collectors and determiners held within, "Portuguese Decapoda Crustacea - Museu Nacional de História Natural e da Ciência collections". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/c00d00cf-db3a-4b78-89f1-12dd20abf4c1">https://bionomia.net/dataset/c00d00cf-db3a-4b78-89f1-12dd20abf4c1</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/c00d00cf-db3a-4b78-89f1-12dd20abf4c1">https://gbif.org/dataset/c00d00cf-db3a-4b78-89f1-12dd20abf4c1</a>. Formatted as a Frictionless Data package.
COVID19.BR: A Dataset of Misinformation about COVID-19 in Brazilian Portuguese WhatsApp Messages
<p>COVID19.BR is provided in a csv file where the columns are date, hour, phone number, international phone code, if the user is Brazilian its state, the text content of the message, word count, character count, and if the message contained media (audio, image, or video). Each row represents a WhatsApp message.</p>
paperChain - Video Animation - Circular Case 1 (Portugal) - Portuguese version
<p>This animation video summarises in approximately two minutes the most relevant results of the implementation of paperChain's Circular Case.</p>
Text-fig. 2. Portuguese Early Cretaceous formations in the Torres Vedras and in the Figueira da Foz region. The Torres Vedras mesofossil flora was collected in the lower Almargem Formation. The Buarcos mesofossil flora, and several other younger mesofossil floras from Portugal, are from the lower part of the Figueira da Foz Formation. Based on information in Rey et al. (2006). in The Early Cretaceous Mesofossil Flora Of Torres Vedras (Ne Of Forte Da Forca), Portugal: A Palaeofloristic Analysis Of An Early Angiosperm Community
Text-fig. 2. Portuguese Early Cretaceous formations in the Torres Vedras and in the Figueira da Foz region. The Torres Vedras mesofossil flora was collected in the lower Almargem Formation. The Buarcos mesofossil flora, and several other younger mesofossil floras from Portugal, are from the lower part of the Figueira da Foz Formation. Based on information in Rey et al. (2006).
Project Gutenberg Self-Publishing Press: Portuguese Books
<p>A collection of 40 books written in Portuguese and published at <a href="http://self.gutenberg.org/">Project Gutenberg Self-Publishing - eBooks | Read eBooks online | Free eBooks</a>.</p> <ul> <li>The <em>gutenberg_selfpub_tagged.vrt</em> file contains all the books. The texts were tagged using Spacy large trained model for Portuguese (<a href="https://spacy.io/models/pt">https://spacy.io/models/pt</a>)</li> <li>The corpus metadata is listed in <em>gutenberg_selfpub_metadata.tsv</em></li> <li><em>gutenberg_selfpub_xml_untagged.zip</em> contains all the untagged texts.</li> </ul> <p>The XML files contain the following tags and attributes:</p> <ul> <li>chapter: n</li> <li>dedication</li> <li>text:id</li> <li>title</li> <li>author</li> <li>subtitle</li> <li>part: n</li> <li>acknowledge</li> <li>summary: lang</li> <li>publisher</li> <li>short: n</li> <li>biography</li> </ul>
Does Taylor approximation really works? (Portuguese version)
<p>Didactic video that shows how Taylor approximation can be used to approximate the exponential function. It is illustrated that the approximation becomes more accurate as the polynomial degree increases.</p>
Fig. 4. Scombrolabrax heterolepis Roule, 1921 in Updated checklist of marine fishes (Chordata: Craniata) from Portugal and the proposed extension of the Portuguese continental shelf
Fig. 4. Scombrolabrax heterolepis Roule, 1921.
Fig. 3 in Updated checklist of marine fishes (Chordata: Craniata) from Portugal and the proposed extension of the Portuguese continental shelf
Fig. 3. Fistularia petimba Lacepède, 1803.
Fig. 2 in Updated checklist of marine fishes (Chordata: Craniata) from Portugal and the proposed extension of the Portuguese continental shelf
Fig. 2. Bajacalifornia megalops (Lütken, 1898).
Portuguese kiosk
This is a kiosk I made for a university project inspired by the ones we have here in Portugal. Source: Objaverse 1.0 / Sketchfab
The Portuguese Large Wildfire Spread Database (PT-FireSprd)
<p>The <strong>Portuguese Large Wildfire Spread Database (PT-FireSprd v2.0 ) </strong>includes the reconstruction of the spread of 155 large wildfires that occurred in Portugal between 2015 and 2024. It includes a detailed set of fire behaviour descriptors, such as rate-of-spread (ROS), fire spread direction, fire growth rate (FGR) and fireline intensity.</p> <p>The wildfires were reconstructed by converging evidence from complementary data sources, such as satellite imagery/products, airborne and ground data collected by fire personnel, official fire data and information in external reports. We then implemented a digraph-based algorithm to estimate the fire behaviour descriptors. Fireline intensity was estimated using Byram's equation assuming full fuel load consumption as provided by national level fuel maps.</p> <p><strong>PT-WFireSprd database is organized in 3 levels: </strong></p> <ul> <li> <p><strong>L1: Wildfire Progression, </strong>representing the spatial and temporal evolution of the wildfire spread (i.e. where and when).</p> </li> <li> <p><strong>L2: Wildfire Behavior,</strong> including quantitative behavior descriptors of how a wildfire burned, such as the rate-of-spread (ROS), fire growth rate (FGR) and fire line intensity (FLI)</p> </li> <li> <p><strong>L3: Simplified Wildfire Behavior,</strong> averaging fire behavior over longer periods that represent relatively homogenous fire runs.</p> </li> </ul> <p>The data from the different levels is composed by a large set of maps that can be useful for several applications and target users.</p> <p>The PT-FireSprd is the first open access fire progression and behaviour database in Mediterranean Europe, dramatically expanding the extant information. Updating the PT-FireSprd database will require a continuous joint effort by researchers and fire personnel.</p>
ALIMENTOPIA Project - CETAPS The Portuguese Vegetarian Database
<pre>One of the main goals of the ALIMENTOPIA/Utopian Foodways Project (2016-2020) was to rescue from the collective amnesia an unprecedented phenomenon in Portugal in the early 20th century: the Portuguese Vegetarian Society. As a priceless cultural artifact, <em>O Vegetariano: Mensário Naturista Ilustrado</em>, published uninterruptedly between 1909 and 1935, was more than just the official organ of the Vegetarian Society of Portugal (SVP), founded in 1911. It was indeed an aggregating platform for a community that reached, at its peak, more than 3,500 subscribers. All the entries of all the publications are indexed in this searchable database, making it a useful tool for any researcher working on Food, Utopian movements, vegetarianism, naturism and Portuguese History.</pre>
Meat Quality from Two Portuguese Autochthonous Lamb Breeds
<p>Recording of the talk “Meat quality from two Portuguese autochthonous lamb breeds”, presented by Ursula Gonzales-Barron at the 71<sup>st</sup> Annual Meeting of the European Federation of Animal Science, EAAP, Online virtual meeting (1-4 Dec 2020).</p>
A Survey of Pathogens on Lamb Carcasses from Portuguese Local Breeds
<p>Recording of the talk “A survey of pathogens on lamb carcasses from Portuguese local breeds”, presented by Ursula Gonzales-Barron at the 71<sup>st</sup> Annual Meeting of the European Federation of Animal Science, EAAP, Online virtual meeting (1-4 Dec 2020).</p>
PT2vec - A Portuguese text corpus created from online newspapers
<p>A Portuguese text corpus created from online newspapers with a total of 394,825,480 tokens and 33,089,734 sentences.</p> <p>If you use this repository, please cite this paper:</p> <p>Pinto JP, Viana P, Teixeira I, Andrade M. 2022. Improving word embeddings in Portuguese: increasing accuracy while reducing the size of the corpus. PeerJ Computer Science 8:e964 <a href="https://doi.org/10.7717/peerj-cs.964">https://doi.org/10.7717/peerj-cs.964</a></p>
English and Portuguese CBOW Models from Europarl Corpus, version 7, using FastText with Subwords Option
<p>The models were trained using FastText, model CBOW, 40 epochs, and subwords. Each *.BIN file has its *.VEC file with the vocabulary ordered by frequency. The *.BIN file can return a vector to represent an out-of-vocabulary (OOV) word if the necessary parts of the OOV word were used in training. FastText and Gensim can use these files. The English and Portuguese models are identified in the file name, <strong>_en_</strong> and <strong>_pt_</strong> respectively.</p> <p>An Excel file has the neighborhood changes of some selected words during training on each epoch. A previous exercise to find words with more than one meaning.</p>
2023 Portuguese Energy Tariffs: Retail and Feed-In Tariff Data
<p><strong>Introduction</strong></p> <p>This dataset contains historic data for 20 electric grid consumption and energy injection tariffs in Portugal, during 2023. The data includes both retail tariffs for grid consumption energy and feed-in income tariffs for energy injected back into the grid.</p> <p> </p> <p><strong>Data Overview</strong></p> <p>The dataset comprises the following key elements:</p> <p> 1. Grid Consumption (Retail Tariffs):</p> <ul> <li>Retail tariffs refer to the cost per unit of electricity consumed from the grid.</li> <li>Tariff values are given in euro (€) per kWh and vary across different retailers and hour schemes.</li> </ul> <p> 2. Energy Injection (Feed-in Tariffs):</p> <ul> <li>Feed-in tariffs represent the compensation per unit of electricity generated by solar systems and fed back into the grid.</li> <li>Tariff values are presented in € per kWh and vary across different retailers and conditions.</li> </ul> <p><br><strong>Assumptions</strong></p> <p><br>The retail tariffs in this dataset refer to the consumed energy only and do not include taxes, contracted power costs, and other fees.</p> <p>The dataset starts on Jan 1st, 2023, and ends on Dec 31st, 2023.</p> <p>The dataset uses a 15min time step for each line. The price per hour is applied to each of the four 15 minutes steps.</p> <p><br><strong>Reference data used to create this dataset:</strong></p> <ul> <li>ERSE Regulatory data simulator - https://simulador.precos.erse.pt/eletricidade/</li> <li>2023 Energy Cost Description - https://www.erse.pt/en/activities/market-regulation/tariffs-and-prices-electricity</li> <li>GALP - https://casa.galp.pt/planos-eletricidade-e-gas</li> <li>Eni Plenitude - https://eniplenitude.pt/eletricidade</li> <li>Repsol - https://www.repsol.pt/particulares/casa/eletricidade-gas/</li> <li>EDP - https://www.edp.pt/empresas/energia/tarifarios/</li> </ul> <p> </p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications.</p>
Autores portugueses impressos no estrangeiro no século XVI
<p>Fazemos uma viagem ao século XVI com João Alves Dias e João Costa para acompanhar as suas descobertas sobre os autores portugueses então editados fora de Portugal, em português e outras línguas, em particular na República de Veneza, Lyon, Salamanca, Antuérpia e Colónia. Serão cerca de quatro mil livros, além de outros tipos de textos. Trata-se de obras de medicina, catecismo e sermões, literatura, filosofia, direito, matemática, gramática e histórias de Portugal, entre outros. Filipe Dias, Fernão Mendes Pinto e Aquiles Estácio são exemplos entre os muitos autores impressos no estrangeiro.</p> <p><br>João Alves Dias é investigador integrado do CHAM e Professor Auxiliar com agregação na NOVA FCSH. Director da revista Fragmenta Historica, os seus interesses de investigação centram-se na história do livro impresso, história de Portugal Moderno, a Paleografia e Diplomática. João Costa é doutorado em História Medieval e investigador integrado do CHAM. É investigador contratado dos National Library and Archives de Abu Dhabi.</p> <p> </p> <p>A entrevista é conduzida por José Alberto Catalão.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.