Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

566

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

566 results for “analytics”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset for publication "Multi-phase quantitative compositional mapping by LA-ICP-MS: analytical approach and data reduction protocol implemented in XMapTools"

<p>Datasets for the publication &quot;Multi-phase quantitative compositional mapping by LA-ICP-MS: analytical approach and data reduction in XMapTools&quot;</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

BIObec's social media analytics

<p>In this folder can be found all the LinkedIn and X (formerly Twitter) analytics from the BIObec project's accounts, since beginning of the project.&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Hate Speech and Bias against Asians, Blacks, Jews, Latines, and Muslims: A Dataset for Machine Learning and Text Analytics

<h1>Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University Dataset on bias against Asians, Blacks, Jews, Latines, and Muslims&nbsp;</h1> <div> <h2>&nbsp;</h2> <h2>Description&nbsp;</h2> </div> <div> <p>The dataset is a product of a research project at Indiana University on biased messages on Twitter against ethnic and religious minorities. We scraped all live messages with the keywords "Asians, Blacks, Jews, Latinos, and Muslims" from the Twitter archive in 2020, 2021, and 2022.</p> <p>Random samples of 600 tweets were created for each keyword and year, including retweets. The samples were annotated in subsamples of 100 tweets by undergraduate students in Professor Gunther Jikeli's class 'Researching White Supremacism and Antisemitism on Social Media' in the fall of 2022 and 2023. A total of 120 students participated in 2022. They annotated datasets from 2020 and 2021. 134 students participated in 2023. They annotated datasets from the years 2021 and 2022. The annotation was done using the <a href="https://annotationportal.com/" target="_blank" rel="noreferrer noopener">Annotation Portal</a> (Jikeli, Soemer and Karali, 2024). The updated version of our portal, <a href="https://portal2.annotationportal.com/" target="_blank" rel="noreferrer noopener">AnnotHate</a>, is now publicly available. Each subsample was annotated by an average of 5.65 students per sample in 2022 and 8.32 students per sample in 2023, with a range of three to ten and three to thirteen students, respectively. Annotation included questions about bias and calling out bias.&nbsp;&nbsp;</p> </div> <div> <p>Annotators used a scale from 1 to 5 on the bias scale (confident not biased, probably not biased, don't know, probably biased, confident biased), using definitions of bias against each ethnic or religious group that can be found in the research reports from <a href="https://isca.indiana.edu/publication-research/social-media-project/Research-Report-BIAS-on-Twitter-against-Asians--Blacks-Jews-Latinos-Muslims-final-002.pdf" target="_blank" rel="noreferrer noopener">2022</a> and <a href="https://isca.indiana.edu/documents/BIAS%20Against%20Asian-Black-Hispanic-Jewish-and-%20Muslim-People%20on%20X-Twitter%20in%202021%20and%202022.pdf" target="_blank" rel="noreferrer noopener">2023</a>. If the annotators interpreted a message as biased according to the definition, they were instructed to choose the specific stereotype from the definition that was most applicable. Tweets that denounced bias against a minority were labeled as "calling out bias".&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>The label was determined by a 75% majority vote. We classified &ldquo;probably biased&rdquo; and &ldquo;confident biased&rdquo; as biased, and &ldquo;confident not biased,&rdquo; &ldquo;probably not biased,&rdquo; and &ldquo;don't know&rdquo; as not biased.&nbsp;</p> </div> <div> <p>The stereotypes about the different minorities varied. About a third of all biased tweets were classified as general 'hate' towards the minority. The nature of specific stereotypes varied by group. Asians were blamed for the Covid-19 pandemic, alongside positive but harmful stereotypes about their perceived excessive privilege. Black people were associated with criminal activity and were subjected to views that portrayed them as inferior. Jews were depicted as wielding undue power and were collectively held accountable for the actions of the Israeli government. In addition, some tweets denied the Holocaust. Hispanic people/Latines faced accusations of being undocumented immigrants and "invaders," along with persistent stereotypes of them as lazy, unintelligent, or having too many children. Muslims were often collectively blamed for acts of terrorism and violence, particularly in discussions about Muslims in India.&nbsp;</p> </div> <div> <p>The annotation results from both cohorts (Class of 2022 and Class of 2023) will not be merged. They can be identified by the "cohort" column. While both cohorts (Class of 2022 and Class of 2023) annotated the same data from 2021,* their annotation results differ. The class of 2022 identified more tweets as biased for the keywords "Asians, Latinos, and Muslims" than the class of 2023, but nearly all of the tweets identified by the class of 2023 were also identified as biased by the class of 2022.&nbsp;&nbsp; The percentage of biased tweets with the keyword 'Blacks' remained nearly the same.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>*Due to a sampling error for the keyword "Jews" in 2021, the data are not identical between the two cohorts. The 2022 cohort annotated two samples for the keyword Jews, one from 2020 and the other from 2021, while the 2023 cohort annotated samples from 2021 and 2022.The 2021 sample for the keyword "Jews" that the 2022 cohort annotated was not representative. It has only 453 tweets from 2021 and 147 from the first eight months of 2022, and it includes some tweets from the query with the keyword "Israel". The 2021 sample for the keyword "Jews" that the 2023 cohort annotated was drawn proportionally for each trimester of 2021 for the keyword "Jews".&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Content</h2> <h3>Cohort 2022&nbsp;</h3> </div> <div> <p>This dataset contains 5880 tweets that cover a wide range of topics common in conversations about Asians, Blacks, Jews, Latines, and Muslims. 357 tweets (6.1 %) are labeled as biased and 5523 (93.9 %) are labeled as not biased. 1365 tweets (23.2 %) are labeled as calling out or denouncing bias.&nbsp;&nbsp;</p> </div> <div> <p>1180 out of 5880 tweets (20.1 %) contain the keyword "Asians," 590 were posted in 2020 and 590 in 2021. 39 tweets (3.3 %) are biased against Asian people. 370 tweets (31,4 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1160 out of 5880 tweets (19.7%) contain the keyword "Blacks," 578 were posted in 2020 and 582 in 2021. 101 tweets (8.7 %) are biased against Black people. 334 tweets (28.8 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>1189 out of 5880 tweets (20.2 %) contain the keyword "Jews," 592 were posted in 2020, 451 in 2021, and &ndash;&ndash;as mentioned above&ndash;&ndash;146 tweets from 2022. 83 tweets (7 %) are biased against Jewish people. 220 tweets (18.5 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1169 out of 5880 tweets (19.9 %) contain the keyword "Latinos," 584 were posted in 2020 and 585 in 2021. 29 tweets (2.5 %) are biased against Latines. 181 tweets (15.5 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1182 out of 5880 tweets (20.1 %) contain the keyword "Muslims," 593 were posted in 2020 and 589 in 2021. 105 tweets (8.9 %) are biased against Muslims. 260 tweets (22 %) call out bias against Muslims.&nbsp;&nbsp;</p> </div> <div> <h3>Cohort 2023&nbsp;</h3> </div> <div> <p>The dataset contains 5363 tweets with the keywords &ldquo;Asians, Blacks, Jews, Latinos and Muslims&rdquo; from 2021 and 2022. 261 tweets (4.9 %) are labeled as biased, and 5102 tweets (95.1 %) were labeled as not biased. 975 tweets (18.1 %) were labeled as calling out or denouncing bias.&nbsp;</p> </div> <div> <p>1068 out of 5363 tweets (19.9 %) contain the keyword "Asians," 559 were posted in 2021 and 509 in 2022. 42 tweets (3.9 %) are biased against Asian people. 280 tweets (26.2 %) call out bias against Asians.&nbsp;&nbsp;</p> </div> <div> <p>1130 out of 5363 tweets (21.1 %) contain the keyword "Blacks," 586 were posted in 2021 and 544 in 2022. 76 tweets (6.7 %) are biased against Black people. 146 tweets (12.9 %) call out bias against Blacks.&nbsp;&nbsp;</p> </div> <div> <p>971 out of 5363 tweets (18.1 %) contain the keyword "Jews," 460 were posted in 2021 and 511 in 2022. 49 tweets (5 %) are biased against Jewish people. 201 tweets (20.7 %) call out bias against Jews.&nbsp;</p> </div> <div> <p>1072 out of 5363 tweets (19.9 %) contain the keyword "Latinos," 583 were posted in 2021 and 489 in 2022. 32 tweets (2.9 %) are biased against Latines. 108 tweets (10.1 %) call out bias against Latines.&nbsp;&nbsp;</p> </div> <div> <p>1122 out of 5363 tweets (20.9 %) contain the keyword "Muslims," 576 were posted in 2021 and 546 in 2022. 62 tweets (5.5 %) are biased against Muslims. 240 tweets (21.3 %) call out bias against Muslims.&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>File Description</h2> </div> <div> <p>The dataset is provided in a csv file format, with each row representing a single message, including replies, quotes, and retweets. The file contains the following columns:&nbsp;&nbsp;</p> <p>'TweetID': Represents the tweet ID.&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Username': Represents the username who published the tweet (if it is a retweet, it will be the user who retweetet the original tweet.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Text': Represents the full text of the tweet (not pre-processed).&nbsp;&nbsp;</p> </div> <div> <p>'CreateDate': Represents the date the tweet was created.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>'Biased': Represents the labeled by our annotators if the tweet is biased (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Calling_Out': Represents the label by our annotators if the tweet is calling out bias against minority groups (1) or not (0).&nbsp;&nbsp;</p> </div> <div> <p>'Keyword': Represents the keyword that was used in the query. The keyword can be in the text, including mentioned names, or the username.&nbsp;&nbsp;&nbsp;&nbsp;</p> </div> <div> <p>&nbsp;&lsquo;Cohort&rsquo;: Represents the year the data was annotated (class of 2022 or class of 2023)&nbsp;</p> </div> <div> <h2>&nbsp;</h2> <h2>Acknowledgements&nbsp; &nbsp;</h2> </div> <div> <p>We are grateful for the technical collaboration with Indiana University's Observatory on Social Media (OSoMe). We thank all class participants for the annotations and contributions, including Kate Baba, Eleni Ballis, Garrett Banuelos, Savannah Benjamin, Luke Bianco, Zoe Bogan, Elisha S. Breton, Aidan Calderaro, Anaye Caldron, Olivia Cozzi, Daj Crisler, Jenna Eidson, Ella Fanning, Victoria Ford, Jess Gruettner, Ronan Hancock, Isabel Hawes, Brennan Hensler, Kyra Horton, Maxwell Idczak, Sanjana Iyer, Jacob Joffe, Katie Johnson, Allison Jones, Kassidy Keltner, Sophia Knoll, Jillian Kolesky, Emily Lowrey, Rachael Morara, Benjamin Nadolne, Rachel Neglia, Seungmin Oh, Kirsten Pecsenye, Sophia Perkovich, Joey Philpott, Katelin Ray, Kaleb Samuels, Chloe Sherman, Rachel Weber, Molly Winkeljohn, Ally Wolfgang, Rowan Wolke, Michael Wong, Jane Woods, Kaleb Woodworth, Aurora Young, Sydney Allen, Hundre Askie, Norah Bardol, Olivia Baren, Samuel Barth, Emma Bender, Noam Biron, Kendyl Bond, Graham Brumley, Kennedi Bruns, Leah Burger, Hannah Busche, Morgan Butrum-Griffith, Zoe Catlin, Angeli Cauley, Nathalya Chavez Medrano, Mia Cooper, Suhani Desai, Isabella Flick, Samantha Garcez, Isabella Grady, Macy Hutchinson, Sarah Kirkman, Ella Leitner, Elle Marquardt, Madison Moss, Ethan Nixdorf, Reya Patel, Mickey Racenstein, Kennedy Rehklau, Grace Roggeman, Jack Rossell, Madeline Rubin, Fernando Sanchez, Hayden Sawyer, Diego Scheker, Lily Schwecke, Brooke Scott, Megan Scott, Samantha Secchi, Jolie Segal, Katherine Smith, Constantine Stefanidis, Cami Stetler, Madisyn West, Alivia Yusefzadeh, Tayssir Aminou, Karen Fecht, Luciana Orrego-Hoyos, Hannah Pickett, and Sophia Tracy.&nbsp;</p> </div> <div> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;</p> </div> <div> <p>&nbsp;</p> </div>

opencc-by-4.0Mar 2023View details →
dryad40/100

A high-resolution three-year dataset supporting rooftop photovoltaics (PV) generation analytics

<p>This dataset includes measured photovoltaic (PV) power generation data and on-site weather data collected from 60 grid-connected rooftop PV stations in Hong Kong over a three-year period (2021-2023). The PV power generation data was collected at 5-minute intervals. The meteorological data was collected at 1-minute intervals from an on-site weather station. The metadata was represented using Brick schema was developed, which simplifies the data comprehension and the development of smart analytics applications. The detailed Brick model is stored in the .ttl file format, which can be accessed for retrieving metadata through the use of SPARQL queries.This dataset can be used in various applications - PV generation benchmarking, PV degradation analysis, PV fault detection, solar radiation and PV power generation forecasting, and the simulation and design of PV systems.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Replication Package for "Software Quality Assurance Analytics: Enabling Software Engineers to Reflect on QA Practices" Paper (SCAM 2024)

<p>Welcome to our artifact!<br>In here we provide additional information for you to retrace our steps in the interview analysis.<br>It has the following contents:</p> <ul> <li><code>codebook.xlsx</code>: Our full codebook with our open codes, structured after the axial codes that emerged. <code>codebook-statistics.xlsx</code> lists for each code in which participant's interview it can be found.</li> <li><code>generate-figures</code>: The plain data and scripts used to generate the figures in the paper.</li> <li><code>survey.pdf</code>: An printout of our whole online questionnaire that guided the participants through the pretest-posttest study and the interview.</li> <li><code>survey-answers.xlsx</code>: The complete data for our participants answers in the online survey during the interviews.</li> <li><code>repoinsights-dashboard-software</code>: The code of our prototype repoinsights. As it is under active development, this is not yet documented for replicating the study setup or extending it. Still, we are providing the source code for transparency and will publish a version with comprehensive setup instructions later.</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Figure 3 in A review of molecular genetic markers and analytical approaches that have been used for delimiting marine mammal subspecies and species

Figure 3. Published values of percent divergence between cetacean subspecies (black bars), species (white bars), and taxa of uncertain taxonomic status (gray bars). Values are based on mtDNA control region sequence data. Not all values represent net sequence divergence. See Table 1 for list of papers corresponding to each value. Since completing this work, Sousa species have been supported ((Mendez et al. 2013) and Inia subspecies changed.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 1 in A review of molecular genetic markers and analytical approaches that have been used for delimiting marine mammal subspecies and species

Figure 1. Sample sizes used in publications of molecular genetic studies of marine mammals at different taxonomic levels. Graphs present the proportion of studies at each taxonomic level that fall into each sample size category. (A) minimum total sample size per focal taxon; (B) maximum sample size per single sampling locality. Papers were categorized as examining taxonomic questions at: species = subspecies/species boundary; subspecies = population/subspecies boundary; uncertain = taxonomic boundary uncertain (see text).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 2 in A review of molecular genetic markers and analytical approaches that have been used for delimiting marine mammal subspecies and species

Figure 2. Types of molecular genetic data used in published studies examining questions at the species-level, subspecies-level, or undefined taxonomic level for marine mammals. Note that studies may have used more than one data type. Mitochondrial DNA sequence data (MtDNASeq), nuclear DNA sequence data (NuSeq), microsatellites (Msats), morphological data (Morph).

opencc-by-4.0Jun 2017View details →
zenodo40/100

A semi-analytical solution for heat transport in rock with parallel fractures and a heat source in both fracture and matrix

<p>In this study, we propose a two-dimensional semi-analytical solution framework based on a Green&rsquo;s function approach for a flexible heat source definition, including source dimensions, energy delivery strength and duration, and the presence of a heat source in the matrix and/or fracture. The solution fully accounts for heat conduction, advection, dispersion and transient heat exchange between the mobile and immobile phases in a system of parallel fractures. The solution having a strip heat source extending from a fracture into the matrix indicates that one-dimensional heat conduction in the matrix underestimates and overestimates temperature responses at early and later times, respectively.</p> <p>The dataset is for the figures 2-7 in the journal paper.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

CUSP - UBC Workshop: Analytics: Characterization and quantification / enumeration of particles in the environment and in tissue

<p>The CUSP-UBC Workshop was held online on 28th January 2022.</p> <p>69 participants took part from across Europe and British Columbia to share experiences, exchange knowledge, and to discuss challenges and solutions as part of a great collaboration between the two clusters.</p> <p><strong>Acknowledgements:</strong></p> <p><strong>Co-Organisation and Cluster Presentations:</strong></p> <p>Lesley Tobin (CUSP Working Group 6 Communication and Dissemination, PlasticsFatE) <a href="mailto:lesley.tobin@optimat.co.uk">&nbsp;</a><a href="mailto:lesley.tobin@optimat.co.uk">lesley.tobin@optimat.co.uk</a></p> <p>Mahdi Takaffoli (Coordinator, Cluster for Microplastics, Health and the Environment, The University of British Columbia) <a href="mailto:mahdi.takaffoli@ubc.ca">mahdi.takaffoli@ubc.ca</a></p> <p><strong>Presenters:</strong></p> <p><strong>Florian Meirer </strong>(Associate Professor, Inorganic Chemistry and Catalysis research group&nbsp;at Utrecht University; Polyrisk &amp; Aurora)</p> <p>&ldquo;Characterizing Nanoplastics with Force Microscopy &ndash; An Update&rdquo;) <a href="mailto:F.Meirer@uu.nl">F.Meirer@uu.nl</a></p> <p><strong>Anna Costa</strong> (Environmental Nanotechnology and Nano-Safety group of CNR-ISTEC; PlasticsFatE) &ldquo;Strategies for MP/NP simulated samples-laboratory tests&rdquo; <a href="mailto:anna.costa@istec.cnr.it">anna.costa@istec.cnr.it</a></p> <p><strong>Tao Huan</strong> (Assistant Professor, Chemistry, The University of British Columbia)</p> <p>&ldquo;Pilot Study of the Impact of Microplastics on Cell Liability and Potential Application of Metabolomics in Understanding the Biological Mechanisms&rdquo;&nbsp;<a href="mailto:thuan@chem.ubc.ca">thuan@chem.ubc.ca</a></p> <p><strong>&nbsp;Edward Grant</strong> (Professor, UBC Chemistry) &ldquo;The challenge of representative microplastic analysis&rdquo; edgrant@chem.ubc.ca</p> <p>Thank you to Michelle Epstein, Doctor of Allergy and Clinical Immunology, MedUni Vienna, for such a useful, stimulating idea, and to everyone who took part despite the unsocial hours!</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
dryad40/100

Analytical kinetic model of native tandem promoters in E. coli

<p><span>Closely spaced promoters in tandem formation are abundant in bacteria. We investigated the evolutionary conservation, biological functions, and the RNA and single-cell protein expression of genes regulated by tandem promoters in <i>E. coli</i>. We also studied the sequence (distance between transcription start sites '<i>d<sub>TSS</sub>'</i>,<i> </i>pause sequences, and distances from oriC) and potential influence of the input transcription factors of these promoters. From this, we propose an analytical model of gene expression based on measured expression dynamics, where RNAP-promoter occupancy times and <i>d<sub>TSS</sub> </i>are the key regulators of transcription interference due to TSS occlusion by RNAP at one of the promoters (when <i>d<sub>TSS</sub> </i>≤ 35 bp) and RNAP occupancy of the downstream promoter (when <i>d<sub>TSS</sub> </i>&gt; 35 bp). Occlusion and downstream promoter occupancy are modeled as linear functions of occupancy time, while the influence of <i>d<sub>TSS</sub> i</i>s implemented by a continuous step function, fit to <i>in vivo</i> data on mean single-cell protein numbers of 30 natural genes controlled by tandem promoters. The best-fitting step is at 35 bp, matching the length of DNA occupied by RNAP in the open complex formation. This model accurately predicts the squared coefficient of variation and skewness of the natural single-cell protein numbers as a function of <i>d<sub>TSS</sub></i>. Additional predictions suggest that promoters in tandem formation can cover a wide range of transcription dynamics within realistic intervals of parameter values. By accurately capturing the dynamics of these promoters, this model can be helpful to predict the dynamics of new promoters and contribute to the expansion of the repertoire of expression dynamics available to synthetic genetic constructs.</span></p>

opencc-zeroFeb 2022View details →
zenodo40/100

Application of multi-analyte / multi-matrix screening method for pesticide residues in fruits and vegetables to Interaboratory Comparison Study on Pesticide Residues in Food (ILC)

<p>The suitability of multi-analyte / multi-matrix screening method for pesticide residues in fruits and vegetables and related products, developed within activities of task (Multi-analyte / multi-matrix screening method for pesticide residues in fruits and vegetables (including tea) and fruit juices), was evaluated by the Interaboratory Comparison Study on Pesticide Residues in Food (ILC). &nbsp;Test material (&ldquo;Pesticide Residues in green tea&rdquo;), prepared from the batch used for another proficiency test (PT), was provided by Fapas (Fera Science Ltd, York, UK).</p> <p>Data set obtains (i) information about performance characteristics of the analytical method and (ii) compilation of results of interlaboratory study.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Covid Data Analytics: Repositório de Dados Provenientes de Múltiplas Fontes sobre a Pandemia de COVID-19 no Brasil

<p>Uma estrat&eacute;gia para melhor compreender as diversas facetas e poss&iacute;veis impactos da pandemia de COVID-19 na sociedade consiste na extra&ccedil;&atilde;o de informa&ccedil;&atilde;o e conhecimento a partir de dados provenientes de diversas fontes oficiais e n&atilde;o oficiais.</p> <p>A import&acirc;ncia desse tema fomentou a publica&ccedil;&atilde;o de diversos artigos cient&iacute;ficos que investigam aspectos relacionados &agrave; pandemia de COVID-19 no Brasil por meio de an&aacute;lises de dados. Alguns trabalhos, por exemplo, fornecem caracteriza&ccedil;&otilde;es e descri&ccedil;&otilde;es da evolu&ccedil;&atilde;o da doen&ccedil;a no pa&iacute;s~\cite{ranzani2021characterisation}, considerando, inclusive, a subnotifica&ccedil;&atilde;o de casos pelas ag&ecirc;ncias oficiais. Outros modelam e preveem a evolu&ccedil;&atilde;o da COVID-19, utilizando dados referentes aos primeiros meses da pandemia e empregando diferentes m&eacute;todos&nbsp;ou mesmo utilizando dados de geolocaliza&ccedil;&atilde;o e de din&acirc;mica populacional.</p> <p>Nesse contexto, &eacute; importante que, sempre que poss&iacute;vel, os dados utilizados para as pesquisas sejam disponibilizados &agrave; comunidade cient&iacute;fica, seja para fins de &nbsp;replicabilidade dos resultados encontrados, seja para a promo&ccedil;&atilde;o de novas investiga&ccedil;&otilde;es.</p> <p>Os dados disponibilizados no reposit&oacute;rio CDA &nbsp;se referem ao per&iacute;odo entre 23 de fevereiro de 2020 e 8 de maio de 2021.<br> Esse reposit&oacute;rio agrega 1.508 arquivos, classificados em dois tipos principais: (i) bases de dados e tabelas extra&iacute;das das fontes descritas anteriormente; e (ii) artigos, relat&oacute;rios, mapas e gr&aacute;ficos produzidos pelos integrantes do projeto a partir da an&aacute;lise dos dados coletados</p> <p><strong>Dados de Fontes Externas</strong></p> <p>Estes arquivos representam 8\% do total de arquivos que comp&otilde;em o reposit&oacute;rio e est&atilde;o distribu&iacute;dos da seguinte maneira:</p> <ul> <li>S&eacute;ries temporais com indicadores econ&ocirc;micos das Unidades Federativas do Brasil e da Uni&atilde;o em formato .csv, com aproximadamente 18.400 registros;</li> <li>7 scripts de tratamento de dados em formato .py.</li> <li>5 arquivos com a contagem do n&uacute;mero de tweets e retweets coletados semanalmente utilizando 13 palavras-chave (&ldquo;corona&rdquo;, &ldquo;covid&rdquo;, &ldquo;coronavirus&rdquo;, &ldquo;covid19&rdquo;, &ldquo;quarentena&rdquo;,&ldquo;hidroxicloroquina&rdquo;, &ldquo;cloroquina&rdquo;, &ldquo;confinamento&rdquo;, &ldquo;distanciamento social&rdquo;, &ldquo;aglomera&ccedil;&atilde;o&rdquo;, &ldquo;aglomera&ccedil;&otilde;es&rdquo;, &ldquo;sars&rdquo; e &ldquo;covid-19&rdquo;) formato .csv&nbsp;</li> <li>3 arquivos do Google Trends no formato .csv com 249 registros contendo 124 termos pr&eacute;-selecionados que t&ecirc;m rela&ccedil;&atilde;o com a pandemia e o percentual relativo de buscas na web nos n&iacute;veis regional e nacional. %\ana{de novo, o que tem nestes csvs?}.</li> <li>7 arquivos como dados anonimizados do Instagram no formato .csv com 90.787 hashtags contendo os termos \#demito, \#demitida, \#desempregada, \#desempregado, \#desemprego, \#falido, \#redu&ccedil;&atilde;odejornada.</li> <li><strong>An&aacute;lises e Relat&oacute;rios</strong></li> <li>Os arquivos com as an&aacute;lises e relat&oacute;rios representam 92\% do total de arquivos do reposit&oacute;rio. Al&eacute;m de documentos de texto, tamb&eacute;m foram disponibilizados materiais visuais, como mapas e gr&aacute;ficos, em diversos formatos. Os arquivos est&atilde;o distribu&iacute;dos da seguinte maneira.</li> </ul> <p>&nbsp; &nbsp;&nbsp;</p> <ul> <li>23 gr&aacute;ficos comparativos de indicadores sociais e econ&ocirc;micos, an&aacute;lises descritivas dos ocupados em atividades essenciais e n&atilde;o ess&ecirc;ncias por regi&otilde;es em formato .svg;</li> <li>409 gr&aacute;ficos de novos casos e &oacute;bitos (02 a 09 de setembro) em formato .png;</li> <li>&nbsp;522 mapas e gr&aacute;ficos de linhas e barras acerca dos casos e &oacute;bitos acumulados de COVID-19 em todo o pa&iacute;s entre as semanas epidemiol&oacute;gicas 9 e 32 de 2020 (23/02/2020 a 08/08/2020) em formato .png;</li> <li>15 arquivos de medidas provis&oacute;rias em formato .pdf;</li> <li>1 gr&aacute;fico interativo gerado a partir do c&aacute;lculo da mortalidade (&oacute;bitos acumulados por 100 mil habitantes) no Brasil em formato .html;</li> <li>25 anima&ccedil;&otilde;es mostrando a evolu&ccedil;&atilde;o da letalidade (mortes acumuladas / casos acumulados) em \% em todos estados do Brasil a cada semana epidemiol&oacute;gica da 9&ordf; &agrave; 31&ordf; em formato .gif;</li> <li>1 relat&oacute;rio sobre an&aacute;lises das informa&ccedil;&otilde;es dispon&iacute;veis para coleta na ferramenta Google Trends em formato .pdf;</li> <li>4 relat&oacute;rios sobre an&aacute;lises das informa&ccedil;&otilde;es dispon&iacute;veis dos grupos de pesquisa em formato .pdf;&nbsp;&nbsp;</li> </ul> <p><strong>Limita&ccedil;&otilde;es nas Bases de Dados Disponibilizadas</strong><br> Devido a quest&otilde;es de privacidade, algumas bases de dados, obtidas atrav&eacute;s da extra&ccedil;&atilde;o de informa&ccedil;&otilde;es das redes sociais online n&atilde;o foram integralmente disponibilizadas no reposit&oacute;rio.<br> Nestes casos, disponibilizamos an&aacute;lises extra&iacute;das a partir destas bases, realizadas com o prop&oacute;sito de responder algumas das perguntas de pesquisa do projeto. As an&aacute;lises realizadas durante o projeto est&atilde;o dispon&iacute;veis em &nbsp;\url{https://covid.dcc.ufmg.br/}.</p> <p>A disponibiliza&ccedil;&atilde;o dos dados ocorreu por meio do&nbsp; padr&atilde;o <em>Open Data Standards</em> e, a partir dele, foram criados e organizados os arquivos, de acordo com o respectivo formato e tipo de informa&ccedil;&atilde;o. Eles foram integrados ao drive do grupo tecnol&oacute;gico por interm&eacute;dio de um formul&aacute;rio, e utilizaram-se de um script para transformar os dados do formul&aacute;rio em um arquivo XML. Como resultado, p&ocirc;de-se modelar e preencher o banco de dados a partir do arquivo XML e, por fim, integr&aacute;-lo a um <a href="https://covid.dcc.ufmg.br/buscador.php">buscador</a> criado em Wordpress, que fica dispon&iacute;vel para download no portal do projeto CDA maiores informa&ccedil;&otilde;es: <a href="https://covid.dcc.ufmg.br/linhas/dados">https://covid.dcc.ufmg.br/linhas/dados/&nbsp;</a></p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Analytical modeling of an hybrid power module based on diamond and SiC devices

<p>This dataset contains the raw data used for the publication (available here : <a href="https://doi.org/10.1016/j.diamond.2022.108936">10.1016/j.diamond.2022.108936</a> ).</p> <p><strong>Analytical modeling of an hybrid power module based on diamond and SiC devices</strong></p> <p>Marine Couret, Anne Castelan, Nazareno Donato, Florin Udrea, Julien Pernot, Nicolas Rouger</p> <p>Detailed descriptions for each file&nbsp;can be found in &quot;Dataset_Description.docx&quot;.</p>

opencc-by-sa-4.0Aug 2021View details →
zenodo40/100

The Piraeus AIS Dataset for Large-scale Maritime Data Analytics

<p><strong>AIS data collected by the&nbsp;University of Piraeus&#39; AIS receiver</strong></p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>The advent of Big Data and streaming technologies has resulted in a swarm of voluminous, heterogeneous information, especially in the domains of Internet of Things (IoT) and transportation. Focusing on the maritime field, we present a dataset that contains vessel position information transmitted by vessels of different types and collected via the Automatic Identification System (AIS). The AIS dataset comes along with spatially and temporally correlated data about the vessels and the area of interest, including weather information. It covers a time span of over 2.5 years, from May 9<sup>th</sup>, 2017 to December 26<sup>th</sup>, 2019 and provides anonymised vessel positions within the wider area of the port of Piraeus (Greece), one of the busiest ports in Europe and worldwide. The dataset consists of over 244 million AIS records, an average of more than 10,000 records per hour, which makes it an ideal input for large-scale mobility data processing and analytics purposes.</p> <p>&nbsp;</p> <p><strong>Dataset related to the following publication</strong></p> <blockquote> <p>Andreas Tritsarolis, Yannis Kontoulis, Yannis Theodoridis, The Piraeus AIS dataset for large-scale maritime data analytics, Data in Brief, Volume 40, 2022, 107782, ISSN 2352-3409,&nbsp;<a href="https://doi.org/10.1016/j.dib.2021.107782">https://doi.org/10.1016/j.dib.2021.107782</a>.</p> </blockquote> <p>&nbsp;</p> <p><strong>Files Description</strong></p> <ul> </ul> <ul> <li><strong>ais_static</strong>: CSV flat files&nbsp;containing vessels&#39; static information and their corresponding types</li> </ul> <ul> <li><strong>geodata</strong>: ESRI Shapefiles containing several geographic-related data (e.g. harbours, islands, etc.)</li> </ul> <ul> <li><strong>noaa_weather</strong>: ESRI Shapefiles containing weather forecast from GRIB files (as provided by NOAA)</li> </ul> <ul> <li><strong>unipi_ais_dynamic</strong>: CSV flat files containing AIS kinematic information&nbsp;</li> </ul> <ul> <li><strong>unipi_ais_dynamic_synopses</strong>: CSV flat files containing metadata (i.e. synopses) regarding vessels&#39; AIS positions</li> </ul> <p>&nbsp;</p> <p><strong>Privacy Statement</strong></p> <p><strong>For privacy-related queries, please contact the authors</strong></p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Datasample - Social Media Analytics and Metrics of Facebook Performance of Libraries, Archives and Museums

<p>The current dataset describes Facebook pages performance for 220 Libraries, Archives and Museums from all over the world. The performance is measured through 9 different social media metrics. That is, number of posts, link-posts, picture-posts, video-posts, total reactions, comments and shares, number of reactions, comments per post and reactions per post. The data harvesting process has been conducted through the use of FanPageKarma API. The gathered metrics and their values depict the performance for each Facebook page in a time-period of 30 days.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Analytical Center of University Cultural Productions in the Context of the Conflict (caPAZ)

<p>This dataset comprises a collection of journalistic articles written by young university students in Colombia, which is a product of the project: Analytical Center of University Cultural Productions in the Context of the Conflict (caPAZ), funded by the Ministry of Science, Technology and Innovation (Minciencias) and the National Center for Historical Memory (CNM) of Colombia (under the code: 1349-872-76354, agreement 872 of 2020) This corpus includes news written by the 8 colleges media of the Colombian Network of College Journalism from 2001 to 2021. The dataset includes digital&nbsp;news, for a total of 2373 news items related to the armed conflict, the memory of the victims and the peace process in Colombia.&nbsp;These news items were collected through a web-scraping technique, using 3 lemmatized keywords (conflicto armado, memoria de las v&iacute;ctimas y proceso de paz), with the aim of identifying these regular expressions in the logical operators that run through the HTML structure of each Web page</p>

opencc-by-4.0Mar 2022View details →
dryad40/100

Data, analytical, and plotting codes needed to reproduce meta-analyses on within-season divorce in birds

<p>This data package contains data (effect sizes and other variables) and analytical codes needed to reproduce three meta-analyses on within-season divorce and breeding success in socially monogamous birds (Culina &amp; Brouwer: No evidence of immediate fitness benefits of within-season divorce in monogamous birds. 2022. Biology Letters.) It also contains codes to reproduce Figures from the main text and from the Supplement. Readme file and Licence are provided too.</p>

opencc-zeroMay 2022View details →
zenodo40/100

Dataset supporting manuscript entitled 'Partitioning of Small Hydrophobic Molecules into Polydimethylsiloxane in Microfluidic Analytical Devices'

<p>This is the dataset supporting the manuscript &#39;Surface and bulk modifications of polydimethylsiloxane to reduce absorption/adsorption of small molecules in lab-on-a-chip&#39; published in Micromachines</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Equivariant analytical mapping of first principles Hamiltonians to accurate and transferable materials models

<p>Supporting data for&nbsp;<a href="https://arxiv.org/abs/2111.13736">https://arxiv.org/abs/2111.13736</a>.</p> <p>ACEhamiltonians.jl code</p> <p>This is an archived copy of the ACEhamiltonians.jl code to accompany the paper&nbsp;<a href="https://arxiv.org/abs/2111.13736">arXiv:2111.13736</a>.</p> <p>See&nbsp;<a href="https://github.com/ACEsuit/ACEhamiltoniansExamples">https://github.com/ACEsuit/ACEhamiltoniansExamples</a>&nbsp;for examples of how to use this code.</p> <p>The code is written in&nbsp;<a href="https://julialang.org/">Julia</a>&nbsp;and requires v1.6 or later. To install the Julia depenendencies:</p> <pre><code><code>$ cd ACEhamiltonians.jl $ julia julia&gt; import Pkg julia&gt; Pkg.activate(&quot;.&quot;) julia&gt; Pkg.instantiate() </code></code></pre> <p>The scripts&nbsp;<code>test/plots.jl</code>,&nbsp;<code>test/fcc-to-bcc.jl</code>&nbsp;and&nbsp;<code>test/vacancy.jl</code>&nbsp;which produce all the plots in the paper can then run as, e.g.</p> <pre><code><code>julia --project=. test/plots.jl </code></code></pre> <p>Training data</p> <p>The&nbsp;<code>training_data</code>&nbsp;folder contains the atomic structure, Hamiltonian and overlap matrices stored in HDF5 format with the following schema:</p> <ul> <li>Data Group :&nbsp;<strong>aitb/</strong></li> <li>Datasets : <ul> <li><strong>H</strong>&nbsp;: Real-space Hamiltonian Matrix. Type: Float64. Shape: Tensor(# of TB Cells, # of Rows, # of Columns)</li> <li><strong>S</strong>&nbsp;: Real-space Overlap Matrix. Type: Float64. Shape: Tensor(# of TB Cells, # of Rows, # of Columns)</li> <li><strong>energy</strong>&nbsp;: Energy. Unit: eV. Type: Float64. Shape: Scalar</li> <li><strong>freeenergy</strong>&nbsp;: Free Energy. Unit: eV. Shape: Scalar</li> <li><strong>unitcell</strong>&nbsp;: Unit cell vectors. Type: Float64. Shape: Matrix(3,3)</li> <li><strong>positions</strong>&nbsp;: Atom positions. Type: Float64. Shape: Array(3)</li> <li><strong>forces</strong>&nbsp;: (Optional, if available) Forces. Type: Float64. Shape: Array(3)</li> <li><strong>metadata</strong>&nbsp;: JSON String including dictionary of information of FHIaims calculation (k-points, basis sets), TB Cells, Cutoff, Orbital definitions.,</li> </ul> </li> </ul> <p>The molecular dynamics and FHI-aims parameters are described in the manuscript.</p> <p>On-site models</p> <p>The&nbsp;<code>onsite_models_ord2</code>&nbsp;folder contains our correlation order 2 models for the on site blocks of the Hamiltonian, in a JSON format readable by the&nbsp;<a href="https://github.com/acesuit/ACE.jl">ACE.jl</a>&nbsp;and&nbsp;<a href="https://github.com/ACEsuit/ACEhamiltonians.jl">ACEhamiltonians.jl</a>&nbsp;Julia packages. There are separate files for the Hamiltonian (<code>*_H.json</code>) and overlap (<code>*_S.json</code>) models. The JSON files also contain training and test sets and associated errors as plotted in Figure 3 in our manuscript.</p> <p>Models have a unique identifier (UUID) which is a hash of the input parameters and training data. The mapping from (order, max_degree) to UUID is as follows:</p> <pre><code><code>(2,4) - 13427527590286463256 (2,5) - 10538156191357510769 (2,6) - 1646489440533135164 (2,7) - 12130775482127724115 (2,8) - 12487060958610974041 (2,9) - 2653067664384673997 (2,10) - 1143382251563115664 (2,11) - 4564001820340015372 (2,12) - 9474261500251782658 </code></code></pre> <p>Off-site models</p> <p>The&nbsp;<code>offsite_models_ord1</code>&nbsp;and&nbsp;<code>offsite_models_ord2</code>&nbsp;folders contain our order 1 and order 2 offsite models for Hamiltonian and overlap matrices. The mapping from (H_order, H_max_degree) + (S_order, S_max_degree) to UUID is as follows:</p> <pre><code><code>(1,6) + (1,8) - 7014526518680934587 (1,7) + (1,9) - 8594416159488562244 (1,8) + (1,10) - 10204186688118368371 (1,9) + (1,11) - 13078304848585360574 (1,10)+ (1,12) - 14750835312950641338 (1,11)+ (1,13) - 9883802224093245794 (1,12)+ (1,14) - 3907899412408606585 (1,13)+ (1,15) - 201683837542179657 (1,14)+ (1,16) - 277744202775070779 (2,6) + (1,8) - 4699475053563592071 (2,7) + (1,9) - 489637409713831432 (2,8) + (1,10) - 18034631670613263469 (2,9) + (1,11) - 720654516759450160 (2,10)+ (1,12) - 15214900801060024044 (2,11)+ (1,13) - 13798832597295943078 (2,12)+ (1,14) - 13162803789413134473 </code></code></pre> <p>FCC only</p> <p>Onsite models:</p> <pre><code><code>2 6 5311732756869418284 2 7 13030014632886405308 2 8 5820099621734447846 2 9 10161014511878227635 2 10 11298425190201843107 2 11 9932031839231628354 2 12 9447261873515969583 </code></code></pre> <p>Optimised FCC model&nbsp;<code>16110190062237887798</code></p> <p>BCC only</p> <p>Onsite models:</p> <pre><code><code>2 6 8949023800586845770 2 7 8045797268444730200 2 8 6919809282139600809 2 9 9935027806122780319 2 10 6376963380608532713 2 11 5001375576268070883 2 12 9678585765722197901 </code></code></pre> <p>Optimised BCC model&nbsp;<code>10293566074413000591</code></p> <p>FCC+BCC optimised models</p> <p>Onsite models</p> <pre><code><code>2 6 2154760103892646619 2 7 6450474921309693835 2 8 14227277988574899288 2 9 476820595195218567 2 10 5364136683220082110 2 11 14619519825012606580 2 12 14181614899005838824 </code></code></pre> <p>Offsite FCC+BCC optimised model -&nbsp;<code>4570230078043807257</code></p> <p>Model errors</p> <p>The&nbsp;<code>model_errors</code>&nbsp;directory contains summarised model errors for the training and testing errors for the models listed above.</p> <p>Reference data</p> <p>Reference electronic structure data computed for the BCC and FCC crystals, along the Bain path and for the relaxed vacancy is stored in the&nbsp;<code>reference_data</code>&nbsp;folder. The Hamiltonian and overlap matrices are stored as compressed binary HDF5 files. The format and metadata can be viewed with the&nbsp;<code>h5dump</code>&nbsp;utility, or read in using the supplied Julia code (or indeed from other languages).</p> <p>Predicted data</p> <p>The&nbsp;<code>predicted_data/FCC</code>&nbsp;and&nbsp;<code>predicted_data/BCC</code>&nbsp;folders contain HDF5 files with the results of all model predictions shown in the manuscript on the FCC and BCC crystal structures.&nbsp;<code>predicted_data/FCC-to-BCC</code>&nbsp;contains the results of predictions along the Bain path with the optimized model described in the manuscript and&nbsp;<code>predicted_data/vacancy</code>&nbsp;contains the vacancy calculations.</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record