Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,036
datasets available to search
ShareScore release 0.7.1
Dataset results
21,036 results for “Thematic”
Data for The Generally Curious: Thematically Distinct Datasets of 4chan's /pol/ Discussion Forum's 'General Threads'
<p>Over the second half of the 2010s, the /pol/ (‘politically incorrect’) forum on the 4chan image board has emerged as a space within which various extreme political ideologies are discussed and cultivated, occasionally informing off-site acts of political extremism. While previous research has often studied this space as a unified whole, it is relevant to more specifically demarcate different publics within 4chan’s /pol/ board, apart from studying it as an ‘amorphous blob’. This paper focuses specifically on ‘generals’ - recurring threads with a specific thematic focus identified by a particular vernacular phrase or tag. By identifying them it is possible to subset the board’s archive into multiple distinct datasets comprising discussions about a particular topic, such as Donald Trump, the Syria war, or British politics. We provide a dataset containing 58,841 opening posts and 13,697,738 replies to those, divided over 329 thematically distinct ‘general thread’ collections. In this paper we outline our data collection and query protocol, the structure of the data and its rationale, as well as a number of suggested research uses for this new data.</p> <p> </p>
Publishing Reproducible Research Outputs - Thematic coding of interview findings
<p>The spreadsheet in the present dataset (CSV format) includes the anonymised thematic coding that has been applied to our interview findings. A list of interviewees and interview questions is available <a href="https://doi.org/10.5281/zenodo.5141665">here</a>.</p> <p>The thematic coding has been applied by using <a href="https://www.qsrinternational.com/nvivo-qualitative-data-analysis-software/home">NVivo</a>, a professional qualitative analysis software, and then exported in spreadsheet form for public sharing. The findings of this analysis have been used to inform our final report, which is available in our <a href="https://zenodo.org/communities/ke-prro/?page=1&size=20">Zenodo project Community</a>.</p>
Sharing research data and findings relevant to the novel coronavirus (COVID-19) outbreak - Thematic coding of qualitative research findings
<p>The spreadsheet in the present dataset (CSV format) includes the anonymised thematic coding that has been applied to our interview and literature review findings to inform the preparation of the report: From intent to impact: Investigating the effects of open sharing commitments.</p> <p>The thematic coding has been applied by using <a href="https://www.qsrinternational.com/nvivo-qualitative-data-analysis-software/home">NVivo</a>, a professional qualitative analysis software, and then exported in spreadsheet form for public sharing.</p> <p>Find out more about this project in our dedicated <a href="https://zenodo.org/communities/data-sharing-in-public-health-emergencies">Zenodo project community</a>.</p>
Acoustic Guitar Timbre Thematic Analysis
<p>A perceptual study was conducted to investigate listener perceptions of acoustic guitar timbre, encompassing descriptive and preference analysis, and the impact of guitar playing style on perceived timbre similarity and preference.</p> <p>The listening test was based on recordings of ten different steel-string acoustic guitars at various price points, sourced from the online music retailer Thomann (https://www.thomann.de/). For each guitar, recordings of three different songs were used, each with a different playing style: picking (mainly individual notes played with a combination of fingers and pick), strumming (mainly chords played with a pick), and fingerstyle (strings plucked with fingers rather than a pick).</p> <p>The study was completed by 27 participants (8 female, 19 male, mean age: 27) of 14 different nationalities. Participants had an advanced musical proficiency (Goldsmiths Musical Sophistication Index General Sophistication score of 97.85) and 11 of them played guitar as their primary instrument. Participants listened to each guitar in each of the three playing styles and were asked to describe the instrument's timbre, what they liked and what they disliked.</p> <p>We conducted a thematic analysis of the participant answers to the three questions (timbre description, timbre like, timbre dislike) for the ten guitars using a combination of deductive and inductive approaches. This dataset provides the analysis codebook, theme and code statistics.</p> <p>Details about the thematic analysis and the study can be found in the accompanying publication currently under revision (the reference will be added in due course).</p>
A Thematic Synthesis on Empathy in Software Engineering based on the Practitioners' Perspective - Supplementary Material
<p>This repository contains the supplementary material of the paper "A Thematic Synthesis on Empathy in Software Engineering based on the Practitioners' Perspective" <br> (DOI https://doi.org/10.1145/3613372.3613407) accepted at the Research Track of the <br> XXXVII Brazilian Symposium on Software Engineering (SBES 2023).</p> <p>The artifacts are a result of a thematic synthesis of grey literature <br> performed to investigate the meaning, importance, practices, and effects of empathy <br> from the perspective of software practitioners. <br> The analysis was based on web articles from DEV, an online community used by software developers. <br> The data were collected and stored in the repository to preserve the evidence and ensure the study’s replicability.<br> <br> The repository contains the following material:</p> <p>1- <all codes.ods> and <all codes.xlsx><br> All codes generated in the data extraction process, considering research questions RQ1-RQ5:<br> The two files have the same content in different formats - ODS and XLSX.</p> <p>2 - <dataset.csv> <br> The list of web articles collected from the DEV in CSV format with all inclusion and <br> exclusion information, plus demographic data.</p> <p>3 - <empathy-framework.jpg><br> Figure 3 of the paper: A conceptual map of the meaning (boxes in orange) and <br> the value (boxes in blue) of empathy according to the software practitioners</p> <p>4 - <empathy-model.jpg> <br> Figure 4 of the paper: A conceptual framework for communication and collaboration (A), <br> management and leadership (B), coding (C), and code review (D).</p> <p>5 - <extraction.ods> and <extraction.xlsx>. The data extracted from the web articles, <br> including codes and quotes for each research question. <br> The two files have the same content in different formats - ODS and XLSX.</p>
EUNIS Habitat Maps: Enhancing Thematic and Spatial Resolution for Europe through Machine Learning
<p>The EUNIS habitat classification is essential for categorising European habitats and supporting European policy on nature conservation and to implement the Nature Restoration Law. As such, to meet the growing demand for detailed and accurate habitat information, we provide spatial predictions for 260+ EUNIS habitat types at EUNIS level 3, together with validation and uncertainty analyses. </p> <p>More specifically, using ensemble machine learning models together with high-resolution satellite imagery and other climatic, terrain and soil variables, we produced an European habitat map at a 100-m resolution indicating the most likely EUNIS habitat at level 3 for every location across Europe. Predictions were validated for three independent countries, namely for France, the Netherlands and Austria. We also provide information on uncertainty and the most probable habitats at level 3 within each EUNIS level 1 formation. Products can be further refined with accurate and local land cover data. This product is thus likely to be particularly useful for restoration but also conservation purposes. </p> <p>Figure: <strong>Wall-to-wall map of EUNIS habitats at level 3 - (color coded at level 2 for visibility)</strong></p> <p></p>
Dataset: Global X Thematic Growth ETF (GXTG) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Fig. 3 in The correlations between certain features of the journal Neotropical Ichthyology and its impact factor: a comparative analysis at the thematic and national levels
Fig. 3. Correlation between average IF and uncitedness rate of journals on zoology in Sample 1 between 2006 and 2010. The highlighted represents the data for Neotropical Ichthyology.
Domain-Driven Design in Microservices-Based Systems Development: A Systematic Literature Review and Thematic Analysis [Dataset]
<p>This repository contains all artifacts related to the study: Domain-Driven Design in Microservices-Based Systems Development: A Systematic Literature Review and Thematic Analysis</p>
Thematic Layers of Causative Factors of Landslide in Palungtar Municipality, Gorkha, Nepal
<p><span>Eleven causative parameters responsible for instability in the area: Slope, Aspect, Land Use Land Cover, Plan Curvature, Profile Curvature, Distance from River, Distance from Road, Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), Topographic Wetness Index (TWI), and Geology. </span></p>
Wikidata Thematic Subgraph Selection
<p><strong>Wikidata Thematic Subgraph Selection</strong></p> <p>These datasets have been designed to train and evaluate algorithms to select thematic subgraphs of interest in a large knowledge graph from seed entities of interest. Specifically, we consider Wikidata. Given a set of seed QIDs of interest, a graph expansion is performed following P31, P279, and (-)P279 edges. Traversed classes that thematically deviates from seed QIDs of interest should be pruned. Datasets thus consist of classes reached from seed QIDs that are labeled as "to prune" or "to keep".</p> <p><strong>Available datasets</strong></p> <table> <tbody><tr> <th>Dataset</th> <th># Seed QIDs</th> <th># Labeled decisions</th> <th># Prune decisions</th> <th>Min prune depth</th> <th>Max prune depth</th> <th># Keep decisions</th> <th>Min keep depth</th> <th>Max keep depth</th> <th># Reached nodes up</th> <th># Reached nodes down</th> </tr> </tbody><tbody> <tr> <td><a href="data/dataset1">dataset1</a></td> <td>455</td> <td>5233</td> <td>3464</td> <td>1</td> <td>4</td> <td>1769</td> <td>1</td> <td>4</td> <td>1507</td> <td>2593609</td> </tr> <tr> <td><a href="data/dataset2">dataset2</a></td> <td>105</td> <td>982</td> <td>388</td> <td>1</td> <td>2</td> <td>594</td> <td>1</td> <td>3</td> <td>1159</td> <td>1247385</td> </tr> </tbody> </table> <p>Each dataset folder contains</p> <ul> <li><code>datasetX.csv</code>: a CSV file containing one seed QID per line (not the complete URL, just the QID). This CSV file has no header.</li> <li><code>datasetX_labels.csv</code>: a CSV file containing one seed QID per line and its label (not the complete URL, just the QID)</li> <li><code>datasetX_gold_decisions.csv</code>: a CSV file with seed QIDs, reached QIDs, and the labeled decision (1: keep, 0: prune)</li> <li><code>datasetX_Y_folds.pkl</code>: folds to train and test models based on the labeled decisions</li> </ul> <p><code>dataset1-2</code> consists of using <code>dataset1</code> for training and <code>dataset2</code> for testing.</p> <p><strong>License</strong></p> <p>Datasets are available under the <a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC</a> license.</p>
The need to develop tailored tools for improving the quality of thematic bibliometric analyses: Evidence from papers published in Sustainability and Scientometrics (Dataset)
<p>This dataset contains the data used to completed the article under peer review:</p> <p>References:<br> Cabezas, A.; Milanés, Y.;Alba, R.; Delgado, A.M. (2023). The need to develop tailored tools for improving the quality of thematic bibliometric analyses: Evidence from papers published in Sustainability and Scientometrics. (Article under peer review)</p> <p>Institutions: Spain (Universidad Internacional de La Rioja, Universidad Pablo de Olavide, Hospital Universitario Virgen de las Nieves)</p>
A thematic synthesis on the adoption of regression testing techniques in Android projects
<p>In software testing, employing regression techniques is a viable strategy to deal with the complexity and the constant evolution of applications since its primary goal is to ensure that changes made between versions do not change the system’s behavior. Although the literature has dedicated efforts to developing new regression testing techniques suitable for the Android mobile platform, studies are limited concerning demonstrating which techniques software developers employ in practice. This study aims to report on a thematic synthesis of adopting regression testing techniques in Android projects. The research encompassed four stages: (i) conducting a structured literature review on regression testing techniques for the Android platform, (ii) carrying out an expert survey, (iii) conducting interviews with industry professionals, and (iv) building a thematic synthesis. The thematic synthesis presented a model from analyzing the results obtained in this multimethod study on regression testing techniques. With such a stud, we could present empirical evidence on how professionals perform regression testing in Android projects, identify the commonly used regression testing techniques, and leverage the requirements for automating Android applications through regression testing.</p>
Fig. 4 in The correlations between certain features of the journal Neotropical Ichthyology and its impact factor: a comparative analysis at the thematic and national levels
Fig. 4. Uncitedness rate of articles published in the Brazilian journals in Sample 2.
Fig. 1 in The correlations between certain features of the journal Neotropical Ichthyology and its impact factor: a comparative analysis at the thematic and national levels
Fig. 1. Brazilian journals and their corresponding self-cited rates between 2006 and 2011.
Fig. 2 in The correlations between certain features of the journal Neotropical Ichthyology and its impact factor: a comparative analysis at the thematic and national levels
Fig. 2. Brazilian journals and their corresponding self-citing rates between 2006 and 2011.
Multi-parameter thematic maps for VRT
<p>"Optimization of VRT technology to the resilient tomato lines", during year 2020 an extensive set of field surveys has been conducted by Casella using the MECS-CROP sensors made available to the TOMRES project, in order to describe and map spatial variability of field-grown tomato. All collected data have been statistically processed, analyzed and interpreted in order to assess the potential in increasing tomato stress resilience, production and quality by means of the tested smart farming techniques (variable rate fertilization, irrigation and spraying).</p>
Multi-parameter thematic maps for VRT (2019)
<p>“Optimization of VRT technology to the resilient tomato lines”, during year 2019 an additional set of field surveys has been conducted by Casella using the MECS-CROP sensors made available to the TOMRES project, in order to describe and map spatial variability of field-grown tomato. All collected data have been statistically processed, analyzed and interpreted in order to assess the potential in increasing tomato stress resilience, production and quality by means of the tested smart farming techniques (variable rate fertilization, irrigation and spraying).</p>
Multi-parameter thematic maps for VRT (2018)
<p>"Optimization of VRT technology to the resilient tomato lines", during years 2017 and 2018 an extensive set of field surveys has been conducted by Casella using the MECS-CROP sensors made available to the TOMRES project, in order to describe and map spatial variability of field-grown tomato. All collected data have been statistically processed, analyzed and interpreted in order to assess the potential in increasing tomato stress resilience, production and quality by means of the tested smart farming techniques (variable rate fertilization, irrigation and spraying).</p>
Land cover classification using Landsat Enhanced Thematic Mapper (ETM) data - year 2000
This land cover classification map was created using Landsat Enhanced Thematic Mapper (ETM) data from the year 2000. The map covers the area of the Central Arizona-Phoenix Long Term Ecological Research study.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.