Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “Data Scientists”
Data and analysis script for "The (non)effect of personalization in climate texts on credibility of climate scientists: A case study on sustainable travel"
<p>Dataset and analysis script for the article "<strong>The (non)effect of personalization in climate texts on credibility of climate scientists</strong><strong>: A case study on sustainable travel</strong>", under review at Geoscience Communication (https://doi.org/10.5194/egusphere-2024-543)</p>
Missouri reservoir water quality data from the Statewide Lake Assessment Program (SLAP), the Lakes of Missouri Volunteer Program (LMVP), and the Reservoir Observer Student Scientists (ROSS) program
This dataset of limnological water quality data continues from Jones et al., 2024, starting in 2017 until 2021. It is from 195 reservoirs, the majority of which are in the state of Missouri (MO) in the USA collected by the University of Missouri Limnology Lab. Water quality parameters analyzed in the MU Limnology Lab during this time frame include: areal pigment absorption coefficient, alkalinity, alpha (light utilization efficiency P-E parameter), ammonium (NH4), ammonium-debt, anatoxin, chlorophyll a (corrected and uncorrected for pheophytins), chloride, cylindrospermopsin, seston d13C, seston d15N, dissolved turbidity, dissolved organic carbon, Ek (light saturation P-E parameter), FVFM (maximum quantum yield of PSII for photochemistry), gross primary production, microcystin, nitrate & nitrite (NO3), nitrate-debt, particulate nitrogen, particulate phosphorus, phosphorus-debt, pheophytin, particulate carbon, particulate inorganic matter, particulate organic matter, phycocyanin (PHYCO), saxitoxin, Secchi disk depth, silica, soluble reactive phosphorus, total dissolved nitrogen (TDN), total dissolved phosphorus (TDP), total nitrogen (TN), total phosphorus (TP), total suspended solids (TSS), and urea. Most of the samples were collected during the summer months (May-September) when the reservoirs were thermally stratified, but a few were taken during the rest of the year (October-April). The majority of samples were taken at the deepest point in the reservoir directly up-reservoir of the dam. Sampling was conducted from a boat most of the time, but a few samples were taken from shorelines and drinking water treatment intake pipes. Most of the data come from the Statewide Lake Assessment Project (SLAP) and the Lakes of Missouri Volunteer Program (LMVP) funded by the Missouri Department of Natural Resources. This data represents duplicate or triplicate water samples collected from either the water surface, integrated over the depth of the epilimnion, or from discrete dep
Open Research Data in Medicine - Polish scientists' attitudes towards data sharing
<p>The survey on the attitudes and beliefs of research staff has been carried out at selected Polish medical universities. The purpose of the questionnaire was to collect respondents' opinions on opening research data created during their scientific work. The research was aimed at preparing the necessary educational, technical and legal support for scientists after launching the Polish Medical Platform, when Polish scientists will be asked to deposit their data in local repositories.</p>
Results of the poll in the study "Information Scientists' Motivations for Research Data Sharing and Reuse"
<p>This is a dataset with results of the poll conducted in the study “Information Scientists’ Motivations for Research Data Sharing and Reuse”.</p> <p>In terms of the Uses and Gratifications Theory (Questions 1 and 2), the most popular uses relate to the categories of research support and information. Researchers share, or would share, their research data in general for any reusability purposes and especially for combination of different datasets to produce new evidence. Also, the vast majority of study participants associate research data sharing with possibilities to accelerate scientific progress and to increase research efficiency. In case of research data reuse, all the researchers indicated that they use, or would use, others’ data first of all for inspiration. Interestingly, study participants put relatively high the category of recognition in case of sharing, but at the same time they do not associate increased recognition among colleagues and other researchers with research data reuse. The remaining categories belonging to the categories of self-esteem and social interaction, i.e. increased citation level and visibility of the research as well as enhanced scientific reputation, possible cooperations and co-authorship, were selected only by few respondents. Also remarkably, data reuse is more frequently linked to entertainment then data sharing. </p> <p>In terms of the Self-Determination Theory (Questions 3 and 4), all but one of the interviewees indicated that they have shared or would share their research data because it can accelerate scientific progress which they consider important and would like to contribute to it (i.e., identified regulation). The second most popular motivation turned out to be the obligation by employer, project funder and/or journals (i.e., external regulation). The third most popular option was social influence, i.e. because many other researchers participate in data sharing and they feel obligated to do the same (i.e., external regulation).This way, the participants demonstrate a mixture of identified motivation and external regulation, both material and social. In the case of data reuse, the participants demonstrate more homogeneous results with identification and intrinsic motivation having most of the votes. The role of external regulation seems to be much less important as in the case with data sharing. So, researchers reuse, or would reuse, research data because it can accelerate scientific progress which is important for them. Additionally, researchers enjoy exploring and using third party research data. Thus, interviewees participate or would participate in data sharing because they consider it important, but also feel or are obliged to do so. At the same time, study participants do not feel pressure from outside when deciding whether to reuse data or not.</p> <p>For more information about the study and its results, please read the article “Information Scientists’ Motivations for Research Data Sharing and Reuse” by Shutsko and Stock (2023).</p>
Data from: More than a token photo: humanising scientists enhances student engagement
Open the record for dataset details and reuse information.
Data from: Continent‐scale phenotype mapping using citizen scientists' photographs
Field investigations of phenotypic variation in free‐living organisms are often limited in scope owing to time and funding constraints. By collaborating with online communities of amateur naturalists, investigators can greatly increase the amount and diversity of phenotypic data in their analyses while simultaneously engaging with a public audience. Here, we present a method for quantifying phenotypes of individual organisms in citizen scientists' photographs. We then show that our protocol for measuring wing phenotypes from photographs yields accurate measurements in two species of Calopterygid damselflies. Next, we show that, while most observations of our target species were made by members of the large and established community of amateur naturalists at iNaturalist.org, our efforts to increase recruitment through various outreach initiatives were successful. Finally, we present results from two case studies: (1) an analysis of wing pigmentation in male smoky rubyspots (Hetaerina titia) showing previously undocumented geographical variation in a seasonal polyphenism, and (2) an analysis of variation in the relative size of the wing spots of male banded demoiselles (Calopteryx splendens) in Great Britain questioning previously documented evidence for character displacement. Our results demonstrate that our protocol can be used to create high quality phenotypic datasets using citizen scientists' photographs, and, when combined with metadata (e.g., date and location), can greatly broaden the scope of studies of geographical and temporal variation in phenotypes. Our analyses of the recruitment and engagement process also demonstrate that collaborating with an online community of amateur naturalists can be a powerful way to conduct hypothesis‐driven research aiming to elucidate the processes that impact trait evolution at landscape scales.
Evaluation data for: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study
<p>All of the evaluation data for the simulations in the paper: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study. We considered the impact of five adaptive sampling methods on the performance of species distribution models (SDMs), please see the paper for more information. Contained in this repository are the evaluation metrics (AUC, mean square error (MSE) and correlation) for SDMs before and after adaptive sampling has taken place. The MSE and correlation evaluation metrics were calculated against the true distributions of the species. These files are those with "combined_outputs" in the titles. The repository also contains the observations of all the species in the simulations both before and after adaptive sampling (the files with "all_observations" in the title.</p> <p>These datasets are to be used with the plotting and evaluation scripts in the GitHub repository associated with the paper.</p>
Explorant Oportunitats: Anàlisi d'Ofertes d'Ocupació per a Data Scientists a Europa a través de Glassdoor.
<div> <div> <div> <p>L'exploració d'oportunitats professionals per a científics de dades a Europa s'ha realitzat mitjançant l'anàlisi de dades extretes de Glassdoor, una plataforma de recerca d'ocupació. Aquest estudi s'ha centrat en examinar les ofertes d'ocupació disponibles per a aquesta ocupació en el mercat laboral europeu. Mitjançant tècniques d'scraping de dades, s'han recopilat i analitzat detalladament les característiques clau de les ofertes, com ara requisits, ubicacions, salaris i beneficis, proporcionant una visió completa de les tendències i oportunitats per als científics de dades a Europa. Aquesta anàlisi exhaustiva pot ser útil tant per als professionals que busquen noves oportunitats laborals com per a les empreses i institucions interessades en comprendre millor la demanda i les expectatives en aquest camp en creixement.</p> </div> </div> </div>
Data of "Using current research information systems to investigate data acquisition and data sharing practices of computer scientists"
<p>This study describes a methodology where departmental academic publications are used to analyse the ways in which computer scientists share research data.</p> <p>Without sufficient information about researchers’ data sharing, there is a risk of mismatching FAIR data service efforts with the needs of researchers. This study describes a methodology where departmental academic publications are used to analyse the ways in which computer scientists share research data. The advancement of FAIR data would benefit from novel methodologies that reliably examine data sharing at the level of multidisciplinary research organisations. Studies that use CRIS publication data to elicit insight into researchers’ data sharing may therefore be a valuable addition to the current interview and questionnaire methodologies.</p> <p><strong>Data was collected from the following sources:</strong></p> <p>All journal articles published by researchers in the computer science department of the case study’s university during 2019 were extracted for scrutiny from the current research information system. For these 193 articles, a coding framework was developed to capture the key elements of acquiring and sharing research data. Article DOIs are included in the research data.</p> <p>The scientific journal articles and theirs DOIs are used in this study for the purpose of academic expression.</p> <p>The raw data is compiled into a single CSV file. Rows represent specific articles and columns are the values of the data points described below. Author names and affiliations were not collected and are not included in the data set. </p> <p> </p> <p>The following data points were used in the analysis:</p> <p><strong>Data points</strong></p> <ul> <li><strong>Main study types</strong></li> <li>Literature-based study (e.g. literature reviews, archive studies, studies of social media)</li> <li>yes/no</li> <li>Novel computational methods (e.g. algorithms, simulations, software)</li> <li>yes/no</li> <li>Interaction studies (e.g, interviews, surveys, tasks, ethnography)</li> <li>yes/no</li> <li>Intervention studies (e.g., EEG, MRI, clinical trials)</li> <li>yes/no</li> <li>Measurement studies (e.g. astronomy, weather, acoustics, chemistry)</li> <li>yes/no</li> <li>Life sciences (e.g. “omics”, ecology)</li> <li>yes/no</li> <li><strong>Data acquisition</strong></li> <li>Article presents a data availability statement</li> <li>yes/no</li> <li>Article does not utilise data</li> <li>yes/no</li> <li>Original data was collected</li> <li>yes/no</li> <li>Open data from prior studies were used</li> <li>yes/no</li> <li>Open data from public authorities, companies, universities and associations</li> <li>yes/no</li> <li><strong>Data sharing</strong></li> <li>Article does not use original data</li> <li>yes/no</li> <li>Data of the article is not available for reuse</li> <li>yes/no</li> <li>Article used openly available data</li> <li>yes/no</li> <li>Authors agree to share their data to interested readers</li> <li>yes/no</li> <li>Article shared data (or part of) as supplementary material</li> <li>yes/no</li> <li>Article shared data (or part of) via open deposition</li> <li>yes/no</li> <li>Article deposited code or used open code</li> <li>yes/no</li> </ul>
Towards an ELSA curriculum for Data Scientists
<p>Review of the existing approaches and identification of the challenges with respect to the context of developing a curriculum for practitioners, such as variance in experience, education, and cultural background.</p>
Data Scientists and ELSA:Describing the Landscape
<p>A brief description of the status quo: ELSA as part of the Data Science curriculum in tertiary education; ELSA awareness demands in industry; determining the Data Scientist profile now and in the future</p>
Towards creating an ELSA Curriculum for Data Scientists - Introduction
<p>An introduction to the first ELSA workshop within the FAIR Data Spaces project. After a short overview of the project, the aim of the ELSA curriculum and of this workshop are presented.</p>
Data collected by citizen scientists reveal the role of climate and phylogeny on the frequency of shelter types used by frogs across the Americas
<p><strong>Supplementary Material </strong> of the manuscript on frog shelters across americas. Raw data and figures.</p> <p><strong>Table S1.</strong> The number of all anuran observations per species per country in the iNaturalist dataset on July 1, 2021.</p> <p><strong>Table S2.</strong> Distribution of shelter types among frog species retrieved from the iNaturalist platform across the Americas. Species with shelter first time reports are marked with *, shelter types are: AH = Holes in an artificial material; AS = Artificial surface; HG = Hole in the ground; HR = Among rocks; VBa = Hollow in bamboo; VBr = Bromelias or bromelia-like plants; VL = Live vegetation; VT = Tree trunk; VD = Leaf litter; WA = Water</p> <p><strong>Table S3.</strong> Anuran shelter sites around the world based on our literature review. *indicates the reference is secondary literature</p>
Explorant Oportunitats: Anàlisi d'Ofertes d'Ocupació per a Data Scientists a Europa a través de Glassdoor amb dades l'índex de felicitat, el PIB per càpita i l'esperança de vida.
<p>L'exploració d'oportunitats professionals per a científics de dades a Europa s'ha realitzat mitjançant l'anàlisi de dades extretes de Glassdoor, una plataforma de recerca d'ocupació. Aquest estudi s'ha centrat en examinar les ofertes d'ocupació disponibles per a aquesta ocupació en el mercat laboral europeu amb dades l'índex de felicitat, el PIB per càpita i l'esperança de vida. Mitjançant tècniques d'scraping de dades, s'han recopilat i analitzat detalladament les característiques clau de les ofertes, com ara requisits, ubicacions, salaris i beneficis, proporcionant una visió completa de les tendències i oportunitats per als científics de dades a Europa. Aquesta anàlisi exhaustiva pot ser útil tant per als professionals que busquen noves oportunitats laborals com per a les empreses i institucions interessades en comprendre millor la demanda i les expectatives en aquest camp en creixement.</p>
Data Science Tasks used in "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"
<p>This entry contains the supplementary files for a scientific article. </p> <p>The dataset contains the necessary files for the two data science tasks used in the experiment study from scientific article "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"</p> <p> </p>
Data from Evaluating the impact of climate communication activities by scientists: what is known and necessary?
<p>This dataset contains three text files in RIS format and two pdf files. They represent the analysis for "Evaluating the impact of climate communication activities by scientists: what is known and necessary?" (https://doi.org/10.5194/gc-7-91-2024). </p> <p>There are three stages</p> <ol> <li>The initial search: <ul> <li> <p>The <em>Literature search PE on CC.pdf</em> file provides the initial search terms</p> </li> <li> <p>The <em>Search PE on CC 'Outreach added'.pdf</em> file expands that initial search to include outreach</p> </li> <li>The <em>Initial search (outreach included).ris</em> file contains references to all 819 articles included in the initial search</li> </ul> </li> <li>Excluded first round (based on title and abstract) <ul> <li>The <em>Excluded first round.ris</em> file contains references to all 755 articles excluded in the first round</li> </ul> </li> <li>Included papers <ul> <li>The <em>Included papers.ris </em>file contains references to all 7 articles that were included in the final analysis</li> </ul> </li> </ol>
Data for Ferretto et al, 2021 "Naturally-detached fragments of the endangered seagrass Posidonia australis collected by citizen scientists can be used to successfully restore fragmented meadows"
<p>Please find attached the data for the manuscript "Ferretto et al, 2021" and a brief description of each file.</p>
Students as citizen scientists: project-based learning through the iNaturalist platform could provide useful biodiversity data
<p>This is a compilation of data from iNaturalist platform based on the geographic limits of brazilian semiarid until 28 october, 2022.</p>
Data and code accompanying "'Safe spaces' and community building for climate scientists, exploring emotions through a case study", Haddaway and Duggan 2023
<p>Data and code accompanying "‘Safe spaces’ and community building for climate scientists, exploring emotions through a case study", Haddaway and Duggan 2023</p>
Welcome-Introduction to the Workshop " A first approach to an ELSA Curriculum for Data Scientists The FAIR Data Spaces Project as a Use Case
<p>Welcome-Introduction to the Workshop " A first approach to an ELSA Curriculum for Data Scientists The FAIR Data Spaces Project as a Use Case", contains a brieff description of the FAIR Data Spaces project</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.