Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “Science of science”
SolSysELTs2022 Part I: ESA Planetary Science missions + Q&A
<p>Presentation and video recording</p>
SolSysELTs2022 Part II: Puzzling origin of water on Earth: frontiers in cometary science with METIS/EELT
<p>Contributed talk: presentation and video recording</p>
Supplementary material 1 from: Patterson DJ (2022) The scope and scale of the life sciences ('Nature's envelope'). Research Ideas and Outcomes 8: e96132. https://doi.org/10.3897/rio.8.e96132
Patterson Nature's envelope (template)
Visualization of the Twitter Science Graph: 2022-11-15
<p><strong>The Twitter Science Graph</strong></p> <p>State 2022-11-15, 1652 nodes, 17235 edges.</p>
Citizen science project descriptions as science communication texts - the good, the bad, and the ugly
<p>This study aimed to determine to what extent do CS project descriptions actually contain the kinds of information relevant to prospective participants and whether this information is conveyed in a comprehensible and attractive manner. To this end, we conducted a qualitative content analysis of a random sample of 120 English-language project descriptions stored in the CS Track database. The coding rubric used for this study is based on the ten-step template for writing engaging project descriptions we recently designed and published. The sample was produced as follows: After creating a dataset containing only English-language project descriptions, we excluded all descriptions which consist of less than 100 or more than 500 words. Texts of less than 100 words cannot be expected to contain a significant amount of information. Project descriptions of more than 500 words are less likely to be read in their entirety than shorter texts and thus ill-suited to the task of capturing the readers’ interest and prompting them to join the project in question. Finally, we applied the ‘random’ function of RStudio to randomly select 120 texts from the resulting dataset of 1283 descriptions. The qualitative content analysis was performed in two consecutive steps. First, in order to ensure that the coding rubric is fit for purpose and all categories within it well-defined and demarcated, all three members of the research team independently coded 40 project descriptions. After discussing the results and making slight modifications to the coding rubric, each team member coded roughly one third of the remaining 80 descriptions. </p> <p>Preliminary results suggest that the majority of project descriptions in our sample fail to mention how citizen scientists will benefit from participating, what kind of training they will receive, how their contributions will be acknowledged, and whether they will have access to project results. Furthermore, the project’s goals, its target audience, and the tasks volunteers will be expected to complete are very often not described explicitly and clearly enough. For instance, very few project descriptions contain concrete information on required skills and equipment or on the time commitment associated with participation.</p> <p>This dataset contains the coding rubric used to analyse project descriptions and a visualisation of preliminary results.</p>
How to make a difference in science policy with your research - Webinar
<p>As part of Project Ô's communication efforts, two international events were organised aimed at extended networks with a focus on political communication and science advocacy relevant to water reuse. These online public events were recorded and edited for publication on the YouTube channel.</p> <p>--</p> <p>On Monday 21 November 2022, a Project Ô-organised webinar was offered on the topic of how to make a difference in science policy with research. It was organised by the Institute for Methods Innovation on behalf of the project, working in collaboration with renowned science policy strategist Dr Andrew George (Sigma Xi). This live event was aimed at water sustainability researchers, professionals and others interested in the links between research and public policy.</p>
Data attached to Radio Science paper "Model to Scale Rain Attenuation Time Series with Link Elevation Angle for LEO Satellite Based Systems"
<p>Data attached to Radio Science paper "Model to Scale Rain Attenuation Time Series with Link Elevation Angle for LEO Satellite Based Systems"</p>
Supplementary material 2 from: Baskauf SJ, Girón Duque JC, Nielsen M, Cobb NS, Singer R, Seltmann KC, Kachian Z, Pérez M, Agosti D, Klompen AML (2023) Implementation Experience Report for Controlled Vocabularies Used with the Audubon Core Terms subjectPart and subjectOrientation. Biodiversity Information Science and Standards 7: e94188. https://doi.org/10.3897/biss.7.94188
Views Controlled Vocabularies Implementation Reporting Form
Supplementary material 1 from: Baskauf SJ, Girón Duque JC, Nielsen M, Cobb NS, Singer R, Seltmann KC, Kachian Z, Pérez M, Agosti D, Klompen AML (2023) Implementation Experience Report for Controlled Vocabularies Used with the Audubon Core Terms subjectPart and subjectOrientation. Biodiversity Information Science and Standards 7: e94188. https://doi.org/10.3897/biss.7.94188
Views Controlled Vocabularies testing notes
Materials Science Optimization Benchmark Dataset for Multi-fidelity Hard-sphere Packing Simulations
Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a "Turing test" of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In the fields of materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 438371 random hard-sphere packing simulations representing 279 CPU days worth of computational overhead were performed across nine input parameters with linear constraints and two discrete fidelities each with continuous fidelity parameters and results were logged to a free-tier shared MongoDB Atlas database. Two core tabular datasets resulted from this study: 1. a failure probability dataset containing unique input parameter sets and the estimated probabilities that the simulation will fail at each of the two steps, and 2. a regression dataset mapping input parameter sets (including repeats) to particle packing fractions and computational runtimes for each of the two steps. These two datasets can be used to create a surrogate model as close as possible to running the actual simulations by incorporating simulation failure and heteroskedastic noise. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This is in contrast with a more traditional approach that imposes a-priori assumptions such as Gaussian noise e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.
Dataset about sharks in French Polynesia extracted from the ORP citizen science network
<p><span>Observers of the Polynesian Shark Observatory (ORP), a citizen science network organized mainly through the Polynesian dive centers, collected an unprecedented amount of data in 43% of the islands of French Polynesia between July 8, 2011, and April 11, 2018, during 13,916 dives. The objective of such a data collection, not accessible to standard research resources, was to provide a unique dataset, offering the opportunity to explore the specific diversity, distribution, seasonality and abundance of many elasmobranch species spread out throughout the overall French Polynesia territory. Indeed, since the data are based on random citizen observation, the spatial distribution was biased toward the most frequented sites and islands where the scuba diving activity is mostly developed. Overall, the increase in observed abundance of rays and sharks observed in French Polynesia and the three most sampled islands, as well as the high specific diversity recorded for the region, provide the first evidence of the effectiveness of the Shark Sanctuary established in 2006. These data, collected randomly by the volunteers, also give insights about potential movement patterns and site fidelity of some species more often seen. While not leading to final conclusion, the network of volunteers filling with information to the Polynesian Shark Observatory is essential in giving preliminary results and research perspectives and directions for future projects on sharks and rays in French Polynesia.</span></p>
De-identified article and author characteristics for a large data set of Web of Science
<p>This data set contains article and author characteristics for all records in the Web of Science, 2000-2020. Standard article identifiers have been removed and replaced with a document ID (`doc_id`), as linking to the original ID is not permitted.</p>
Open Science and Open Innovation database
<p>Open Science and Open Innovation database </p>
The Brill Knowledge Graph: A Database of Bibliographic References and Index Terms extracted from Books in Humanities and Social Sciences
<p>We present a complete dataset of linked bibliography and index data, partially disambiguated and augmented with references to external resources, extracted from the Brill’s archive in the field of Classics. Processed book identifiers are listed in a separate text file. Text fragments extracted from different books via this process are then parsed and compared using a string-based similarity metric to form clusters of bibliographic references to the same published work or (variants of) the same subjects discussed in these books. The entire set of references was then disambiguated using Google Books and Crossref APIs.</p> <p><a href="https://jdmdh.episciences.org/11062">Paper about extraction pipeline</a></p> <p><a href="https://www.nkokash.com/documents/KIEM-RDJ.pdf">Paper about extracted KG</a></p> <p> </p>
Dataset from web of science about public revenue in Indonesia
<p>This is the dataset of meta-data of papers used for the samples of the study</p>
OpenAlex slices for "Collaboration and topic switches in Science"
<p>OpenAlex slices stored as zipped parquet files. Needs pandas >= 2, pyarrow >= 7.</p>
Codes and data related to the article: Renard (2023). Use of a National Flood Mark Database to Estimate Flood Hazard in the Distant Past. Hydrological Sciences Journal.
<p>This package contains R codes and data related to the article:</p> <p>B. Renard. 2023. Use of a National Flood Mark Database to Estimate Flood Hazard in the Distant Past<em>. Hydrological Sciences Journal</em>. DOI: <a href="https://doi.org/10.1080/02626667.2023.2212165">10.1080/02626667.2023.2212165</a></p> <p><strong>CONTENTS</strong></p> <p>R scripts used to set up models, analyse results and prepare figures.</p> <ul> <li>France207_MAX.RData contain monthly maxima at 207 stations. See references below for the original sources.</li> <li>Folder data_raw/ contain exports from the flood marks database. See references below for the original sources.</li> <li>funk.R is a library of functions used by other scripts.</li> <li>0_setUp.R prepares the runs to estimate the model for one component. Estimation is performed using the RSTooDs package (<a href="https://zenodo.org/record/5075760">https://zenodo.org/record/5075760</a>). This script extracts the data, specifies the probabilitic models and write STooDs configuration files.</li> <li>1_margin.R analyses the runs that have been performed and estimates marginal distributions at each station. Note that the runs set up in the previous point need to have been performed before using this script.</li> <li>2_probMap.R computes estimated flood probabilities over the whole 1705-2015 period. The previous script needs to have been run before.</li> <li>3_sensitivityAnalysis.R prepares configuration files for sensitivity analysis experiments.</li> <li>4_ZenodoRelease.R prepares the released data and probability maps. The previous scripts needs to have been run before.</li> <li>Scripts named fig_XXX.R create the figures shown in the article.</li> </ul> <p><strong>REFERENCES</strong><br> Streamflow series at stations: <a href="https://www.hydro.eaufrance.fr/">https://www.hydro.eaufrance.fr/</a><br> developed by the "Service central d’hydrométéorologie et d’appui à la prévision des inondations" (Schapi) from the Ministry of the Ecological Transition</p> <p>Flood marks database: <a href="https://www.reperesdecrues.developpement-durable.gouv.fr/">https://www.reperesdecrues.developpement-durable.gouv.fr/</a><br> developed by the flood forecasting network "Vigicrues" from the Ministry of the Ecological Transition</p> <p> </p>
Knowledge organization systems as enablers to the conduct of science
<p>The sophistication of knowledge organization systems (KOS) has evolved rapidly over the past thirty years, largely driven by information technology innovations. Two key assumptions have been (a) that KOS-work is the preserve of information professionals acting as skilled intermediaries, and (b) that it is largely focused on enabling the finding and discovery of information. This paper challenges both assumptions with reference to the conduct of science in the 21st century, by describing the ways in which access to KOS skills and tools is already broadening beyond information professionals to scientists, and by describing how knowledge organization systems enable sense-making of trends within science and new knowledge creation, beyond simple access and discovery roles. It closes with remarks on the implications for information professionals engaged in KOS-related work.</p>
Business process management concept. Bibliographic collection from Web of Science (March 21, 2023).
<p>The search query: https://www.webofscience.com/wos/woscc/summary/f99d7b5d-a7bb-4c35-a8ba-d45042c4e0a9-7ab6bd8b/relevance/1. Contains 95 articles (adjusted after screening the relevance of titles, abstracts, keywords for the research purposes).</p>
Supplementary material 3 from: Colombari F, Battisti A (2023) Citizen science at school increases awareness of biological invasions and contributes to the detection of exotic ambrosia beetles. In: Jactel H, Orazio C, Robinet C, Douma JC, Santini A, Battisti A, Branco M, Seehausen L, Kenis M (Eds) Conceptual and technical innovations to better manage invasions of alien pests and pathogens in forests. NeoBiota 84: 211-229. https://doi.org/10.3897/neobiota.84.95177
Number of individuals of each species captured at each school
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.