Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “Science of science”
The Commercial Potential of Science
<div> <div> <div> <div> </div> </div> </div> <div> <p>[<strong>c</strong><strong>oming soon: commercial and scientific potential predictions for over 30 million articles, published worldwide</strong>]</p> <p> </p> <p>This dataset introduces a novel index designed to predict the commercial potential of scientific articles. The index captures the probability that an article will be used by firms for the development of marketable products or processes. In addition to commercial potential, the dataset also introduces an index to predict scientific potential—the likelihood that an article will be relevant for the advance of science, regardless its commercial application. </p> <p>The indices are crucial for researchers focused on understanding 1) the production of science with commercial potential and 2) the pathway from academic research to market innovations and the factors that influence the commercial viability of scientific discoveries.</p> <p> </p> <p><strong>Citation Information:</strong> If you use this dataset, please cite the article: “Masclans-Armengol, R., Hasan, S., & Cohen, W. M. (2024). Measuring the Commercial Potential of Science. NBER working paper”</p> <p> </p> <p><strong>Components of the Dataset: </strong>The dataset encompasses indices for over 5.2 million articles that meet the following criteria:</p> <ul> <li>Publication year: 2000 to 2020</li> <li>Published under 126 U.S. universities</li> <li>Articles in the applied and natural sciences and engineering fields</li> </ul> <p>Data is delivered via a single csv file. Each row contains information for a scientific article, with the following variables:</p> <ul> <li>‘<em>doi’</em>: Digital Object Identifier—unique article identifier that can be used to match to other data sources, such as OpenAlex, Dimensions, or Web of Science.</li> <li>‘<em>compot</em>’: commercial potential index.</li> <li>‘<em>scipot</em>‘: scientific potential index.</li> </ul> <p>To develop the commercial potential index, we employed SciBert (Beltagy et al., 2019), a Large Language Model for scientific understanding. We fine tune SciBert with deep neural networks to classify scientific articles based on their potential for commercial application. We trained 20 predictive models, one per year, using the text of an academic article’s abstract to generate ex-ante, out-of-sample, and out-of-training-time-period predictions of any given scientific article’s commercial potential.</p> <p>Following the same methodology, we compute the scientific potential index.</p> <p> </p> <p><strong>Licensing and Contact Information: </strong>The dataset and its components are distributed under a Creative Commons Attribution Non-Commercial license.</p> <p> </p> <p><strong>Acknowledgments</strong>: We thank The Technology Opportunity Lab at Duke University and the Kauffman Foundation for funding the creation of this dataset.</p> </div> </div>
Retractions in Humanities and Social Sciences: A Study of Retracted Papers from China
<p>This dataset presents a comprehensive study of retractions in the field of Humanities and Social Sciences (HSS), focusing specifically on retracted papers originating from China. </p>
Assessing the Overlap of Science Knowledge Graphs: A Quantitative Analysis — exact and related matches
<p>Results of the 'Assessing the Overlap of Science Knowledge Graphs: A Quantitative Analysis' papers. There are 2 datasets:</p> <ul> <li>'exact_matches.csv': contains detailed information about the concepts present both in OpenAlex and OpenAIRE.</li> <li>'related_matches.csv': contains detailed information about the concepts from OpenAlex and OpenAIRE that were not present in both KGs but got aligned following the algorithm presented in the paper.</li> </ul> <p>The detailed information refers to the following column:</p> <ul> <li>Category1: name of the first category</li> <li>Source1: source of the first category ('OpenAlex' or 'OpenAIRE')</li> <li>Category2: name of the second category</li> <li>Source2: source of the first category ('OpenAlex' or 'OpenAIRE')</li> <li>Similarity: semantic similarity value of the two categories</li> <li>PapersInC1: number of papers from the collected dataset belonging to the first category</li> <li>PapersInC2: number of papers from the collected dataset belonging to the second category</li> <li>PapersInBoth: number of papers from the collected dataset belonging to both of the categories</li> <li>Agreement: the value of the agreement of the categories in the tw KGs (Intersection over Union)</li> </ul>
SUPPLEMENTARY DATA TO: Using a citizen science approach to assess nanoplastics pollution in remote high-altitude glaciers
<p>This is the repository of the supplementary data, and it contains the following files: </p> <p>Raw data files as the original output of TD-PTR-ToF-MS for all the samples, all the blanks, all the spikes and all the calibration runs (.h5 files in three zip arcives)</p> <p>Polymer library files (a zip archive including csv files.</p> <p>A data analysis file including raw data, blank subtraction and LOD correction of all measurements (xlsx file).</p> <p>A fingerprinting result file for each plastic type (xlsx file)</p> <p>A data analysis file after plastic fingerprinting (xlsx file). </p> <p> </p> <p> </p>
Dataset for: Mechanisms of uropathogenic E. coli mucosal association in the gastrointestinal tract (Science Advances)
<p>This dataset contains the plotted data for Figures 1B, 1B, 1F, Figure 4C, Figures 5C, 5E, 5G, 5H, 5J, 5K, Figure S5B, Figure S6A, S6B, S6C, Figure S7 A, S7C, S7D, and Figure S8C.</p>
Data supporting research on journalistic production on science by press offices of universities and university centers in Rio Grande do Sul (2016)
<p>These data are part of a final paper called Scientific Journalism in Community Higher Education Institutions in Rio Grande do Sul. The work was developed at the University of Vale do Taquari - Univates, between 2016 and 2017. The data is in Portuguese. The abstract of the paper is available below. Science occupies an important place in today's society. It is what has allowed us to reach our current stages of intellectual development and also to achieve memorable feats as a species. Science is produced, to a large extent, in the academic environment. This monograph focuses its efforts on trying to understand how 15 universities and university centers in the state of Rio Grande do Sul, partners in the Consortium of Community Universities of Rio Grande do Sul (Comung), carry out the dissemination of their academic production. It is understood that it is important to disseminate scientific information to the population so that individuals can critically evaluate the actions developed in the field of science. The general objective of this study is to investigate the production of scientific news in higher education institutions linked to Comung, as well as to characterize the relationships between the actors and processes related to the journalistic dissemination of science produced in these same institutions. The qualitative and quantitative analysis of the scientific dissemination texts was directed to the news published by the organizations, through an exploratory study that mapped all the production of the press offices between January and August 2016. Subsequently, a qualitative analysis of the discourse of press officers was carried out, after applying a questionnaire. Throughout the investigation, the hypothesis that the institutions disseminate their scientific production was confirmed, but they do so in markedly different ways. The study is available at this link: https://www.univates.br/bdu/items/4cdf0835-1fbe-425a-a216-2a511e9aabf8. </p>
VABB-SHW: Dataset of Flemish Academic Bibliography for the Social Sciences and Humanities (edition 14)
<p>This dataset contains the fourteenth edition of the <em>Flemish Academic Bibliography for the Social Sciences and Humanities (VABB-SHW)</em>, a database of academic publications from the social sciences and humanities authored by researchers affiliated to Flemish universities (<a href="https://www.ecoom.be/en/data-collections/vabb-shw">more information</a>). Publications in the database are used as one of the parameters of the Flemish performance-based research funding system. Only approved publications are included in this dataset.</p>
Data from "The academic impact of Open Science: a scoping review"
<p>These files include the data from the "The societal impact of Open Science - a scoping review", part of a series of studies conducetd within the PathOS Horizon Europe project on the academic, economic, and societal impacts of Open Science. This study was conducted in two phases. In phase 1 an academic database search was conducted. For phase 2 an automatic snowball search was performed based on results from phase 1 ( and grey literature was searched manually.</p> <p>The upload contains five files:</p> <ol> <li>Main file with extracted information for the 485 studies included in the review ("academic_impact_included_all_data.csv").</li> <li>Excel file documenting the grey literature search ("Grey_Literature_Search.xlsx").</li> <li>R project for the snowball search ("academic_impact_snowball.zip").</li> <li>R project for cleaning data and producing summaries and figures ("academic_impact_stats.zip").</li> <li>Excel file mapping the frascati codes used in file (1) to their textual representations ("Frascati definitions.xlsx")</li> </ol> <p>For more details on the methods see the <a href="https://osf.io/m4rnc">protocol</a> and its <a href="https://osf.io/3b6xj">addendum</a>. For the background, results and discussion see the <a href="../records/7883699">deliverable</a> (reporting on phase 1) and the pre-print.</p>
Dataset - A Python-Based Approach to Sputter Deposition Simulations in Combinatorial Materials Science
<p>This dataset accompanies the publication <em>"A Python-Based Approach to Sputter Deposition Simulations in Combinatorial Materials Science,"</em> which presents and validates pySIMTRA, a Python wrapper for the Monte Carlo-based SIMTRA simulation tool. The dataset includes all measured and simulated data shown in the publication, as well as additional animations visualizing the compositions in the multinary composition space.</p> <p>The dataset contains the compositions for each of the seven materials libraries (in at.%) in the quaternary Ni-Pd-Pt-Ru system. Additionally, it provides the simulated number of particles as outputted by SIMTRA, which serve as the basis for composition estimation. Both the compositional data and the particle counts are supplied in .csv format. To supplement the results, 3D animations of the quaternary compositional spaces are included, showing the comparison between simulated compositions (red dots) and measured compositions (blue dots). These animations offer a more intuitive visualization of the data compared to the static Figures in the publication and are supplied as .gif files.</p> <p>Due to the in-depth analysis of cathode tilt discussed in the paper, the dataset also includes simulation results for the ternary Pd-Pt-Ru library, highlighting the effect of varying the cathode tilt angle. Simulations were conducted for tilt angles of 10°, 9.5°, 9°, and 8.5°.</p>
Radio Science observations of Mars Express and Tianwen-1 spacecraft from the 2021 Maritan Solar Conjunction
<p>Phase scintillation dataset from the 2021 Martian solar conjunction is contained in the DATA.zip file. The Read Me pdf file contiants information useful to understanding the file naming conventions and their contents.</p>
Data and code for: "Global Sampling Decline Erodes Science Potential of Natural History Collections"
<p># GBIF Specimen Data Analysis and Forecasting<br><br>## Version 2 - modified date ranges for figures 1 and 2 in response to reviewer comments</p> <p>This repository contains the code and data for analysing and forecasting trends in Global Biodiversity Information Facility (GBIF) specimen records across three major taxonomic groups: Chordata, Arthropoda, and Plantae. <br>The analysis pipeline includes data cleaning, anomaly detection, primary analyses, and forecasting based on historical database snapshots.</p> <p>These scripts and data correspond to analyses in the following manuscript:</p> <p>Global Sampling Decline Erodes Science Potential of Natural History Collections</p> <p>Authors:<br>Owen Forbes<br>Andrew G. Young<br>Peter H. Thrall</p> <p><br>## Repository Structure</p> <p>The repository consists of three main Quarto (.qmd) scripts and associated data files:</p> <p>1. `1_DataCleaning_Forbes-et-al_2025.qmd`: Data cleaning and anomaly detection<br>2. `2_PrimaryAnalyses_Forbes-et-al_2025.qmd`: Primary analyses and visualisation<br>3. `3_SnapshotsForecasting_Forbes-et-al_2025.qmd`: Historical snapshot analysis and forecasting</p> <p>## Requirements</p> <p>- R (version 4.3.2 or later)<br>- Required R packages:<br> - tidyverse (v2.0.0) - for data manipulation and visualization<br> - readr (v2.1.5) - for reading CSV/TSV files<br> - ggplot2 (v3.4.0 or v3.5.0) - for creating visualizations<br> - rnaturalearth (v1.0.1) - for accessing natural earth map data<br> - dplyr (v1.1.0 or v1.1.4) - for data manipulation<br> - countrycode (v1.6.0) - for converting country names and codes<br> - spdep (v1.3-3) - for spatial dependence modeling<br> - sp (v1.6-0 or v2.1-3) - for spatial data manipulation<br> - sf (v1.0-15 or v1.0-16) - for simple features access<br> - data.table (v1.14.8) - for fast aggregation of large data<br> - lubridate (v1.9.2) - for date-time manipulation<br> - viridis (v0.6.3) - for color palettes<br> - gridExtra (v2.3) - for arranging multiple plots<br> - ggpubr (v0.6.0) - for creating publication-ready plots<br> - zoo (v1.8-12) - for time series, including moving averages<br> - scales (v1.3.0) - for graphical scales<br> - forecast (v8.22.0) - for ARIMA forecast models<br> - purrr (v1.0.2) - for mapping custom forecast function onto each dataset<br> - arrow - for working with parquet files</p> <p>Install these packages before running the scripts.</p> <p>## How to Use</p> <p>1. Download this repository to your local machine.<br>2. Set your working directory to the location of the scripts.<br>3. Download raw datasets from GBIF (as required)<br>4. Ensure all required R packages are installed.<br>5. Run the scripts in RStudio or your preferred R environment.</p> <p>### Data Cleaning (`1_DataCleaning_Forbes-et-al_2025.qmd`)</p> <p>This script cleans the raw GBIF data and identifies anomalies. It produces files containing indexes of dataset records to be removed, which are used in subsequent analyses.</p> <p>**Note**: The raw GBIF exported datasets for contemporary records are not included in this repository due to file size constraints. Download them from the GBIF links provided in the script and place them in the `data/` directory.</p> <p>### Primary Analyses (`2_PrimaryAnalyses_Forbes-et-al_2025.qmd`)</p> <p>This script performs the main analyses and generates visualisations. It uses the outputs from the data cleaning script to filter anomalous records.</p> <p>To reproduce all analysis stages from the original raw .csv files:<br>- Start at the chunks labelled "DATA LOAD AND FILTERING".<br>- Run the pipeline for non-spatial analyses before spatial analyses.<br>- Due to memory constraints, it's recommended to run analyses for one taxonomic group and one analysis stream at a time.</p> <p>To skip to plot generation:<br>- Navigate to sections tagged as "@! SKIP TO PLOTTING !@".<br>- Ensure all required analysis output files are in the `data/` directory.</p> <p>### Forecasting (`3_SnapshotsForecasting_Forbes-et-al_2025.qmd`)</p> <p>This script analyses historical GBIF database snapshots and forecasts future growth. It uses the cleaned snapshot data produced by the data cleaning script.</p> <p>## Data Files</p> <p>### GBIF Exports - Raw Data (not included on Zenodo due to file size, please download directly from GBIF)<br>- `0016915-240425142415019.csv` for Chordata - https://www.gbif.org/occurrence/download/0016915-240425142415019</p> <p>- `0016914-240425142415019.csv` for Plantae - https://www.gbif.org/occurrence/download/0016914-240425142415019 </p> <p>- `0016913-240425142415019.csv` for Arthropoda - https://www.gbif.org/occurrence/download/0016913-240425142415019</p> <p>### Included Data Files</p> <p>#### Raw Data<br>- `GBIF_snapshots.parquet` # Historical snapshots RAW dataset (arrow/parquet format)<br>- `GBIF_integer_to_datasetKey.tsv` # Mapping old dataset IDs onto new datasetKey field</p> <p>#### Contemporary Datasets - data cleaning outputs<br>- `chordata_counts_to_highlight_030724` # List of anomalous Chordata dataset + year indexes to filter<br>- `arthropoda_counts_to_highlight_OG_030724` # List of anomalous Arthropoda dataset + year indexes to filter<br>- `plantae_counts_to_highlight_030724` # List of anomalous Plantae dataset + year indexes to filter</p> <p>#### Cleaned Snapshots<br>- `plantae_snapshots_filter_threshold_IN_040924` # Cleaned Plantae snapshots<br>- `arthropoda_snapshots_filter_threshold_IN_040924` # Cleaned Arthropoda snapshots<br>- `chordata_snapshots_filter_threshold_IN_040924` # Cleaned Chordata snapshots<br>- `gbif_dates_df_anomaly_filtered_090724` # Anomaly-filtered snapshots (combined dataset)<br>- `gbif_dates_df_anomalies_highlighted_090724` # Anomalies highlighted snapshots (combined dataset)</p> <p>#### Analysis Outputs - for skipping straight to plot/figure generation<br>- `arthropoda_specimens_per_year_080724` # Arthropoda specimen counts per year<br>- `arthropoda_unique_species_per_year_080724` # Arthropoda unique species counts per year<br>- `arthropoda_grid_counts_080724` # Arthropoda grid counts<br>- `chordata_specimens_per_year_080724` # Chordata specimen counts per year<br>- `chordata_unique_species_per_year_080724` # Chordata unique species counts per year<br>- `chordata_grid_counts_080724` # Chordata grid counts<br>- `plantae_specimens_per_year_080724` # Plantae specimen counts per year<br>- `plantae_unique_species_per_year_080724` # Plantae unique species counts per year<br>- `plantae_grid_counts_080724` # Plantae grid counts<br>- `chordata_continent_count_080724` # Chordata continent-specific counts<br>- `arthropoda_continent_count_080724` # Arthropoda continent-specific counts<br>- `plantae_continent_count_080724` # Plantae continent-specific counts</p> <p> </p>
Information literacy in the area of Library and Information Science. A bibliometric analysis in Latin America, from the Lens database (2001-2020).
<p>The results of scientific production on ALFIN (2001-2020) in the areas of Library and Information Science are shown. All BIC journals were identified from Latindex. Then it was verified whether these journals were contained in the following databases: Web of Science (Core Collection and Scielo Citation Index), Scopus, Lens and Dimensions. The Lens database was chosen for retrieving records on ALFIN and performing the bibliometric analysis, as it has the highest coverage of BIC journals in Latindex. The trend and growth of scientific production were evaluated according to authors and year of publication; the productivity of authors was analyzed using Lotka's Law and the dispersion of the literature according to Bradford's Law. The degree, index and coefficient of collaboration were determined and collaboration networks were identified according to authors. The results show that scientific production on ALFIN in Latin America, reached a peak between 2017 and 2018, presenting a decrease from 2019 onwards. It was also observed that the production, collaboration between authors and the number of journals is predominantly Brazilian.</p>
FIG. 3 in The d'Orbigny Palaeontological Collection of the National Museum of Natural History and Science, Lisbon, Portugal: Historical perspective and revision of Cretaceous Cephalopoda
FIG. 3. — Cretaceous ammonites of the d'Orbigny Collection of the National Museum of Natural History and Science (Museu Nacional de História Natural e da Ciência): A-D, Neolissoceras grasianum (d'Orbigny, 1840) in ventral (A), lateral (B) and oral (C) views, and original label (D): Nº 357/Ammonites grasanus (d'Orb), Andar 17º Neocomiense, Terreno Cretaceo, Localidade S.t Julien (Hautes Alpes); E-G, Pleurohoplites (Pleurohoplites) renauxianus (d'Orbigny, 1840) in lateral (E) and ventral (F) views, and original label (G): Nº 464/Ammonites Renauxianus (d'Orb), Andar 20º Cenomaniense, Terreno Cretaceo, Localidade Mont-Blainville (Meuse); H-K, Acanthoceras rhotomagense (Brongniart, 1822) in oral (H), lateral (I) and ventral (J) views, and original label (K): Nº 463/Ammonites rhotomagensis (Lamarck), Andar 20º Cenomaniense, Terreno Cretaceo, Localidade Rouen (Seine inf.re). Scale bar: 2 cm.
FIG. 2 in The d'Orbigny Palaeontological Collection of the National Museum of Natural History and Science, Lisbon, Portugal: Historical perspective and revision of Cretaceous Cephalopoda
FIG. 2. — Cretaceous nautiloid and ammonites of the d'Orbigny Collection of the National Museum of Natural History and Science (Museu Nacional de História Natural e da Ciência): A-C, Angulithes triangularis de Montfort, 1808 in oral (A) and lateral (B) views, and original label (C): Nº 459/Nautilus triangularis (Montf), Andar 20º Cenomaniense, Terreno Cretaceo, Localidade Fouras (Charente inf.re); D-G, Phylloceras (Hypophylloceras) tethys (d'Orbigny, 1840) in ventral (D), lateral (E) and oral (F) views, and original label (G): Nº 360/Ammonites Tethys (d'Orb), Andar 17º Neocomiense, Terreno Cretaceo, Localidade Arredores de [environs of] Sisteron (Basses Alpes); H-K, Ptychophylloceras (Semisulcatoceras) semisulcatum (d'Orbigny, 1840) in ventral (H), lateral (I) and oral (J) views, and original label (K): Nº 359/Ammonites semisulcatus (d'Orb), Andar 17º Neocomiense, Terreno Cretaceo, Localidade Sisteron (Basses Alpes). Scale bar: 2 cm.
Citizen science at public libraries: Data on librarians and users perceptions of participating in a citizen science project in Catalunya, Spain
<p>As libraries struggle to keep pace with the changing societal landscape, emerging practices such as citizen science (CS) initiatives are being incorporated to reinforce the idea of public libraries as gathering, meeting, and collaboration spaces within the context of shared community and shared learning resources. However, there is little empirical evidence of whether the most open and participatory ways that CS puts forward can converge with and be nurtured by the essence of public libraries. Also, the roles of librarians and users in the ‘next generation public library’ have been under-developed. As the number of CS initiatives at public libraries grows, so does the need to collect evidence on the impact and the capacity of assimilation of CS practices. The data describes librarians and users' perceptions of participating in a citizen science project. Two hands-on activities for librarians of the Barcelona Network of Public Libraries were implemented. One was a training course for 30 librarians from 24 libraries which allowed them to envisage citizen science implementation in each library. The second activity consisted in the co-creation of a citizen social science project. 40 library users, 7 librarians from 3 different cities, and professional scientists, were involved. The data on librarians and users' perception was collected through participant observation, surveys, and a focus group to identify strengths and challenges of implementing citizen science at public libraries. The data covers librarians and users attitudes towards citizen science, their motivations to participate, their perceived ability to implement a citizen science project (as for librarians) or to contribute to science (as for library users), and the participants intention to keep engaged with citizen science, drawing on the Theory of Planned Behavior. Responses to closed-ended survey questions are analyzed at a descriptive level. The qualitative feedback from the focus group and the open-ended survey question on motivations is subjected to a thematic analysis. The data offers interesting insights to identify opportunities and challenges of implementing citizen science at public libraries, contributing to the debate over the public library's mission as local community hub.</p> <p>The dataset is formed by 5 tables:</p> <ol> <li>Librarians_pre.csv: data on librarians profiles, attitudes towards citizen science, expected impact of the project and self-efficacy collected at the beginning of the Citizen Science Lab.</li> <li>Librarians_post.csv: data on librarians profiles, attitudes towards citizen science, perceived impact of the project and self-efficacy collected at the end of the Citizen Science Lab.</li> <li>Users_first_phase.csv: data on users profiles, motivation, attitudes towards the library, confidence to perform scientific tasks and self-efficacy collected at the beginning of the Science and Citizen Action.</li> <li>Users_second_phase.csv: data on users profiles and motivation collected at the middle of the Science and Citizen Action.</li> <li>Users_last_phase.csv: data on users profiles, attitudes towards the library, confidence to perform scientific tasks and perceived impact of the project collected at the end of the Science and Citizen Action.</li> </ol> <p><strong>Citizen Science Lab Questionnaire (Librarians_pre)</strong></p> <table> <tbody> <tr> <td> <p><strong>Personal information</strong></p> </td> </tr> <tr> <td> <p>1. [rol_1] What is your role at the library?</p> </td> <td> <ul> <li>Director</li> <li>Library technician</li> <li>Support technician</li> <li>Service support</li> </ul> </td> </tr> <tr> <td> <p>2. [years_1] How long have you been working at the library?</p> </td> <td> <ul> <li>2 or less</li> <li>3 to 5 years</li> <li>6 to 10 years</li> <li>11 to 20 years</li> <li>more than 20 years</li> </ul> </td> </tr> <tr> <td> <p>3. [back_1] Do you have a scientific background?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>4. [know_1] Have you already heard about citizen science?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>5. [part_1] Have you already participated in a citizen science project?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p><strong>Attitudes towards users engagement</strong></p> </td> </tr> <tr> <td> <p>6. [att_lib_pre1] Do you believe that library users are able to participate in a citizen science project?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> <tr> <td> <p>7. [att_lib_pre2] Do you believe that library users will commit to participating in a citizen science project?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Expected impacts</strong></p> </td> </tr> <tr> <td> <p>8. [exp_lib_pre] To what extent do you believe that citizen science may bring positive impacts to your library?</p> <p> </p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Self-efficacy</strong></p> </td> </tr> <tr> <td> <p>9. [se_lib_pre1] Right now, do you feel able to recommend any citizen science project to library users?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>10. [se_lib_pre2] Right now, do you feel able to implement yourself and lead a citizen science project?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> </tbody> </table> <p><strong>Citizen Science Lab Questionnaire (Librarians_post)</strong></p> <table> <tbody> <tr> <td> <p><strong>Personal information</strong></p> </td> </tr> <tr> <td> <p>1. [years_2] How long have you been working at the library?</p> </td> <td> <ul> <li>2 or less</li> <li>3 to 5 years</li> <li>6 to 10 years</li> <li>11 to 20 years</li> <li>more than 20 years</li> </ul> </td> </tr> <tr> <td> <p>2. [back_2] Do you have a scientific background?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>3. [sat_1] To what extent does the project meet your initial expectations?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Attitudes towards users engagement</strong></p> </td> </tr> <tr> <td> <p>4. [att_lib_post1] Do you believe that library users will commit to participating in a citizen science project?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> <tr> <td> <p>5. [att_lib_post2] What are/could be the potential barriers to users engagement in citizen science?</p> </td> <td> <p>[open]</p> </td> </tr> <tr> <td> <p><strong>Perceived impact</strong></p> </td> </tr> <tr> <td> <p>6. [imp_lib] What do you believe that citizen science may bring to public libraries and users?</p> <p>1 [Not at all] …… 5 [Totally]</p> <p> </p> </td> <td> <p>a. Knowledge of the scientific process</p> <p>b. New connections among participants </p> <p>c. Fun</p> <p>d. New knowledge of the local environment</p> <p>e. Scientific evidence on a common concern</p> <p>f. Social cohesion</p> <p>g. Positive attitudes towards science</p> <p>h. Willingness to learn</p> <p>i. Critical thinking and self-efficacy</p> </td> </tr> <tr> <td> <p><strong>Self-efficacy</strong></p> </td> </tr> <tr> <td> <p>8. [se_lib_post1] Right now, do you feel able to recommend any citizen science project to library users?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>9. [se_lib_post2] Right now, do you feel able to implement yourself and lead a citizen science project?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Intention to keep engaged</strong></p> </td> </tr> <tr> <td> <p>10. [eng_lib] To what extent are you motivated to keep engaged with citizen science?</p> <p> </p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> </tbody> </table> <p> </p> <p><strong>Science and Citizens Action Focus group guide (Librarians)</strong></p> <p><strong>Opening questions</strong></p> <p><strong>1.</strong> To start with…. Are you satisfied with the project?</p> <p><strong>Probe</strong><strong>:</strong> Yes, no, why? Was it fun/interesting/challenging/enriching….</p> <p><strong>2.</strong> Do you feel you have learned something new?</p> <p><strong>Probe</strong><strong>:</strong> About your library’s environment, users, science and citizen science... Is there anything special that you will take with you after the project?</p> <p><strong>Reflections on the cocreation process</strong></p> <p><strong>3.</strong> At what time during the project have you felt most comfortable?</p> <p><strong>Probe:</strong> For example, has it been easier to lead the activity and/or involve and retain the community? Did you find it entertaining?</p> <p><strong>4.</strong> At what time during the project have you felt less at ease?</p> <p><strong>Probe</strong><strong>:</strong> What was challenging during the cocreation process?</p> <p><strong>5.</strong> To what extent do you feel more capable of implementing and leading a citizen science project in your library right now?</p> <p><strong>Probe: </strong>For example, in the case of both more crowdsourcing and of cocreated projects that actively involve the community</p> <p><strong>Reflections on the perceived impact</strong></p> <p><strong>6. </strong>To what extent does the project meet your initial expectations?</p> <p><strong>Probe</strong><strong>:</strong> in line with what you discussed at the beginning of the project, you expected it to promote participation, new connections among participants, improve the library perceptions and stimulate the participants’ critical thinking...Do you think that citizen science may meet these expectations?</p> <p><strong>Reflections on citizen science at public libraries</strong></p> <p><strong>7.</strong> To what extent can citizen science (in its most ‘extreme’ form of participation) be imagined as an activity within the library that promotes more active user participation?</p> <p><strong>Probe</strong><strong>:</strong> Through for example cocreation, experimentation, and hands-on learning activities...</p> <p><strong>8. </strong>Do you think that the activity has brought new knowledge? What new knowledge has the activity brought from your perspective?</p> <p><strong>Probe</strong><strong>:</strong> Knowledge of the scientific process, knowledge of the community or new users...</p> <p>9. What could be the opportunities and barriers of introducing citizen science at public libraries? And the barriers?</p> <p><strong>Probe:</strong> Like for example improving the perception of the library, actively involving certain users...What could be the ‘return’ for the community? What impact can citizen science projects have on making the environment more dynamic from libraries?</p> <p><strong>10. </strong>More generally, what could be the ‘added value’ of the introduction of citizen science within the library’s range of activities?</p> <p><strong>Probe:</strong> Is it a fun activity that promotes socialization, for example? Or that allows to generate new knowledge? Or that highlights the library’s social value? Or, also, that may offer new uses and new roles to the library? Can it foster a sense of community with the library as a connector? What other impacts can be generated in your environment?</p> <p><strong>Closing</strong></p> <p><strong>11.</strong> Do you see yourselves the next year, implementing a citizen science project as part of the library’s range of activities? And adopting an existing one?</p> <p><strong>Probe: </strong>Are you motivated to get more involved with citizen science projects? What kind of projects? What level of user involvement do you expect? What barriers do you see to users’ involvement? What benefits and opportunities do you think you can bring to the library?</p> <p><strong>12.</strong> We have now reached the end of the discussion. Anyone want to add anything else?</p> <p><strong>Science and Citizens Action Questionnaire (Users_first_phase)</strong></p> <table> <tbody> <tr> <td> <p><strong>Personal information</strong></p> </td> </tr> <tr> <td> <p>1. [gen_1] Are you..?</p> </td> <td> <ul> <li>Woman</li> <li>Man</li> <li>NA</li> </ul> </td> </tr> <tr> <td> <p>2. [years_3] How old are you?</p> </td> <td> <ul> <li>18-25</li> <li>26-35</li> <li>36-45</li> <li>46-55</li> <li>56-65</li> <li>66+</li> </ul> </td> </tr> <tr> <td> <p>3. [rol_2] What is your role at the library?</p> </td> <td> <ul> <li>Library user not associated with local associations</li> <li>Library technician</li> <li>Member of a local association</li> <li>Representative of public administrations</li> <li>Representative of the private sector</li> <li>Others:</li> </ul> </td> </tr> <tr> <td> <p>4. [back_3] Do you have a scientific background?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>5. [part_2] Have you already participated in a citizen science project?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p><strong>Motivations to participate</strong></p> </td> </tr> <tr> <td> <p>6. [mot_us] What did motivate you to participate in the project?</p> </td> <td> <p>[open]</p> </td> </tr> <tr> <td> <p><strong>Attitudes towards the library</strong></p> </td> </tr> <tr> <td> <p>7. [att_us_pre1] To what extent do you believe that your library is responsive to the community needs?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> <tr> <td> <p>8. [att_us_pre2] To what extent to you believe your library is able to face local challenges based on users' active participation?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Confidence to perform scientific tasks</strong></p> </td> </tr> <tr> <td> <p>9. [conf_us_pre] To what extent do you feel able to contribute to perform the following scientific tasks:</p> <p>1 [Not at all] …… 4 [Totally]</p> </td> <td> <p>a. Formulate the research question</p> <p>b. Data collection</p> <p>c. Analysis and interpretation of the results</p> <p>d. Propose concrete actions based on scientific evidence</p> </td> </tr> <tr> <td> <p><strong>Self-efficacy</strong></p> </td> </tr> <tr> <td> <p>10. [se_us_pre] To what extent do you feel able to positively contribute to the library and your community?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4 [Totally]</li> </ul> </td> </tr> </tbody> </table> <p><strong>Science and Citizens Action Questionnaire (Users_second_phase)</strong></p> <table> <tbody> <tr> <td> <p><strong>Personal information</strong></p> </td> </tr> <tr> <td> <p>1. [gen_3] Are you..?</p> </td> <td> <ul> <li>Woman</li> <li>Man</li> <li>NA</li> </ul> </td> </tr> <tr> <td> <p>2. [years_5] How old are you?</p> </td> <td> <ul> <li>18-25</li> <li>26-35</li> <li>36-45</li> <li>46-55</li> <li>56-65</li> <li>66+</li> </ul> </td> </tr> <tr> <td> <p>3. [rol_4] What is your role at the library?</p> </td> <td> <ul> <li>Library user or technician not associated with local associations</li> <li>Member of a local association</li> <li>Representative of public administrations</li> <li>Representative of the private sector</li> <li>Others:</li> </ul> </td> </tr> <tr> <td> <p>3. [back_5] Do you have a scientific background?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>4. [mot_us2] To what extent are you motivated to carry out the experiment?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> </tbody> </table> <p><strong>Science and Citizens Action Questionnaire (Users_last_phase)</strong></p> <table> <tbody> <tr> <td> <p><strong>Personal information</strong></p> </td> </tr> <tr> <td> <p>1. [gen_2] Are you..?</p> </td> <td> <ul> <li>Woman</li> <li>Man</li> <li>NA</li> </ul> </td> </tr> <tr> <td> <p>2. [years_4] How old are you?</p> </td> <td> <ul> <li>18-25</li> <li>26-35</li> <li>36-45</li> <li>46-55</li> <li>56-65</li> <li>66+</li> </ul> </td> </tr> <tr> <td> <p>3. [rol_3] What is your role at the library?</p> </td> <td> <ul> <li>Library user or technician not associated with local associations</li> <li>Member of a local association</li> <li>Representative of public administrations</li> <li>Representative of the private sector</li> <li>Others:</li> </ul> </td> </tr> <tr> <td> <p>3. [back_4] Do you have a scientific background?</p> </td> <td> <ul> <li>Yes</li> <li>No</li> </ul> </td> </tr> <tr> <td> <p>4. [part_3] To how many cocreation sessions have you participated?</p> </td> <td> <ul> <li>None</li> <li>1</li> <li>2</li> <li>3</li> </ul> </td> </tr> <tr> <td> <p>5. [sat_2] To what extent are you satisfied with the experiment?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Attitudes towards the library</strong></p> </td> </tr> <tr> <td> <p>6. [att_us_post] To what extent do you believe that the project has positively changed your perception of the library?</p> </td> <td> <ul> <li>1 [Not at all]</li> <li>2</li> <li>3</li> <li>4</li> <li>5 [Totally]</li> </ul> </td> </tr> <tr> <td> <p><strong>Confidence to perform scientific tasks</strong></p> </td> </tr> <tr> <td> <p>7. [conf_us_post] To what extent do you feel able to contribute to perform the following scientific tasks:</p> <p>1 [Not at all] …… 5 [Totally]</p> </td> <td> <p>a. Formulate the research question</p> <p>b. Data collection</p> <p>c. Analysis and interpretation of the results</p> <p>d. Propose concrete actions based on scientific evidence</p> </td> </tr> <tr> <td> <p><strong>Perceived impact</strong></p> </td> </tr> <tr> <td> <p>8. [imp_us] What do you believe that citizen science may bring to public libraries and users?</p> <p>1 [Not at all] …… 5 [Totally]</p> <p> </p> </td> <td> <p>a. Knowledge of the scientific process</p> <p>b. New connections among participants </p> <p>c. Fun</p> <p>d. New knowledge of the local environment</p> <p>e. Scientific evidence on a common concern</p> <p>f. Social cohesion</p> <p>g. Positive attitudes towards science</p> <p>h. Willingness to learn</p> <p>i. Critical thinking and self-efficacy</p> </td> </tr> </tbody> </table> <p> </p>
The new normal? Redaction bias in biomedical science
<p>A concerning amount of biomedical research is not reproducible. Unreliable results impede empirical progress in medical science, ultimately putting patients at risk. Many proximal causes of this irreproducibility have been identified, a major one being inappropriate statistical methods and analytical choices by investigators. Within this, we formally quantify the impact inappropriate redaction beyond a threshold value in biomedical science. This is effectively truncation of a data-set by removing extreme data points, and we elucidate its potential to accidentally or deliberately engineer a spurious result in significance testing. We demonstrate that the removal of a surprisingly small number of data points can be used to dramatically alter a result. It is unknown how often redaction bias occurs in the broader literature, but given the risk of distortion to the literature involved, we suggest that it must be studiously avoided, and mitigated with approaches to counteract any potential malign effects to the research quality of medical science.</p>
Open Science: Epistemologies, science policy implementation and digitalisation
<p>This workshop explores Open Science and the promises and challenges for science and society. After briefly covering the definition, history and application of Open Science policies, we discuss why they are relevant and what barriers and concerns there are for implementation. The workshop engages with the challenges of universalising access to knowledge in a world with a plurality of epistemic communities and developmental agendas. Lastly, we discuss the role that digitalisation plays in implementing Open Science policies and what this means for sustainable development.</p>
Using convolutional neural networks to efficiently extract immense phenological data from community science images
<p>Community science image libraries offer a massive, but largely untapped, source of observational data for phenological research. The iNaturalist platform offers a particularly rich archive, containing more than 49 million verifiable, georeferenced, open access images, encompassing seven continents and over 278,000 species. A critical limitation preventing scientists from taking full advantage of this rich data source is labor. Each image must be manually inspected and categorized by phenophase, which is both time-intensive and costly. Consequently, researchers may only be able to use a subset of the total number of images available in the database. While iNaturalist has the potential to yield enough data for high-resolution and spatially extensive studies, it requires more efficient tools for phenological data extraction. A promising solution is automation of the image annotation process using deep learning. Recent innovations in deep learning have made these open-source tools accessible to a general research audience. However, it is unknown whether deep learning tools can accurately and efficiently annotate phenophases in community science images. Here, we train a convolutional neural network (CNN) to annotate images of Alliaria petiolata into distinct phenophases from iNaturalist and compare the performance of the model with non-expert human annotators. We demonstrate that researchers can successfully employ deep learning techniques to extract phenological information from community science images. A CNN classified two-stage phenology (flowering and non-flowering) with 95.9% accuracy and classified four-stage phenology (vegetative, budding, flowering, and fruiting) with 86.4% accuracy. The overall accuracy of the CNN did not differ from humans (p = 0.383), although performance varied across phenophases. We found that a primary challenge of using deep learning for image annotation was not related to the model itself, but instead in the quality of the community science images. Up to 4% of A. petiolata images in iNaturalist were taken from an improper distance, were physically manipulated, or were digitally altered, which limited both human and machine annotators in accurately classifying phenology. Thus, we provide a list of photography guidelines that could be included in community science platforms to inform community scientists in the best practices for creating images that facilitate phenological analysis.</p>
Question Bank for Open Syllabus UNESCO Recommendation on Open Science
<p>Question bank of multiple choice, multiple response, and essay format questions to accompany the Open Syllabus: UNESCO Recommendation on Open Science. These questions are intended to be used for low-stakes comprehension checks and discussion. The file format is compatible with the Respondus tool for importing questions into learning management systems.</p>
"Python for Data Science" (AY250; UC Berkeley) Data files
<p>Data files for "Python for Data Science" (AY250; UC Berkeley)</p> <ul> <li><a href="https://zenodo.org/api/files/796932a4-1a76-4467-a66e-ac0d47e029c7/homework1_data.tgz">homework1_data.tgz </a>- Data for HW1</li> </ul> <p>Course website: https://github.com/profjsb/python-seminar</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.