Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
210
datasets available to search
ShareScore release 0.7.1
Dataset results
210 results for “natural science”
Frictionless Tabular Data Package for GC-MS Rose scent profile data for Data published in Nature genetics, June, 2018 & Science, July 2015
<p>This dataset, in the form of a Frictionless Tabular Data Package (https://frictionlessdata.io/specs/tabular-data-package/), holds the measurements of 35 known metabolites(all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in one Rose cultivars (all annotated with resolvable NCBITaxonomy Identifiers) and one organism part (annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable STATO terms. The measurements over these metabolites, which were made in 2 distinct experiments, were extracted from: a supplementary material table, available from https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip and published alongside the Nature Genetics manuscript identified by the following doi: https://doi.org/10.1038/s41588-018-0110-3, published in June 2018 a supplementary material table available as a pdf from 'Biosynthesis of monoterpene scent compounds in roses' by Magnard et al, Science 03 Jul 2015 identified by the following doi: https://doi.org/10.1126/science.aab0696. This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR)and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.It is associated to the following project: https://github.com/proccaserra/rose2018ng-notebook with all the necessaryinformation, executable code and tutorials in the form of Jupyter notebooks.</p>
Data on a citation context analysis focusing on natural sciences and social sciences and humanities
<p>This dataset contains data on citation context analysis between natural sciences (NS) and social sciences and humanities (SSH). In particular, the data were created through manual coding of each citation between papers related to SDG7 (renewable energy) and SDG13 (climate change) and papers cited by them. This dataset consists of 9 files, associated with the article: Nishikawa, K. How and why are citations between disciplines made? A citation context analysis focusing on natural sciences and social sciences and humanities. Scientometrics (2023). <a href="https://doi.org/10.1007/s11192-023-04664-y">https://doi.org/10.1007/s11192-023-04664-y</a></p> <p> </p> <p>The files are numbered as follows:</p> <ul> <li>00 – README</li> <li>01 – Data by citation pair for SDG7 (original)</li> <li>02 – Data by citation pair for SDG13 (original)</li> <li>03 – Data by mention location for SDG7 (original)</li> <li>04 – Data by mention location for SDG13 (original)</li> <li>05 – Data by citation pair for SDG7 (additional)</li> <li>06 – Data by citation pair for SDG13 (additional)</li> <li>07 – Data by mention location for SDG7 (additional)</li> <li>08 – Data by mention location for SDG13 (additional)</li> </ul> <p>See README for more information.</p>
Annual Article Processing Charges (APCs) and number of gold and hybrid open access articles in Web of Science indexed journals published by Elsevier, Sage, Springer-Nature, Taylor & Francis and Wiley 2015-2018
<p><strong>Dataset of annual Article Processing Charges (APCs) for 6,252 journals from 2015 to 2018. </strong>The dataset contains annual APCs for journals indexed in the Web of Science (WoS) and published by the oligopoly of academic publishers (Elsevier, Sage, Springer-Nature, Taylor & Francis, Wiley). It also includes an estimate of the total APCs paid by the academic community based on the number of gold and hybrid articles published between 2015 and 2018. The dataset was created using publication data from WoS, OA status from Unpaywall and annual APC prices from open datasets (<a href="https://doi.org/10.5281/ZENODO.3841568">Matthias, 2020</a>; <a href="https://doi.org/10.5683/SP2/84PNSG">Morrison, 2021</a>) and historical fees retrieved via the Internet Archive Wayback Machine. </p> <p>Detailed methods and findings are reported in the following journal article</p> <p>Butler, L.-A., Matthias, L., Simard, M.-A., Mongeon, P., & Haustein, S. (2023). The Oligopoly's Shift to Open Access. How the Big Five Academic Publishers Profit from Article Processing Charges. <em>Quantitative Science Studies</em>. Preprint: <a href="https://doi.org/10.5281/zenodo.8322555">https://doi.org/10.5281/zenodo.8322555</a></p> <p><strong>Description of included files (v1):</strong></p> <p><em>APCs.csv: </em>contains the annual APCs for gold and hybrid OA journals indexed in Web of Science published by the oligopoly of academic publishers (Elsevier, Sage, Springer-Nature, Taylor & Francis, Wiley) between 2015 and 2018 including the total estimate of APCs paid per journal per year. It contains APC data for 18,846 journal-year-OA status combinations.</p> <p><em>countries.csv</em>: contains the fractionalized number of annual gold and hybrid OA articles by oligopoly publishers between 2015 and 2018 and the total estimate of fractionalized APCs paid per country per journal per year.</p> <p><em>oecd.csv</em>: contains the fractionalized number of annual gold and hybrid OA articles by oligopoly publishers between 2015 and 2018 and the total estimate of fractionalized APCs per discipline per journal per year.</p> <p><em>ReadMe.csv</em>: contains a description of the variables used in <em>APCs.csv</em>, <em>countries.csv</em> and <em>oecd.csv</em>.</p> <p> </p>
Frictionless Tabular Data Package for GC-MS Rose scent profile data for Data published in Nature genetics, June, 2018 & Science, July 2015
<p>This dataset, in the form of a Frictionless Tabular Data Package (<a href="https://frictionlessdata.io/specs/tabular-data-package/">https://frictionlessdata.io/specs/tabular-data-package/)</a>, holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxonomy Identifiers) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable <a href="https://github.com/ISA-tools/stato">STATO</a> terms. </p> <p>The data were extracted from:</p> <ul> <li>a supplementary material table, available from <a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip">https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip</a> and published alongside the Nature Genetics manuscript identified by the following doi: <a href="https://doi.org/10.1038/s41588-018-0110-3">https://doi.org/10.1038/s41588-018-0110-3</a>, published in June 2018</li> <li>a supplementary material table available as a pdf from "Biosynthesis of monoterpene scent compounds in roses" by Magnard et al, Science 03 Jul 2015 identified by the following doi: <a href="https://doi.org/10.1126/science.aab0696">https://doi.org/10.1126/science.aab0696</a></li> </ul> <p>This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR) and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.</p> <p>It is associated to the following project: <a href="https://github.com/proccaserra/rose2018ng-notebook">https://github.com/proccaserra/rose2018ng-notebook</a> with all the necessary information, executable code and tutorials in the form of Jupyter notebooks.</p> <p> </p>
Data/ codes used in the the Natural Hazards and Earth System Sciences (NHESS) publication titled "Wind-Wave Characteristics and extremes along the Emilia-Romagna coast" by Pranavam Ayyappan Pillai et al. (2022)
<p>The archive contains datasets and codes used in the manuscript titled "Wind-Wave Characteristics and extremes along the Emilia-Romagna coast", and published in the journal <em>Natural Hazards and Earth System Sciences</em> (<em>NHESS</em>) by Pranavam Ayyappan Pillai et al., 2022.</p> <p>Pranavam Ayyappan Pillai, U., Pinardi, N., Federico, I., Causio, S., Trotta, F., Unguendoli, S., and Valentini, A.: Wind-Wave Characteristics and extremes along the Emilia-Romagna coast, Nat. Hazards Earth Syst. Sci. Discuss. https://doi.org/10.5194/nhess-2022-103, 2022.</p>
Data and code for: "Global Sampling Decline Erodes Science Potential of Natural History Collections"
<p># GBIF Specimen Data Analysis and Forecasting<br><br>## Version 2 - modified date ranges for figures 1 and 2 in response to reviewer comments</p> <p>This repository contains the code and data for analysing and forecasting trends in Global Biodiversity Information Facility (GBIF) specimen records across three major taxonomic groups: Chordata, Arthropoda, and Plantae. <br>The analysis pipeline includes data cleaning, anomaly detection, primary analyses, and forecasting based on historical database snapshots.</p> <p>These scripts and data correspond to analyses in the following manuscript:</p> <p>Global Sampling Decline Erodes Science Potential of Natural History Collections</p> <p>Authors:<br>Owen Forbes<br>Andrew G. Young<br>Peter H. Thrall</p> <p><br>## Repository Structure</p> <p>The repository consists of three main Quarto (.qmd) scripts and associated data files:</p> <p>1. `1_DataCleaning_Forbes-et-al_2025.qmd`: Data cleaning and anomaly detection<br>2. `2_PrimaryAnalyses_Forbes-et-al_2025.qmd`: Primary analyses and visualisation<br>3. `3_SnapshotsForecasting_Forbes-et-al_2025.qmd`: Historical snapshot analysis and forecasting</p> <p>## Requirements</p> <p>- R (version 4.3.2 or later)<br>- Required R packages:<br> - tidyverse (v2.0.0) - for data manipulation and visualization<br> - readr (v2.1.5) - for reading CSV/TSV files<br> - ggplot2 (v3.4.0 or v3.5.0) - for creating visualizations<br> - rnaturalearth (v1.0.1) - for accessing natural earth map data<br> - dplyr (v1.1.0 or v1.1.4) - for data manipulation<br> - countrycode (v1.6.0) - for converting country names and codes<br> - spdep (v1.3-3) - for spatial dependence modeling<br> - sp (v1.6-0 or v2.1-3) - for spatial data manipulation<br> - sf (v1.0-15 or v1.0-16) - for simple features access<br> - data.table (v1.14.8) - for fast aggregation of large data<br> - lubridate (v1.9.2) - for date-time manipulation<br> - viridis (v0.6.3) - for color palettes<br> - gridExtra (v2.3) - for arranging multiple plots<br> - ggpubr (v0.6.0) - for creating publication-ready plots<br> - zoo (v1.8-12) - for time series, including moving averages<br> - scales (v1.3.0) - for graphical scales<br> - forecast (v8.22.0) - for ARIMA forecast models<br> - purrr (v1.0.2) - for mapping custom forecast function onto each dataset<br> - arrow - for working with parquet files</p> <p>Install these packages before running the scripts.</p> <p>## How to Use</p> <p>1. Download this repository to your local machine.<br>2. Set your working directory to the location of the scripts.<br>3. Download raw datasets from GBIF (as required)<br>4. Ensure all required R packages are installed.<br>5. Run the scripts in RStudio or your preferred R environment.</p> <p>### Data Cleaning (`1_DataCleaning_Forbes-et-al_2025.qmd`)</p> <p>This script cleans the raw GBIF data and identifies anomalies. It produces files containing indexes of dataset records to be removed, which are used in subsequent analyses.</p> <p>**Note**: The raw GBIF exported datasets for contemporary records are not included in this repository due to file size constraints. Download them from the GBIF links provided in the script and place them in the `data/` directory.</p> <p>### Primary Analyses (`2_PrimaryAnalyses_Forbes-et-al_2025.qmd`)</p> <p>This script performs the main analyses and generates visualisations. It uses the outputs from the data cleaning script to filter anomalous records.</p> <p>To reproduce all analysis stages from the original raw .csv files:<br>- Start at the chunks labelled "DATA LOAD AND FILTERING".<br>- Run the pipeline for non-spatial analyses before spatial analyses.<br>- Due to memory constraints, it's recommended to run analyses for one taxonomic group and one analysis stream at a time.</p> <p>To skip to plot generation:<br>- Navigate to sections tagged as "@! SKIP TO PLOTTING !@".<br>- Ensure all required analysis output files are in the `data/` directory.</p> <p>### Forecasting (`3_SnapshotsForecasting_Forbes-et-al_2025.qmd`)</p> <p>This script analyses historical GBIF database snapshots and forecasts future growth. It uses the cleaned snapshot data produced by the data cleaning script.</p> <p>## Data Files</p> <p>### GBIF Exports - Raw Data (not included on Zenodo due to file size, please download directly from GBIF)<br>- `0016915-240425142415019.csv` for Chordata - https://www.gbif.org/occurrence/download/0016915-240425142415019</p> <p>- `0016914-240425142415019.csv` for Plantae - https://www.gbif.org/occurrence/download/0016914-240425142415019 </p> <p>- `0016913-240425142415019.csv` for Arthropoda - https://www.gbif.org/occurrence/download/0016913-240425142415019</p> <p>### Included Data Files</p> <p>#### Raw Data<br>- `GBIF_snapshots.parquet` # Historical snapshots RAW dataset (arrow/parquet format)<br>- `GBIF_integer_to_datasetKey.tsv` # Mapping old dataset IDs onto new datasetKey field</p> <p>#### Contemporary Datasets - data cleaning outputs<br>- `chordata_counts_to_highlight_030724` # List of anomalous Chordata dataset + year indexes to filter<br>- `arthropoda_counts_to_highlight_OG_030724` # List of anomalous Arthropoda dataset + year indexes to filter<br>- `plantae_counts_to_highlight_030724` # List of anomalous Plantae dataset + year indexes to filter</p> <p>#### Cleaned Snapshots<br>- `plantae_snapshots_filter_threshold_IN_040924` # Cleaned Plantae snapshots<br>- `arthropoda_snapshots_filter_threshold_IN_040924` # Cleaned Arthropoda snapshots<br>- `chordata_snapshots_filter_threshold_IN_040924` # Cleaned Chordata snapshots<br>- `gbif_dates_df_anomaly_filtered_090724` # Anomaly-filtered snapshots (combined dataset)<br>- `gbif_dates_df_anomalies_highlighted_090724` # Anomalies highlighted snapshots (combined dataset)</p> <p>#### Analysis Outputs - for skipping straight to plot/figure generation<br>- `arthropoda_specimens_per_year_080724` # Arthropoda specimen counts per year<br>- `arthropoda_unique_species_per_year_080724` # Arthropoda unique species counts per year<br>- `arthropoda_grid_counts_080724` # Arthropoda grid counts<br>- `chordata_specimens_per_year_080724` # Chordata specimen counts per year<br>- `chordata_unique_species_per_year_080724` # Chordata unique species counts per year<br>- `chordata_grid_counts_080724` # Chordata grid counts<br>- `plantae_specimens_per_year_080724` # Plantae specimen counts per year<br>- `plantae_unique_species_per_year_080724` # Plantae unique species counts per year<br>- `plantae_grid_counts_080724` # Plantae grid counts<br>- `chordata_continent_count_080724` # Chordata continent-specific counts<br>- `arthropoda_continent_count_080724` # Arthropoda continent-specific counts<br>- `plantae_continent_count_080724` # Plantae continent-specific counts</p> <p> </p>
FIG. 3 in The d'Orbigny Palaeontological Collection of the National Museum of Natural History and Science, Lisbon, Portugal: Historical perspective and revision of Cretaceous Cephalopoda
FIG. 3. — Cretaceous ammonites of the d'Orbigny Collection of the National Museum of Natural History and Science (Museu Nacional de História Natural e da Ciência): A-D, Neolissoceras grasianum (d'Orbigny, 1840) in ventral (A), lateral (B) and oral (C) views, and original label (D): Nº 357/Ammonites grasanus (d'Orb), Andar 17º Neocomiense, Terreno Cretaceo, Localidade S.t Julien (Hautes Alpes); E-G, Pleurohoplites (Pleurohoplites) renauxianus (d'Orbigny, 1840) in lateral (E) and ventral (F) views, and original label (G): Nº 464/Ammonites Renauxianus (d'Orb), Andar 20º Cenomaniense, Terreno Cretaceo, Localidade Mont-Blainville (Meuse); H-K, Acanthoceras rhotomagense (Brongniart, 1822) in oral (H), lateral (I) and ventral (J) views, and original label (K): Nº 463/Ammonites rhotomagensis (Lamarck), Andar 20º Cenomaniense, Terreno Cretaceo, Localidade Rouen (Seine inf.re). Scale bar: 2 cm.
FIG. 2 in The d'Orbigny Palaeontological Collection of the National Museum of Natural History and Science, Lisbon, Portugal: Historical perspective and revision of Cretaceous Cephalopoda
FIG. 2. — Cretaceous nautiloid and ammonites of the d'Orbigny Collection of the National Museum of Natural History and Science (Museu Nacional de História Natural e da Ciência): A-C, Angulithes triangularis de Montfort, 1808 in oral (A) and lateral (B) views, and original label (C): Nº 459/Nautilus triangularis (Montf), Andar 20º Cenomaniense, Terreno Cretaceo, Localidade Fouras (Charente inf.re); D-G, Phylloceras (Hypophylloceras) tethys (d'Orbigny, 1840) in ventral (D), lateral (E) and oral (F) views, and original label (G): Nº 360/Ammonites Tethys (d'Orb), Andar 17º Neocomiense, Terreno Cretaceo, Localidade Arredores de [environs of] Sisteron (Basses Alpes); H-K, Ptychophylloceras (Semisulcatoceras) semisulcatum (d'Orbigny, 1840) in ventral (H), lateral (I) and oral (J) views, and original label (K): Nº 359/Ammonites semisulcatus (d'Orb), Andar 17º Neocomiense, Terreno Cretaceo, Localidade Sisteron (Basses Alpes). Scale bar: 2 cm.
Fig. 19 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 19.ɹCopidognathus rombus and Halacarellus longus [nomen nudum], glass slide prepared by Dr. Makarova.
Fig. 26 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 26.ɹCopidognathus kamchaticus [nomen nudum], ˂ NSMT-Ac 14711. Body (A), genitoanal region (B) (Phase-contrast micrograph), first leg (C). Scale bars for A=100 µm, BC=50 µm.
Fig. 21 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 21.ɹCopidognathus rombus, paratype ˂ NSMT-Ac 14704. Body (A), gnathosoma (B) (Phase-contrast micrographs). Scale bars for A=100 µm, B=50 µm.
Fig. 16 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 16.ɹCopidognathus globulosus, holotype ˁ NSMT-Ac 14701. Idiosoma (A), genitoanal region (B), gnathosoma (C), first and second legs (D). Scale bars for A=100 µm, BCD=50 µm.
Fig. 18 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 18.ɹCopidognathus pacificus, holotype ˁ NSMT-Ac 14702. Idiosoma (A), genitoanal region (B), first leg (C) (Phase-contrast micrographs). Scale bars for ABC=50 µm.
Fig. 11 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 11.ɹHalacarellus longus [nomen nudum], ˂ NSMT-Ac 14706 (No. 1) mounted on the glass slide of Copidognathus rombus. Body (A), gnathosoma (B) (Phase-contrast micrographs). Scale bars for A=100 µm, B=50 µm.
Fig. 13 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 13.ɹHalacarellus longus [nomen nudum], ˂ NSMT-Ac 14708 (No. 6) mounted on the glass slide of Copidognathus rombus. Body (A), gnathosoma (B) (Phase-contrast micrographs). Scale bars for A=100 µm, B=50 µm.
Fig. 9 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 9.ɹHalacarellus longus [nomen nudum], ˂ NSMT-Ac 14699. Body (A), genitoanal region (B), gnathosoma (C). Scale bars for A=100 µm, BC=50 µm.
Fig. 15 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 15.ɹCopidognathus globulosus, glass slide prepared by Dr. Makarova. A: front side, B: back side.
Fig. 24 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 24.ɹCopidognathus beringiensis [nomen nudum], ˂ NSMT-Ac 14710. Body (A), genitoanal region (B), first leg (C). Scale bars for A=200 µm, BC=50 µm.
Fig. 28 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 28.ɹAgaue kurilensis, holotype ˂ NSMT-Ac 14712. Idiosoma (A), gnathosoma (B), first leg (C). Scale bars for A=200 µm, B=50 µm, C=100 µm.
Fig. 10 in Type Specimens of Halacarid Mites by Dr. N. G. Makarova Relocated in the Collection of the National Museum of Nature and Science, Tsukuba, Japan
Fig. 10.ɹHalacarellus longus [nomen nudum], ˁ NSMT-Ac 14700. Body (A), genitoanal region (B), gnathosoma (C). Scale bars for A=100 µm, BC=50 µm.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.