Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,618
datasets available to search
ShareScore release 0.7.1
Dataset results
13,618 results for “biology”
UVA_VPTS - Vertical profiles of biological targets derived from weather radars in Belgium, Germany and the Netherlands
<p><em>UVA_VPTS - Vertical profiles of biological targets derived from weather radars in Belgium, Germany and the Netherlands</em> is a vertical profile time series dataset published by the <a href="https://www.inbo.be/en">Research Institute for Nature and Forest (INBO)</a>. It contains animal movement data derived from 24 weather radars in Belgium, Germany and the Netherlands, with varying coverage from 2008 to 2023. These data were created by processing weather radar data - provided by the Royal Meteorological Institute of Belgium (<a href="https://www.meteo.be/">RMI</a>), German Meteorological Service (<a href="https://www.dwd.de/">DWD</a>) and Royal Netherlands Meteorological Institute (<a href="https://www.knmi.nl/">KMNI</a>) - with methods optimized for extracting bird targets. The resulting data are vertical profile time series (VPTS), containing the density, speed and direction of biological targets within a weather radar (<code>radar</code>) volume, grouped into altitude bins (<code>height</code>) and measured over time (<code>datetime</code>). The data are also available in the <a href="https://aloftdata.eu/browse/?prefix=uva/">Aloft bucket</a>.</p> <p>See Desmet et al. (2025, <a href="https://doi.org/10.1038/s41597-025-04641-5">https://doi.org/10.1038/s41597-025-04641-5</a>) for a more detailed description of this dataset.</p> <h2>Files</h2> <p>VPTS data in this deposit are organized per country (.tgz file), radar (directory), year (directory) and month (.csv.gz file). Fields in the data follow the <a href="https://aloftdata.eu/vpts-csv/">VPTS CSV</a> format and are described in <code>vpts-csv-table-schema.json</code>. An overview of what data are available is provided in <code>coverage.csv</code>. Radar metadata can be found at <a href="https://aloftdata.eu/radars/">https://aloftdata.eu/radars/</a>.</p> <ul> <li><strong>coverage.csv</strong>: coverage of the VPTS data, representing the number of unique hours, heights, source files and records for each radar and date combination.</li> <li><strong>vpts-csv-table-schema.json</strong>: technical description of the fields in the VPTS data.</li> <li><strong>be.tgz</strong>: VPTS data from 3 radars in Belgium.</li> <li><strong>de.gz</strong>: VPTS data from 18 radars in Germany.</li> <li><strong>nl.gz</strong>: VPTS data from 3 radars in the Netherlands.</li> </ul> <h2>Acknowledgements</h2> <p>This dataset was processed using infrastructure provided by the University of Amsterdam, SURF Cooperative, Ghent University and the Research Institute for Nature and Forest (INBO). It was mainly supported by the <a href="https://globam.science/">GloBAM project</a>, funded through the 2017-18 Belmont Forum and BiodivERsA joint call for research proposals under the BiodivScen ERA-Net COFUND programme.</p>
Data from: Enamel proteins reveal biological sex and genetic variability within southern African Paranthropus
<p>This dataset contains the sequences of Paranthropus robustus, first described in 'Enamel proteins reveal biological sex and genetic variability within southern African Paranthropus', as well as the reference data and all the results from the analysis of those sequences.</p> <p><strong>Folders and Sub-Folders:</strong></p> <p><strong>- Paranthropus_Raw_AA_Sequences_Unaligned: </strong>Contains 2 fasta files. Paranthropus_Unaligned.fasta contains all the Paranthropus robustus sequences that were used for all of the analyses. Paranthropus_Unaligned_UNFILTERED.fasta contains all the Paranthropus robusts sequences <strong>before </strong><strong>filtering </strong>for SAP quality/confidence. These sequences were not used in any of the analyses, but are provided here for openness. </p> <p> </p> <p> </p> <p><strong>-</strong> <strong>Reference_Datasets</strong>: Contains 3 fasta files. Each fasta file is a reference dataset used in at least one analysis. The identity and origin of each sample is described in the supplementary document of the publication.</p> <p> </p> <p> </p> <p><strong>- Phylogenetic_Analysis_Datasets_and_Trees: </strong>Contains the following <strong>five folders</strong></p> <p> - <strong>Paranthropus_Alignments_All_Datasets</strong>: Contains three folders. Each folder contains the aligned and I/L corrected MSAs (Multiple Sequence Alignments) of Paranthropus robustus and a reference dataset.</p> <p> - <strong>Paranthropus_Diversity_Dataset_Trees_Results</strong>: Contains all analysis done using the 'diversity' reference dataset. Contains one folder for each protein, which includes the protein alignment and the phylogenetic tree of that protein. Additionally a folder named 'CONCATENATED' contains the concatenated alignemnts and trees. The BEAST2-STARBEAST3 folder contains the Starbeast3 analysis, including the xml, output log file, output trees and the input taxon set file.</p> <p> - <strong>Paranthropus_Representative_Dataset_Trees_Results:</strong> Contains all analysis done using the 'representative' reference dataset. Contains one folder for each protein, which includes the protein alignment and the phylogenetic tree of that protein. Additionally a folder named 'CONCATENATED' contains the concatenated alignemnts and trees. The BEAST2 folder contains the time-calibrated BEAST2 analysis, including the xml, output log file, output trees. The folder Distance_Matrix contains the generated distance matrix and the Rscript used to generate the heatmap from it.</p> <p> - <strong>Paranthropus_Independent_Dataset_Trees_Results: </strong>Contains all nexus files and tree-figures used in the analysis of the 'independent' reference dataset. </p> <p> - <strong>Tree_Figures: </strong>Contains three sub-folders and an additional figure. Each sub-folder contains the phylogenetic tree figures generated using one of the three reference datasets.</p>
Data for "Detection of metabolite-protein interactions in complex biological samples by high-resolution relaxometry: towards interactomics by NMR"
<p>Raw NMR data for relaxometry experiments, divided by donor sample. For every donor sample 2 or 3 different samples were used in order to record data at 19 different magnetic fields.</p> <p>Data from fast field-cycling relaxometry. All the data is in one xlsx file, divided by donor sample.</p> <p>Relaxometry results for alanine, lactate, creatinine and glutamine, obtained from the fitting of their relaxation decays recorded at 19 different fields, divided by donor sample.</p>
Nicotiana benthamiana as a model organism for plant biology study
<p><em>Nicotiana benthamiana</em> is an amenable model organism for plant biology study. Several functional genomics tools, including viral vectors, RNAi, ethylmethanesulfonate (EMS) mutagenesis, CRISPR-mediated genome editing, and agroinfiltration, are available in the <em>N. benthamiana</em> experimental system. These tools can be applied to research in genomics, biochemistry, metabolomics, cell biology and pathology as well as other topics in plant biology.</p> <p>*This is an updated graphical abstract for commnetary article "Dude, where is my mutant? <em>Nicotiana benthamiana</em> meets forward genetics" (Derevnina et al., 2019, New Phytologist 221(2):607-610).</p>
Dataset for: Multi-scale approach to biodiversity proxies of biological control service in European farmlands
<p>Dataset for the BiodivERsA COFUND Woodned project. Information on which spatio-temporal factors are simultaneously affecting crop pests and their natural enemies is required to improve conservation biological control practices. The study was conducted in 80 winter wheat crop fields distributed in three regions of North-western Europe (Brittany, Hauts-de-France and Wallonia), along intra-regional gradients of landscape complexity. Five taxa : aphids, slugs, spiders, carabids, and parasitoids were sampled for two consecutive years. We analysed the influence of regional, landscape and local factors on the abundance and species richness of crop-dwelling organisms, as proxies of the service/disservice they provide. Firstly, there was higher biocontrol potential in areas with mild winter climatic conditions. Secondly, natural enemy communities were less diverse and had lower abundances in landscapes with high crop and wooded continuities, contrary to slugs and aphids. Finally, field boundaries with grass strips were more favourable to spiders and carabids than boundaries formed by hedges, while the opposite was found for crop pests, with the latter being less abundant towards the centre of the fields. These results are quite unexpected because they show that hedgerows and woodlots should not be the unique cornerstones of agro-ecological landscape design strategies. We point out that combining woody and grassy habitats to take full advantage of the features and ecosystem services they both provide may promote sustainable agricultural ecosystems. It may be possible to both reduce pest pressure and promote natural enemies by accounting for taxa-specific antagonistic responses to multi-scale environmental characteristics.</p>
Extended data for Manuscript: Identification of potential biological targets of oxindole scaffolds via in silico repositioning strategies
<p>This is the Extended Data for the manuscript "<strong>Identification of potential biological targets of oxindole scaffolds via <em>in silico</em> repositioning strategies" </strong>submitted to F1000 Research.</p> <p>Extended Data include a list of all the accession codes as mentioned in the text, the results of 2D fingerprint-based similarity analyses and ligand-protein complexes predicted by rigid docking and Induced Fit Docking calculations.</p>
Biological data science courses at UMONS, Belgium: student's activity for 2019-2020
<p>Progression of the students in the different exercises of the biological data science courses at the University of Mons, Belgium for the academic year 2019-2020.</p> <p>Activity of the students was recorded to monitor their individual progression in asynchronous exercises. The courses were taught in flipped classroom by Philippe Grosjean (<a href="mailto:philippe.grosjean@umons.ac.be">philippe.grosjean@umons.ac.be</a>) and Guyliann Engels (<a href="mailto:guyliann.engels@umons.ac.be">guyliann.engels@umons.ac.be</a>) the University of Mons. These authors designed almost all the teaching material, the exercises, and the related software. The courses were also taught at the Campus Charleroi by Raphaël Conotte (<a href="mailto:raphael.conotte@umons.ac.be">raphael.conotte@umons.ac.be</a>) that also contributed to a part of the learnr exercises and of the inline course.</p> <p><strong>How to use these data?</strong></p> <p>The README file provides detailed information on the purpose, collection and management of the data. The data are presented in tabular format in CSV files. Metadata in the `datapackage.json` document the different tables and their fields. It is in the Frictionless data format (<a href="https://frictionlessdata.io/">https://frictionlessdata.io</a>). You can get a view of a part of these metadata by uploading the file `datapackage.json` into the inline data package creator at <a href="https://create.frictionlessdata.io/">https://create.frictionlessdata.io</a>. There is a large set of libraries and tools for different programming languages available at <a href="https://frictionlessdata.io/tooling/libraries/">https://frictionlessdata.io/tooling/libraries/</a>. Otherwise, any CSV library should import the data in your favourite software. Please, note that encoding is UTF8. For R, the {learnitdown} package provides specific functions to import these data and/or convert them in a SQLite database (<a href="https://www.sciviews.org/learnitdown/">https://www.sciviews.org/learnitdown/</a>).</p> <p>For any question, send an email at <a href="mailto:sdd@sciviews.org">sdd@sciviews.org</a>.</p>
Biological data science courses at UMONS, Belgium: student's activity for 2020-2021
<p>Progression of the students in the different exercises of the biological data science courses at the University of Mons, Belgium for the academic year 2020-2021.</p> <p>Activity of the students was recorded to monitor their individual progression in asynchronous exercises. The courses were taught in flipped classroom by Philippe Grosjean (<a href="mailto:philippe.grosjean@umons.ac.be">philippe.grosjean@umons.ac.be</a>) and Guyliann Engels (<a href="mailto:guyliann.engels@umons.ac.be">guyliann.engels@umons.ac.be</a>) the University of Mons. These authors designed almost all the teaching material, the exercises, and the related software. The courses were also taught at the Campus Charleroi by Raphaël Conotte (<a href="mailto:raphael.conotte@umons.ac.be">raphael.conotte@umons.ac.be</a>) that also contributed to a part of the learnr exercises and of the inline course.</p> <p><strong>How to use these data?</strong></p> <p>The README file provides detailed information on the purpose, collection and management of the data. The data are presented in tabular format in CSV files. Metadata in the `datapackage.json` document the different tables and their fields. It is in the Frictionless data format (<a href="https://frictionlessdata.io">https://frictionlessdata.io</a>). You can get a view of a part of these metadata by uploading the file `datapackage.json` into the inline data package creator at <a href="https://create.frictionlessdata.io">https://create.frictionlessdata.io</a>. There is a large set of libraries and tools for different programming languages available at <a href="https://frictionlessdata.io/tooling/libraries/">https://frictionlessdata.io/tooling/libraries/</a>. Otherwise, any CSV library should import the data in your favourite software. Please, note that encoding is UTF8. For R, the {learnitdown} package provides specific functions to import these data and/or convert them in a SQLite database (<a href="https://www.sciviews.org/learnitdown/">https://www.sciviews.org/learnitdown/</a>).</p> <p>For any question, send an email at <a href="mailto:sdd@sciviews.org">sdd@sciviews.org</a>.</p>
Biological data science courses at UMONS, Belgium: student's activity for 2018-2019
<p>Progression of the students in the different exercises of the biological data science courses at the University of Mons, Belgium for the academic year 2018-2019.</p> <p>Activity of the students was recorded to monitor their individual progression in asynchronous exercises. The courses were taught in flipped classroom by Philippe Grosjean (<a href="mailto:philippe.grosjean@umons.ac.be">philippe.grosjean@umons.ac.be</a>) and Guyliann Engels (<a href="mailto:guyliann.engels@umons.ac.be">guyliann.engels@umons.ac.be</a>) the University of Mons. These authors designed almost all the teaching material, the exercises, and the related software.</p> <p><strong>How to use these data?</strong></p> <p>The README file provides detailed information on the purpose, collection and management of the data. The data are presented in tabular format in CSV files. Metadata in the `datapackage.json` document the different tables and their fields. It is in the Frictionless data format (<a href="https://frictionlessdata.io/">https://frictionlessdata.io</a>). You can get a view of a part of these metadata by uploading the file `datapackage.json` into the inline data package creator at <a href="https://create.frictionlessdata.io/">https://create.frictionlessdata.io</a>. There is a large set of libraries and tools for different programming languages available at <a href="https://frictionlessdata.io/tooling/libraries/">https://frictionlessdata.io/tooling/libraries/</a>. Otherwise, any CSV library should import the data in your favourite software. Please, note that encoding is UTF8. For R, the {learnitdown} package provides specific functions to import these data and/or convert them in a SQLite database (<a href="https://www.sciviews.org/learnitdown/">https://www.sciviews.org/learnitdown/</a>).</p> <p>For any question, send an email at <a href="mailto:sdd@sciviews.org">sdd@sciviews.org</a>.</p>
Relaxation anisotropy of quantitative MRI parameters in biological tissues
<p>Dataset for the manuscript "Relaxation anisotropy of quantitative MRI parameters in biological tissues" published in Scientific Reports 2022</p>
Testing a biological mechanism of the insurancehypothesis in experimental aquatic communities - Data Deposit
<p>1.The insurance hypothesis predicts a stabilizing effect of increasing species richness on commu-nity and ecosystem properties. Difference among species’ responses to environmental fluctuationsprovides a general mechanism for the hypothesis. Previous experimental investigations of theinsurance hypothesis have not examined this mechanism directly.</p> <p>2.First, responses to temperature of four protist species were measured in laboratory microcosms.For each species, we measured the response of intrinsic rate of increase (r) and carrying capacity(K) to temperature.</p> <p>3.Next, communities containing pairs of species were exposed to temperature fluctuations. Com-munity biomass varied less when correlation inKbetween species (but notr) was more negative,and this resulted from more negative covariances in population sizes, as predicted. Results werecontingent on species identity, with findings differing between analyses including or not includingcommunities containing one particular species.</p> <p>4.These findings provide the clearest support to date for this mechanism of the insurance hypo-thesis. Biodiversity, in terms of differences in species’ responses to environmental fluctuations (i.e.functional response diversity) stabilizes community dynamics.</p>
BioSR+: Dataset Extension of biological images for super-resolution microscopy
<p>BioSR+ dataset is an extension of our pre-published BioSR dataset of biological images for super-resolution microscopy, currently including image pairs of low-and-high resolution images of five biology structures (CCPs, ER, MTs, F-actin, Myosin-IIA) and 8 signal levels for each ROI. The BioSR+ dataset is related to our Nature Methods paper "Evaluation and development of deep neural networks for image super-resolution in optical microscopy" (DOI: 10.1038/s41592-020-01048-5) and Nature Biotechnology paper "Rationalized deep learning super-resolution <br> microscopy for sustained live imaging of rapid subcellular processes" (DOI:10.1038/s41587-022-01471-3). Both BioSR and BioSR+ are freely available and can be used for non-commercial purposes with proper citations of above two papers.</p>
Antibiotic resistant pathogen outbreak investigation: an interdisciplinary module to teach fundamentals of evolutionary biology
<p>The evolution of resistance to antibiotics provides a timely and relevant topic for teaching undergraduate students evolutionary biology. Here, we present a module incorporating modified sequencing data from eight antibiotic resistant pathogen outbreaks in hospital settings with bioinformatics and phylogenetic analyses. This module uses whole genome sequencing data from hospital outbreaks investigated by the Centers for Disease Control and Prevention to provide examples of antibiotic resistance spread. Students work in groups to analyze outbreak data to identify the bacterial species and antibiotic resistance genes, to infer a phylogenetic tree examining relatedness among isolates, and to determine a possible source of the outbreak. Students then compile their results in individual reports and provide recommendations for preventing the further spread of antibiotic resistant organisms. In addition to providing genomic outbreak data, we include a teaching concepts guide discussing three integral components of the module: how evolutionary biology concepts of natural selection and competition impact antibiotic resistance; outbreak investigation information to aid in phylogenetic analysis and creation of recommendations; and instructions for the bioinformatics protocol. Completion of this module provides students an opportunity to think critically about the evolution of resistance, practice bioinformatics techniques, and relate evolutionary biology to current events.</p>
Data: Stability and biological response of PEGylated gold nanoparticles
<p><span>This dataset is focused on thermal stability of PEGylated Au NPs at 4 and 37 °C and after sterilization in autoclave.</span></p>
Quantification of ADHD Medication in Biological Fluids with Liquid Chromatography: A Comprehensive Review - Metadata
<p>This file is the metadata related to the publication "Quantification of ADHD Medication in Biological Fluids with Liquid Chromatography: A Comprehensive Review".</p>
Data from: Collaborative Research: Influence of phosphorus deficiency on enigmatic biological methane production in oxic freshwater lakes
<p>Data from: Collaborative Research: Influence of phosphorus deficiency on enigmatic biological methane production in oxic freshwater lakes</p> <p>NSF Projects 1951002 (PI: Matthew J. Church), 1950963 (PI: John E. Dore)</p>
Data from: BioEncoder: a metric learning toolkit for comparative organismal biology
<p><strong>BioEncoder: a metric learning toolkit for comparative organismal biology</strong></p> <p><strong>Abstract </strong>- In the realm of biological image analysis, deep learning (DL) has become a core toolkit, e.g., for segmentation and classification. However, conventional DL methods are challenged by large biodiversity datasets characterized by unbalanced classes and hard-to-distinguish phenotypic differences between them. Here we present BioEncoder, a user-friendly toolkit for metric learning, which overcomes these challenges by focussing on learning relationships between individual data points rather than on the separability of classes. BioEncoder is released as a Python package, created for ease of use and flexibility across diverse datasets. It features taxon-agnostic data loaders, custom augmentation options, and simple hyperparameter adjustments through text-based configuration files. The toolkit's significance lies in its potential to unlock new research avenues in biological image analysis while democratizing access to advanced deep metric learning techniques. BioEncoder focuses on the urgent need for toolkits bridging the gap between complex DL pipelines and practical applications in biological research.</p> <p><strong>Dataset </strong>- This data repository includes two things: a snapshot of the BioEncoder package (BioEncoder-main.zip, version 1.0.0, downloaded from https://github.com/agporto/BioEncoder on 2024-07-19 at 17:20), and the damselfly dataset used for the case study presented in the paper (bioencoder_data.zip). The dataset archive also encompasses the configuration files and the final model checkpoints from the case study, as well as a script to reproduce the results and figures presented in the paper.</p> <p><strong>How to use - </strong>Get started by consulting the <a href="https://github.com/agporto/BioEncoder?tab=readme-ov-file#quickstart">GithHub repository</a> for information on how to install BioEncoder, then download the <a href="../records/10909614/files/BioEncoder-data.zip?download=1&preview=1">data archive</a> and run the script. Some parts of the script can be executed using the model checkpoints, for orther parts the training rountine needs to be run. </p>
Data from: Normalizing gas-chromatography–mass spectrometry data: method choice can alter biological inference
<p>Gas-Chromatography Mass Spectrometry data from European badger (<em>Meles meles</em>) sub-caudal gland secretion used in:</p> <p>Noonan, M.J., Tinnesand, H.V.,<sup> </sup>and Buesching, C.D. (2018). Normalizing gas-chromatography–mass spectrometry data: method choice can alter biological inference. BioEssays, 40(6): 0-0. DOI: 10.1002/bies.201700210.</p>
Spectral irradiance at Lammi Biological Station Research Forest 2015: for assessing scale-wise similarity of curves with a thick pen
<p>This dataset contains records of the solar spectral energy irradiance (W m<sup>-2</sup> nm<sup>-1</sup>) in the understorey of forest stands at Lammi Biological Station, southern Finland (61◦ 3.24’ N, 25◦ 118 2.23’ E) during the spring of 2015. These spectra allow the change in spectral energy irradiance to be followed through the period of canopy leaf flush. Records are the average of recorded spectra from four points recorded at 40-cm above the forest floor using a Maya 2000 Pro array spectrometer. Spectra were recorded from exactly the same location on three dates, 2015-04-25, 2015-05-22, and 2015-06-05, before, during and after leaf flush. Data were recorded from the understorey of a young Betula stand, an old Betula stand, an old mixed Betula stand, a Quercus stand, and a Picea stand, in three positions: shade, semi-shade from leaves, and full sun in a sunfleck. On each occasion control measurements of spectral energy irradiance in full sun of an open field were also recorded at the beginning, middle and end of each measurement period. All measurements were made during the 2 hours either side of solar noon, on clear-sky days. Details of the sampling method and interpretation are given in the paper, Hartikainen et al., (2018) in Ecology and Evolution, which showcases the use of Thick Pen Transform to compare spectra.</p>
Agriculture - General: biological diversity 6
<p>A database of tree of biological, cultural, ecological or historical interest because of their age, size or condition. National Biodiversity Data Centre (2016). Heritage Trees of Ireland. Occurrence dataset <a href="https://doi.org/10.15468/9athfc">https://doi.org/10.15468/9athfc</a> accessed via GBIF.org</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.