Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
32,629
datasets available to search
ShareScore release 0.7.1
Dataset results
32,629 results for “Dataset”
Harmonized National Land Cover Dataset Values for HydroBASINS Basins
Water quality is largely reflective of processes occurring on the surrounding landscape. While national landcover data are widely available via remotely sensed products, they are usually not aggregated in a manner that is expeditiously merged with basin-level data. To facilitate national-scale analyses of basin-level landcover with co-located water quality data, we present aggregated land cover data for the Contiguous United States. Data are aggregated using the HydroBASINS basin shapefiles. HYBAS_ID is retained to enable merging with HydroBASINS parent datasets.
Interagency Ecological Program: Integrated Dataset of Phytoplankton Enumeration Data in the San Francisco Estuary, 1992-2024
Phytoplankton community composition is an important driver of zooplankton productivity and food supply for higher trophic levels in the San Francisco Estuary. Various monitoring surveys throughout the region collect phytoplankton enumeration data dating back to the 1990s. These include surveys from the CA Department of Water Resources (CADWR), CA Department of Fish and Wildlife (CDFW), the US Bureau of Reclamation (USBR), and the US Geological Survey (USGS). These surveys collect data via various sampling and laboratory methods which are not always directly comparable. This integrated dataset includes both enumeration counts and well-documented metadata to allow for informed decision-making in the integration of these data. It also standardizes taxonomic names between groups via a key list. Note that, in this dataset, we make conservative decisions about taxonomic resolution to ensure maximum compatibility between groups. For more detailed metadata and higher taxonomic resolution, refer to individual surveys’ publications or reach out to their primary contact.
Which multiband factor should you choose for your resting-state fMRI study? The Emory Multiband Dataset
Open the record for dataset details and reuse information.
CARMEN immunopeptidomics publication associated dataset
<h1>CARMEN: CAnceR imMunopeptidogENomics</h1> <blockquote> <div>An immunopeptidomic dataset accompanying the publication "<a href="http://dx.doi.org/10.1101/2025.05.08.651510" target="_blank" rel="noopener">Expanding the definition of MHC Class I peptide binding promiscuity to support vaccine discovery across cancers with CARMEN</a>" containing peptides determined by mass-spectrometry associated with MHC Class I bindings from 72 publications (2323 samples).</div> </blockquote> <div> </div> <div> <p><strong>Authors:</strong> <a href="mailto:aleksander.palkowski@gmail.com" target="_blank" rel="noopener">Aleksander Palkowski</a>*, <a href="mailto:mwaleron@gmail.com" target="_blank" rel="noopener">Michal Waleron</a>*, <a href="mailto:emilia.daghir@gmail.com" target="_blank" rel="noopener">Emilia Daghir-Wojtkowiak</a>*, <a href="mailto:ashwinkallor@gmail.com" target="_blank" rel="noopener">Ashwin Adrian Kallor</a>*, <a href="mailto:javier.alfaro@proteogenomics.ca" target="_blank" rel="noopener">Javier Antonio Alfaro</a></p> <p>* <em>These authors contributed equally to this work</em></p> </div> <div> </div> <div>The entire dataset consist of four table files in the <a href="https://parquet.apache.org" target="_blank" rel="noopener">Apache Parquet</a> data file format:</div> <ul> <li>main</li> <li>mapped-protein-annotations-pogo</li> <li>mapped-protein-annotations-msfragger</li> <li>hla-sequences</li> </ul> <div> </div> <div><strong>Please refer to the README.md file for details.</strong></div> <div> </div>
Damage assessment of a physical beam reinforced with masses - dataset
<p>The dataset beam-signal contains the spectrum vibration signals in the frequency domain measured from a beam reinforced with masses under healthy and faulty conditions. This data is for a commonly used system in various industrial applications. The data can be used for online condition process monitoring to detect and diagnose any anomaly or faulty condition in the system. Hence, the datasets provide the geometric and experimental measurements performed on the beam reinforced with masses for various mass losses considered structural damage. The collected data included the following datasets:</p> <ul> <li>Dataset Mass-position contains 70 sampling positions for the six masses attached to the beam. (<a href="../api/records/8081690/draft/files/Mass%20position.xlsx/content">Mass position</a>)</li> <li>Dataset DI contains 280 damage indexes calculated using the FRAC method. (<a href="../api/records/8081690/draft/files/DI_FRAC_Exp-estimation.xlsx/content">DI_FRAC_Exp-estimation</a>)</li> <li>Dataset beam-signal includes 280 inertances responses magnitudes and respective phases considering 70 samples of healthy and 210 sampled of damaged conditions ( <a href="../api/records/8081690/draft/files/Dataset%20Beam-signal_Healthy.zip/content">Dataset Beam-signal_Healthy, </a><a href="../api/records/8081690/draft/files/Dateset%20Beam-signal_Damaged-2.96.zip/content">Dateset Beam-signal_Damaged-2.96, </a></li> </ul> <p><a href="../api/records/8081690/draft/files/Dataset%20Beam-signal_Damaged-5.92.zip/content"> Dataset Beam-signal_Damaged-5.92, </a><a href="../api/records/8081690/draft/files/Dataset%20Beam-signal_Damaged-8.87.zip/content">Dataset Beam-signal_Damaged-8.87) .</a></p> <p>The dataset beam-signal can be used to develop structural health monitoring techniques for detecting damage and anomalies in the structure. The dataset's Mass-position and DIs can impose parametric uncertainty in the experiment. Stochastic and damage identification metrics can be used for further insights on new monitoring and control techniques. Since the tests include paramedic uncertainty, they can also be employed in uncertainty quantification, stochastic modelling and supervised and unsupervised machine learning techniques. </p> <p>Therefore, the datasets are intended to benefit the scientific community investigating the dynamics of structures and readers interested in experimental practices applied to systems and modelling. These datasets can be used for numerical model validation, identification techniques, uncertainty quantification, machine learning, and structural integrity monitoring algorithms based on experimental measurement samples on the beam reinforced with mass.</p> <p>A detailed description of the experiment can be found in </p> <p>[1] Sousa, A.A.S.R., da Silva Coelho, J., Machado, M.R. et al. Multiclass Supervised Machine Learning Algorithms Applied to Damage and Assessment Using Beam Dynamic Response. J. Vib. Eng. Technol. (2023). https://doi.org/10.1007/s42417-023-01072-7</p> <p>[2] Monitoramento da Integridade Estrutural de Vigas utilizando Técnicas de Aprendizado de Máquina, 2023. Mestrado em Integridade de Materiais da Engenharia - Universidade de Brasília (In Portuguese)</p> <p>[3] Amanda A.S.R. de Sousa, Marcela R. Machado, Experimental vibration dataset collected of a beam reinforced with masses under different health conditions, Data in Brief, 2024, 110043, ISSN 2352-3409, https://doi.org/10.1016/j.dib.2024.110043.</p>
Anyskop Blowout Prehistoric Dataset, Western Cape, South Africa
<p>These Stone Age archaeological datasets were collected in 2001 and 2002 by a team from the Department of Early Prehistory and Quaternary Ecology of the University of Tübingen (Germany) headed by Nicholas J. Conard. Many South African researchers collaborated on this project, with Pippa Haarhoff, John Compton, Dave Roberts, and Stephan Woodborne deserving special mention.</p> <p>The field work took place at the Anyskop Blowout (ANY1) located within the West Coast Fossil Park near Langebaanweg, Western Cape, South Africa. The field work was conducted with the help of students from the universities of Tübingen and Cape Town. The datasets are predominantly in English (with some German as well) and include field data in the MAIN table. Further analytical data for many classes of artifacts include: LITHICS, FAUNA, POTTERY, MODERN, BUCKETS, REFITS.</p> <p>All collected materials are curated by the Iziko South African Museums in Cape Town under accession numbers SAM-AA-8903 (finds collected by other teams before 2001) and SAM-AA-9007 (finds from this study, 2001-2002). Some of the finds are exhibited in the museum at the West Coast Fossil Park.</p> <p>Funding for this research project came mainly from the German Research Foundation (DFG - CO 226/5-1, 5-2, 5-5 and 5-6) and the University of Tübingen. Significant support was provided by the Iziko South African Museums, the West Coast Fossil Park, and the University of Cape Town.</p>
Dataset: Environmental benchmarks for European Cement Industry
<p>This dataset contains the information relative to the article "Environemntal benchmarks for European cement industry".</p> <p><a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.spc.2024.01.020" target="_blank" rel="noopener">Reference paper</a></p> <p><a href="https://www.researchgate.net/publication/377796848_Environmental_benchmarks_for_the_European_cement_industry" target="_blank" rel="noopener">ResearchGate link</a></p>
Dataset for algorithmic thinking skills assessment: Results from the virtual CAT large-scale study in Swiss compulsory education
<p><strong>Overview</strong><br>This dataset was collected during a main study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.<br>As algorithmic thinking becomes increasingly vital in our digital age, this study bridges the gap between traditional assessments and the needs of today's learners by introducing a digital platform. The virtual CAT, a digital adaptation of an unplugged assessment activity, offers scalable, automated assessments with reduced human intervention.</p> <p><strong>Study Context, Location and Participants</strong><br>To comprehensively investigate algorithmic competencies within compulsory education, exploring their variations and determining the factors influencing them, in Spring 2023 we conducted an experimental study with the virtual CAT's.<br>The sample comprises 129 students (65 girls and 64 boys), selected from nine classes across five public schools in Ticino and Solothurn cantons.</p> <p><strong>Data Collection</strong><br>During the data collection process, session and participant details were manually recorded by the administrator. <br>Each session has been assigned a unique identifier, and specific details, such as the date, canton, school name and type, and the students’ HarmoS grade (HG) level, have been recorded. <br>Student information are limited to sex and date of birth, with birth dates used to calculate ages, a significant factor in our demographic analysis. <br>To protect student privacy, unique identifiers have been assigned to each participant, keeping the data anonymous and secure. <br>The assessment tool automatically tracked all user interaction within the platform.<br>All data collected have been pseudonymised, aligning with prevailing open science practices in Switzerland (SNSF, 2021). <br>Data collection was integrated into a validation module of the app. </p> <p><strong>Data Features</strong><br>The dataset comprises the following files:</p> <ul> <li>STUDENTS_SESSIONS.csv</li> <li>RESULTS.csv</li> <li>LOGS.csv</li> <li>CANTONS.csv</li> <li>ALGORITHMS.csv</li> </ul> <p>These files collectively provide insights into the algorithmic actions of the students, demographic details, session logs, results, and more.</p> <p><strong>Usage & Ethics</strong><br>In the spirit of open science, this dataset is made available to the public after meticulous anonymisation to ensure all participants' privacy and ethical treatment. <br>Initial authorisations were secured from school administrators, teachers, and parents. <br>Detailed communication regarding the study's nature, data handling, and objectives was transparently shared with all stakeholders.</p> <p><strong>REFERENCES</strong></p> <p><strong>[1]</strong> A. Piatti, G. Adorni, L. El-Hamamsy, L. Negrini, D. Assaf, L. Gambardella & F. Mondada. (2022). The CT-cube: A framework for the design and the assessment of computational thinking activities. Computers in Human Behavior Reports, 5, 100166. <a href="https://doi.org/10.1016/j.chbr.2021.100166">https://doi.org/10.1016/j.chbr.2021.100166</a></p> <p><strong>[2]</strong> Adorni, G., & Piatti, S., & Karpenko, V. (2023). virtual CAT: An app for algorithmic thinking assessment within Swiss compulsory education. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10027851">https://doi.org/10.5281/zenodo.10027851</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-app/</a></p> <p><strong>[3]</strong> Adorni, G., & Karpenko, V. (2023). virtual CAT programming language interpreter. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10016535">https://doi.org/10.5281/zenodo.10016535</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-programming-language-interpreter/</a></p> <p><strong>[4]</strong> Adorni, G., & Karpenko, V. (2023). virtual CAT data infrastructure. Zenodo Software. <a href="https://doi.org/10.5281/zenodo.10015011">https://doi.org/10.5281/zenodo.10015011</a> On GitHub: <a href="https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure">https://github.com/GiorgiaAuroraAdorni/virtual-CAT-data-infrastructure</a></p> <p> </p>
Multi-Class Depression Detection Dataset
<p>This dataset was created as part of the Master's thesis titled "Multi-Class Depression Detection Through Tweets Using Artificial Intelligence." It contains tweets labeled for five types of depression (Bipolar, Major, Psychotic, Atypical, and Postpartum) using lexicons verified by psychiatrists. </p> <p>Purpose: Designed for multi-class classification of depression using AI, focusing on Explainable AI for highlighting key words in the tweets influencing the predictions.<br>Applications: The dataset is suitable for research in natural language processing, sentiment analysis, mental health prediction, and Explainable AI.</p> <p>This dataset is shared under the Creative Commons Attribution 4.0 International (CC BY) license, requiring proper attribution for any use or modification.</p>
Dataset of reports about MOF-based SERS substrates since 2011 until March 2023. Structure, characteristics, analytes, and performances.
<p>This dataset was generated to aid the creation of a review article addressing the use of Metal-Organic Frameworks (MOF)-based Surface Enhanced Raman Spectroscopy (SERS) platforms for the detection of Volatile Organic Compounds (VOCs).</p> <p>This dataset was generated employing the Web of Science database, encompassing manuscripts published up to March 2023. A literature search was initially conducted using a combination of keywords, including "MOF," "Metal-Organic Framework," "SERS," "Surface Enhanced Raman Spectroscopy," and "Surface Enhanced Raman Scattering." This search spanned the "Topic" category, enabling exploration across title, abstract, author keywords, and keyword-plus fields.</p> <p>From the initial pool of 238 documents, review articles and duplicates were systematically excluded, resulting in a refined collection of 182 articles. Subsequently, articles not concurrently addressing MOF and SERS or those utilizing MOF as sacrificial templates were further excluded, resulting in a final subset of 72 articles. From this curated set, relevant parameters were extracted, resulting in 229 entries for the dataset. </p> <p>Characteristics about the structure (in terms of MOF type and configuration; Plasmonic element type and configuration), target analyte (including type, phase, and incubation time), measurement specifications (in terms of laser, laser power, exposure time), and performance of the MOF-based SERS substrates were collected.</p> <p>Listed references 1-72 correspond with the manuscript number in the dataset.</p> <p>Listed references 73-80 correspond with references for selected examples of MOF pore diameters.</p>
EU-TRHeaDS Conjoint Dataset
<p>The EU-TRHeaDS Conjoint Dataset is a set of 20,920 observations that was gathered, organised and edited in the framework of the research project ‘EU Citizens’ Transnational Rights and Health-related Deservingness at the Street-level - EU-TRHeaDS’ (PI: Roberta Perna), which has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie grant agreement No 101022244.</p> <p>One of EU-TRHeaDS' aims is to investigate which criteria ‘activate’ the category of healthcare (un)deservingness in the context of intra-EU migration among the general public in two EU Member States (Belgium and Spain), and the extent to which these preferences turn into patterns of systematic penalisation towards specific EU nationality groups. It does so by carrying out a conjoint experimental study nested in an online survey run in parallel in Belgium and Spain with a representative sample of the population on the dimensions of gender, age (18 years old), level of education achieved and geographical region of residence. During the four tasks of the experiment, respondents were asked who they would prioritise to access publicly-funded healthcare out of two fictitious patients who differed in four attributes, all randomly assigned: 1) nationality; 2) migration trajectory; 3) responsibility over ill health, and 4) employment status.</p> <p>As a subset of a larger survey on intra-EU mobility and access to healthcare rights, the EU-TRHeaDS Conjoint Dataset specifically includes the socio-demographic variables of the probabilistic sample in each country, the variables of the conjoint experiment and information about the time spent by respondents in completing each of the four experimental tasks.</p> <p>For detailed information and the codebook, see the document 'EU-TRHeaDS_conjoint_Description&Codebook'</p>
Collection of figures to explore intra-regime weather variability of North Atlantic-European year-round weather regimes as Supplementary Dataset for Gerighausen et al. (2024)
<p>This is a supplementary dataset accompanying the publication <strong>Gerighausen et al. (2024) </strong>submitted to Meteorological Applications. It contains a collection of browsable figures, complementing selected regimes, seasons, and countries in the paper. The figures are provided as a zipped archive. The ZIP-File (1.2 GB) contains 4 subfolders and 4 auxiliary files as described in <strong>readme.md </strong>in the main folder. Once downloaded and unpacked, the .html navigation panels can be used in any browser to navigate through the plots. </p> <p>Data and methods used to generate the figures are explained in Gerighausen et al. (2024). In brief the analysis is based on ERA5 reanalysis 1979-2021 at 1° grid spacing and 6h temporal resolution aggregated to daily data. Anomalies are computed with respect to a 31-day running mean climatology. The figures are explained in the table below and in the navigation panel.</p> <p><strong>Gerighausen</strong>, J., J. Dorrington, M. Osman, and C. M. Grams, <strong>2024</strong>: Quantifying intra-regime weather variability for energy applications, <em>submitted to Meteorological Applications.</em> <a href="https://doi.org/10.48550/arXiv.2408.04302">doi:10.48550/arXiv.2408.04302</a></p>
Dataset of "Thermal Stability of Valuable Metals in Lithium-Ion Battery Cathode Materials: Temperature Range 100-400 °C"
<p>Lithium is crucial in lithium-ion batteries (LIBs), serving as a main component of the electrolyte and cathode. Elements such as cobalt, nickel, and manganese are also vital for high performance, energy density, and stability. This study aimed to examine the behaviour of end- of-life cathode material (LiNi0.6Mn0.2Co0.2O2) and its valuable metals after exposure to temperatures between 100 and 400 °C, comparing it with untreated material. The lithium content cannot be reliably determined by conventional analytical methods, so inductively coupled plasma optical emission spectroscopy (ICP-OES) was chosen for this purpose. For ICP-OES measurements, samples were dissolved in different solvents for a specified time, and the concentrations of lithium, nickel, manganese, and cobalt were measured. From the measured values, their theoretical yields were calculated. Due to the annealing at given temperatures and subsequent dissolution, this step can be considered as the first stage of the pyrometallurgical- hydrometallurgical process used in battery recycling. The study was complemented by further analyses to monitor the effect of annealing temperatures on the properties of the material. Based on the results, it was found that the highest theoretical yield in this temperature range was for material annealed at 400 °C and dissolved in 20% nitric acid for 4 hours.</p>
Leaf spectroscopy and active fluorescence datasets for early drought and nitrogen stress diagnosis in tomato
<p>The dataset contains different plant physiological parameters collected during a 14-day stress and recovery experiment on tomato (<em>Solanum lycopersicum</em> L. cv Moneymaker) plants, undergoing a nitrogen deficiency, drought or control treatment. </p> <p>A full description of the experiment, together with the scientific results, is published by Pescador-Dionisio et al. (2024), and can be found through: <a href="https://doi.org/10.1111/nph.20253">https://doi.org/10.1111/nph.20253.</a></p> <p>The goal of the dataset collection was to obtain a non-invasive proximal sensing dataset at leaf level (reflectance, transmittance, upward and downward fluorescence), in parallel to gas exchange and active fluorescence measurements. The leaf spectroscopy dataset was further processed by a pigment spectral unmixing algorithm according to Van Wittenberghe et al. (2024), to calculate fluorescence quantum efficiency (<em><strong>FQE</strong></em>) and effective absorbance (<strong><em>A_eff</em></strong>) changes associated to the activation of regulated heat dissipation (<strong><em>A_eff_535_Xan</em></strong>). The latter absorption feature is linked to the xanthophyll ('<strong>Xan</strong>') absorption in the 500-600 nm range, which is modelled by the sum of three Gaussians. For a full description of this feature, see Van Wittenberghe et al. (2021).</p> <p>Gas exchange and active fluorescence measurements were carried out with a LI-6400 portable photosysthesis system (LI-COR Biosciences, Lincoln, USA) equipped with a 6400-40 leaf chamber fluorometer. Steady-state measurements were done at 300 and 1000 μmol m−2 s−1 ('<strong><em>PAR300</em></strong>' and '<em><strong>PAR1000</strong></em>'), i.e. growing light conditions and light saturating conditions. Light response curves were taken on different days. Common fluorescence parameters (e.g., <em><strong>Fv/Fm, Fo, Fm, NPQ, YNO, YNPQ</strong></em>) are provided together with 'sustained' and reversible' NPQ parameters calculated according Porcar-Castell (2011).</p> <p>Leaf spectroscopy and active steady-state fluorescence measurements were performed on the same measuring days ('<em><strong>d0</strong></em>', '<em><strong>d2</strong></em>', '<em><strong>d4</strong></em>', '<em><strong>d7</strong></em>', '<em><strong>d14</strong></em>') and on the same leaf, both at 300 and 1000 μmol m−2 s−1 ('<em><strong>PAR300</strong></em>' and '<em><strong>PAR1000</strong></em>'), taking into account an adaptation time. We used a LED light source and several filters, placed in front of a FluoWat leaf clip, which was connected to two high-performance VIS-NIR spectroradiometers (QEPRO, Ocean Insight Inc., Orlando, Florida, USA). The spectroscopy measurements are presented in the Matlab structures for each measuring day, e.g. "<strong><em>2023_d0_Leaf_Spec_Tomato_Stress.mat</em></strong>".</p> <p>The outputs of the pigment spectral fitting code are presented by Matlab structures, e.g. "<strong><em>2023_d0_Leaf_Fitting_Tomato_Stress.mat</em></strong>", which contains the effective absorbance fitting (<strong><em>A_eff</em></strong>) of each pigment (<strong>Chl a, Chl b, Carotene-b, Anthocyanins, and Xanthophylls</strong>) for the wavelength range [500-780] nm, the absorbed photosynthetically active radiation by Chlorophyll a ('<em><strong>APAR_Chla</strong></em>') for the wavelength range [400-800] nm, and the fluorescence quantum efficiency, calculated as the ratio of the emitted fluorescence photons and the flux of photons absorbed by Chlorophyll a. </p> <p>Additional metadata from HPLC photosynthetic pigment analyses, xanthophyll-related enzyme expression, biomass and total content of elemental nitrogen are provided.</p> <p>Please follow the README files for more detailed information.</p> <p> </p>
Kleptotrace-micro-dataset
<p>This micro-benchmark dataset was made for evaluation of the proposed pipeline in Koletsis et al. Entity Extraction from High-Level Corruption<br>Schemes via Large Language Models. BDA4FCT@IEEE Big Data 2024. Also available at https://arxiv.org/abs/2409.13704</p> <p>This dataset comprises 15 articles, totaling 441 sentences, focused on topics related to financial corruption. It includes 2 lists of individuals and organizations mentioned within these articles.</p>
Improving Artificial Teachers by Considering How People Learn and Forget: Dataset
<p>This dataset contains the results of the experiment described in <a href="https://dl.acm.org/doi/10.1145/3397481.3450696">Nioche et al. (2021)</a>. </p> <p>This dataset contains 4 data files:</p> <ul> <li><em>data.csv</em>: the main data file.</li> <li><em>stimuli.csv:</em> the description/listing of the stimuli.</li> <li><em>demographic_info.csv</em>: the demographic information about the users.</li> <li><em>data_incl_preliminary_exp.csv</em>: an additional data file that includes the user of the preliminary experiments</li> </ul> <p>The main data file contains the logs of 53 different users using a self-teaching application for one week. The goal of the users was to learn the English meaning of Japanese kanji. Each user completed between 1370 trials and 1608 trials. Each user saw between 85 and 204 characters. </p> <p>Two additional files are also joint to the data files:</p> <ul> <li><em>info.ipynb</em>: A Jupyter notebook that provides information about each data file, a few descriptive plots, and an example of data manipulation.</li> <li><em>info.pdf: </em>A pdf rendering of the notebook.</li> </ul> <p>If you use this dataset, please refer to it by citing <a href="https://dl.acm.org/doi/10.1145/3397481.3450696">Nioche et al. (2021)</a>.</p>
Global dataset of nitrogen fixation rates across inland and coastal waters based on a coordinated synthesis effort
Biological nitrogen fixation converts inert di-nitrogen gas into bioavailable nitrogen and can be an important source of bioavailable nitrogen to organisms. This dataset synthesizes the aquatic nitrogen fixation rate measurements across inland and coastal waters. Data were derived from papers and datasets published by April 2022 and include rates measured using the acetylene reduction assay (ARA), 15N2 labeling, or the N2/Ar technique. The dataset is comprised of 4793 nitrogen fixation rates measurements from 267 studies, and is structured into four tables: 1) a reference table with sources from which data were extracted, 2) a rates table with nitrogen fixation rates that includes habitat, substrate, geographic coordinates, and method of measuring N2 fixation rates, 3) a table with supporting environmental and chemical data for a subset of the rate measurements when data were available, and 4) a data dictionary with definitions for each variable in each data table. This dataset was compiled and curated by the NSF-funded Aquatic Nitrogen Fixation Research Coordination Network (award number 2015825).
Dataset on tree diversity metrics and aboveground carbon storage in Southeastern U.S. oak-pine forests, 2009–2019
This dataset contains measurements of tree structural and taxonomic diversity, stand attributes, and aboveground carbon storage from mixed oak-pine forests in Florida, Georgia, and Alabama, located in the southeastern United States. Data were collected from 946 mixed oak-pine, and 7224 longleaf-slash pine and oak-pine Forest Inventory and Analysis (FIA) plots respectively spanning the years 2009 to 2019. Variables include aboveground carbon, aboveground biomass, tree density, basal area, stand age, and diversity metrics such as Shannon indices for tree species and diameter-based structural classes. Functional diversity metrics—including functional dominance and functional divergence—are also included. These data were used to support a published study examining the interactive effects of diversity metrics on carbon storage using structural equation modeling. The geographic coverage represents humid subtropical forest regions of the southeastern U.S.
Sub-Alpine Lake (>600 m) High-Frequency Water Temperature, DOC (2007-2021), and Weather Station (Fall 2023) Dataset, Maine, USA.
We collected high-frequency surface and bottom water temperature in a set of nine high-elevation lakes in Maine, USA from 2007-2021. High-elevation is defined >600m above sea level. Dissolved organic carbon concentration data for the same time period and set of lakes is modified from Nelson, S.J., R.A. Hovel, J.F. Daly, A.L. Gavin, S. Dykema, and W.H. McDowell. 2021. Northeastern Mountain Ponds Geochemistry Compilation 1978-2019 ver 1. Environmental Data Initiative. https://doi.org/10.6073/pasta/8b51d651da0e0cff8c6ad853ef69ec3b. Air temperature and precipitation data were collected from a weather station deployed in the Mountain Pond watershed in Fall 2023 to aid comparison with low and high resolution PRISM datasets.
Long-term demographic dataset for Cladonia perforata, including fine-scale cover, occupancy, and subpopulation area data, 2011-2024
This dataset includes all data pertaining to a long-term demographic study of Cladonia perforata (perforate reindeer lichen), a federally endangered lichen endemic to Florida, including fine-scale cover, occupancy, and population area data, conducted by the Archbold Biological Station Plant Ecology Program. This includes 13 years of data (2011-2024) from nine subpopulation (including seven at Archbold Biological Station, and two at the Lake Wales Ridge Wildlife and Environmental Area, Royce Unit), all located in rosemary scrub habitat within the Lake Wales Ridge metapopulation. This study sought to characterize the fire ecology and long-term population trends for the species, and thus also includes data on prescribed burn severity and time since fire. Data were collected using a stratified random plot design, with occupancy plots (presence/absence within 1.5 meter radius) throughout the subpopulation and a subset of these designated as cover plots only, with this cover data collected as point intercept hits within a 48x48cm area. Cover data also includes microhabitat data – canopy cover in densiometer reading and dominant ground cover. Cover and occupancy data were taken every 3 years for each subpopulation (subpopulations were on different yearly schedules). Subpopulation area was mapped using a submeter GPS unit every 6 years. Subpopulations were resampled for all metrics as soon as possible following a fire, and the sampling schedule was then reset.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.