Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,283
datasets available to search
ShareScore release 0.7.1
Dataset results
4,283 results for “Database”
Brazilian High School Curricula Database
<p>This is a database of high school curricula in Brazil, in PDF, preprocessed in .txt, and a Pretrained Vector Space Model based on the documents.</p>
River Sediment Database (RivSed)
<p>The River Sediment Database (RivSed) database contains surface suspended sediment concentrations (SSC) derived from Landsat 5, 7, and 8 Level 1 Collection 1 surface reflectance from all rivers in the contiguous USA that are ~60 meters wide or greater. SSC represent spatially integrated "reach" median concentrations over the footprint of NHDPlusV2 centerlines where high quality river water pixels were detected within each Landsat image from 1984-2018. This is built in the River Surface Reflectance database (RiverSR) also in Zenodo (Gardner et al,. 2020 <em>Geophysical Research Letters</em>). </p> <p>The paper associated with RivSed: <strong>Gardner, J., Pavelsky, T. M., Topp, S., Yang, X., Ross, M. R., & Cohen, S. (2023). Human activities change suspended sediment concentration along rivers. <em>Environmental Research Letters. </em><a href="https://iopscience.iop.org/article/10.1088/1748-9326/acd8d8">https://iopscience.iop.org/article/10.1088/1748-9326/acd8d8</a></strong></p> <p> </p> <p> </p> <p><strong>Files:</strong></p> <p>1) Metadata (riverSed_v1.0_metadata.pdf): Description of all data files associated with this repository. </p> <p>2) RiverSed (RiverSed_USA_v1.1.txt). Table of SSC and associated data that is joinable to nhdplusv2_modified_v1.0.shp based on the "ID" column and to the original NHDplusV2 flowlines with the "COMID" column.</p> <p>3) Shapefile of river centerlines to which the reflectance data can be attached (nhdplusv2_modified_v1.0.shp).</p> <p>4) Shapefile of the reach polygons associated with each nhdplusv2_modified reach. (nhdplusv2_polygons_v1.0.shp).</p> <p>5) The look up table for reach IDs of original (COMID) and modified (ID) NHDplusV2 centerlines. (COMID_ID.csv). Short reaches were joined together to optimize for remote sensing data collection and make more consistent reach lengths.</p> <p>6) SSC-Landsat matchup database with extended metadata on locations and in-situ data derived from Aquasat (Ross et al., 2019) (Aquasat_TSS_v1.1.csv)</p> <p>7) The final training data used to build the xgboost machine learning model (train_clean_xgb_v1.1.csv)</p> <p>8) The xgboost model that can make SSC predictions over inland waters in USA using Landsat bands/band combinations (finalmodel_xgb_v1.1.rds and .RData). The model can only be loaded in R for now.</p> <p> </p>
Integration of the Drug-Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts.
<p><strong>ABSTRACT </strong></p> <p>It contains the data of drug targets (gene names), uniprot identifiers, secondary linked data sources (e.g., PharmGKB), market drug name, chembl identifier, and pubchem compound identifier obtained from DGIdb.</p> <p><strong>Instructions: </strong></p> <p>Data were cleaned and duplicates were removed. Data were all categorical features.</p> <p><strong>Inspiration:</strong></p> <p>This dataset uploaded to U-BRITE for "DRG_DEPOT" summer 2023 team project. It is used for constructing R2G dataset, which will map drugs to their drug targets (gene -> protein = drug target)</p> <p><strong>Acknowledgements</strong></p> <p>Freshour SL, Kiwala S, Cotto KC, Coffman AC, McMichael JF, Song JJ, Griffith M, Griffith OL, Wagner AH. Integration of the Drug-Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts. Nucleic Acids Res. 2021 Jan 8;49(D1):D1144-D1151. doi: 10.1093/nar/gkaa1084. PMID: 33237278; PMCID: PMC7778926.</p> <p>https://www.dgidb.org/</p> <p><strong>U-BRITE last update date:</strong> 06/09/2023</p>
Monotonic Flexural Testing of Corroded Reinforced Concrete Beams Database
<p>The database presented here provides a collection of 804 corroded reinforced concrete beams from 54 experimental programs available in the literature. All beam specimens were tested under simply-supported monotonic three/four-point bending conditions and failed in flexure-dominated modes. The database includes 45 independent variables, 11 dependent variables, 649 corroded members, and 155 uncorroded control beams, tested across 14 countries. Of the corroded beams, 11 were naturally corroded, 30 were corroded via long-term environmental exposure (typically in the form of salt spray or fog), and 608 were corroded artificially through the impressed-current method. All observations (individual beam tests) are statistically independent, as each data entry represents one independent test. Highlighted cells indicate non-reported variables.</p> <p>This database was compiled as part of the author's Ph.D. research for the purpose of predictive machine learning. Published articles applying the database can be accessed at:<br><br><a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.dibe.2024.100527" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.dibe.2024.100527</a></p> <p>Please see the accompanying User's Manual PDF for a complete description of all nomenclature, abbreviations, assumptions, and calculations used to derive the input and response parameters.</p> <p>Please feel free to reach out to the authors if you have any queries or concerns. </p>
Database of Digital Technology for Co-creation (2DTC)
<p>This CSV file contains a list of 50 technologies commonly used in the co-creation process. The database is organised around a taxonomy for digital technology used in co-creation, developed by the Health CASCADE consortium. It can be used to select the most adapted digital technology for specific co-creation processes based on detailed functional and non-functional requirements.</p> <p>Futur development will allow taxonomy development, increase the number of classified technologies, and update the available ones.</p>
A large EEG database with users' profile information for motor imagery Brain-Computer Interface research
<p><em><strong>Context </strong></em>: <br> We share a large database containing electroencephalographic signals from 87 human participants, with more than 20,800 trials in total representing about 70 hours of recording. It was collected during brain-computer interface (BCI) experiments and organized into 3 datasets (A, B, and C) that were all recorded following the same protocol: right and left hand motor imagery (MI) tasks during one single day session.<br> It includes the performance of the associated BCI users, detailed information about the demographics, personality and cognitive user’s profile, and the experimental instructions and codes (executed in the open-source platform OpenViBE).<br> Such database could prove useful for various studies, including but not limited to: 1) studying the relationships between BCI users' profiles and their BCI performances, 2) studying how EEG signals properties varies for different users' profiles and MI tasks, 3) using the large number of participants to design cross-user BCI machine learning algorithms or 4) incorporating users' profile information into the design of EEG signal classification algorithms.<br> <br> Sixty participants (Dataset A) performed the first experiment, designed in order to investigated the impact of experimenters' and users' gender on MI-BCI user training outcomes, i.e., users performance and experience, (Pillette & al). Twenty one participants (Dataset B) performed the second one, designed to examined the relationship between users' online performance (i.e., classification accuracy) and the characteristics of the chosen user-specific Most Discriminant Frequency Band (MDFB) (Benaroch & al). The only difference between the two experiments lies in the algorithm used to select the MDFB. Dataset C contains 6 additional participants who completed one of the two experiments described above. Physiological signals were measured using a g.USBAmp (g.tec, Austria), sampled at 512 Hz, and processed online using OpenViBE 2.1.0 (Dataset A) & OpenVIBE 2.2.0 (Dataset B). For Dataset C, participants C83 and C85 were collected with OpenViBE 2.1.0 and the remaining 4 participants with OpenViBE 2.2.0. Experiments were recorded at Inria Bordeaux sud-ouest, France.</p> <p><em><strong>Duration</strong> </em>: Each participant's folder is composed of approximately 48 minutes EEG recording. Meaning six 7-minutes runs and a 6-minutes baseline.</p> <p><br> <strong><em>Documents</em></strong><em> </em><br> <em>Instructions</em>: checklist read by experimenters during the experiments.<br> <em>Questionnaires</em>: the Mental Rotation test used, the translation of 4 questionnaires, notably the Demographic and Social information, the Pre and Post-session questionnaires, and the Index of Learning style. English and french version<br> <em>Performance</em>: The online OpenViBE BCI classification performances obtained by each participant are provided for each run, as well as answers to all questionnaires<br> <em>Scenarios/scripts</em> : set of OpenViBE scenarios used to perform each of the steps of the MI-BCI protocol, e.g., acquire training data, calibrate the classifier or run the online MI-BCI</p> <p><strong><em>Database </em></strong>: raw signals<br> Dataset A : N=60 participants<br> Dataset B : N=21 participants<br> Dataset C : N=6 participants<br> <br> The article that expained the database is available here:<br> Dreyer, P., Roc, A., Pillette, L. <em>et al.</em> A large EEG database with users’ profile information for motor imagery brain-computer interface research. <em>Sci Data</em> <strong>10</strong>, 580 (2023).<br> https://doi.org/10.1038/s41597-023-02445-z<br> </p>
MASCDB, a database of images, descriptors and microphysical properties of individual snowflakes in free fall
<p><strong>Dataset overview</strong></p> <p>This dataset provides data and images of snowflakes in free fall collected with a <a href="https://amt.copernicus.org/articles/5/2625/2012/">Multi-Angle Snowflake Camera (MASC)</a> The dataset includes, for each recorded snowflakes:</p> <ol> <li>A triplet of gray-scale images corresponding to the three cameras of the MASC</li> <li>A large quantity of geometrical, textural descriptors and the pre-compiled output of published retrieval algorithms as well as basic environmental information at the location and time of each measurement.</li> </ol> <p>The pre-computed descriptors and retrievals are available either individually for each camera view or, some of them, available as descriptors of the triplet as a whole. A non exhaustive list of precomputed quantities includes for example:</p> <ul> <li>Textural and geometrical descriptors as in <a href="https://amt.copernicus.org/articles/10/1335/2017/"><em>Praz et al 2017</em></a></li> <li>Hydrometeor classification, riming degree estimation, melting identification, as in <a href="https://amt.copernicus.org/articles/10/1335/2017/"><em>Praz et al 2017</em></a></li> <li>Blowing snow identification, as in <a href="https://tc.copernicus.org/articles/14/367/2020/"><em>Schaer et al 2020 </em></a></li> <li>Mass, volume, gyration estimation<em>, as in <a href="https://amt.copernicus.org/preprints/amt-2021-176/">Leinonen et al 2021</a></em></li> </ul> <p><strong>Data format and structure</strong></p> <p>The dataset is divided into four <em>.parquet</em> file (for scalar descriptors) and a <em>Zarr</em> database (for the images). A detailed description of the data content and of the data records is available <a href="https://pymascdb.readthedocs.io/en/latest/data.html#data">here</a>.</p> <p><strong>Supporting code</strong></p> <p>A python-based API is available to manipulate, display and organize the data of our dataset. It can be found on <a href="https://github.com/ltelab/pymascdb">GitHub</a>. See also the code documentation on <a href="https://pymascdb.readthedocs.io/en/latest/index.html">ReadTheDocs</a>.</p> <p><strong>Download notes</strong></p> <ul> <li>All files available here for download should be stored in the same folder, if the python-based API is used</li> <li><em>MASCdb.zarr.zip</em> must be unzipped after download</li> </ul> <p><strong>Field campaigns</strong></p> <p>A list of campaigns included in the dataset, with a minimal description is given in the following table</p> <table> <tbody> <tr> <td><strong>Campaign_name</strong></td> <td><strong>Information</strong></td> <td> <p><strong>Shielded / Not shielded</strong></p> <p><em>DFIR = Double Fence Intercomparison Reference</em></p> </td> </tr> <tr> <td> <p><em>APRES3-2016 & APRES3-2017</em></p> </td> <td>Installed in Antarctica in the context of the APRES3 project. See for example <a href="https://essd.copernicus.org/articles/10/1605/2018/essd-10-1605-2018.html">Genthon et al, 2018</a> or <a href="https://tc.copernicus.org/articles/11/1797/2017/">Grazioli et al 2017</a></td> <td>Not shielded</td> </tr> <tr> <td><em>Davos-2015</em></td> <td>Installed in the Swiss Alps within the context of <a href="https://public.wmo.int/en/resources/meteoworld/spice-%E2%80%93-improving-snowfall-measurements">SPICE</a> (Solid Precipitation InterComparison Experiment)</td> <td>Shielded (DFIR)</td> </tr> <tr> <td><em>Davos-2019</em></td> <td>Installed in the Swiss Alps within the context of <a href="https://www.envidat.ch/group/about/raclets-field-campaign">RACLETS</a> (<em>Role of Aerosols and CLouds Enhanced by Topography on Snow</em>)</td> <td>Not shielded</td> </tr> <tr> <td><em>ICEGENESIS-2021</em></td> <td>Installed in the Swiss Jura in a MeteoSwiss ground measurement site, within the context of ICE-GENESIS. See for example <a href="https://doi.org/10.1175/BAMS-D-21-0184.1">Billault-Roux et al, 2023</a></td> <td>Not shielded</td> </tr> <tr> <td><em>ICEPOP-2018</em></td> <td>Installed in Korea, in the context of ICEPOP. See for example <a href="https://doi.org/10.5194/essd-13-417-2021">Gehring et al 2021</a>.</td> <td>Shielded (DFIR)</td> </tr> <tr> <td><em>Jura-2019 & Jura-2023</em></td> <td>Installed in the Swiss Jura within a MeteoSwiss measurement site</td> <td>Not shielded</td> </tr> <tr> <td><em>Norway-2016</em></td> <td>Installed in Norway during the High-Latitude Measurement of Snowfall (HiLaMS). See for example <a href="https://doi.org/10.1175/BAMS-D-21-0007.1">Cooper et al, 2022</a>.</td> <td>Not shielded</td> </tr> <tr> <td><em>PLATO-2019</em></td> <td>Installed in the "Davis" Antarctic base during the <a href="https://www.osti.gov/biblio/1524773">PLATO</a> field campaign</td> <td>Not shielded</td> </tr> <tr> <td><em>POPE-2020</em></td> <td>Installed in the "Princess Elizabeth Antarctica" base during the POPE campaign. See for example <a href="https://essd.copernicus.org/articles/15/1115/2023/essd-15-1115-2023.html">Ferrone et al, 2023</a>.</td> <td>Not shielded</td> </tr> <tr> <td><em>Remoray-2022</em></td> <td>Installed in the French Jura.</td> <td>Not shielded</td> </tr> <tr> <td><em>Valais-2016</em></td> <td>Installed in the Swiss Alps in a ski resort.</td> <td>Not shielded</td> </tr> <tr> <td>ISLAS-2022</td> <td>Installed in Norway during the <a href="https://www.uib.no/en/rg/meten/150202/islas2022-field-campaign">ISLAS campaign</a></td> <td>Not shielded</td> </tr> <tr> <td>Norway-2023</td> <td>Installed in Norway during the MC2-ICEPACKS campaign</td> <td>Not shielded</td> </tr> </tbody> </table> <p> </p> <p><strong>Version</strong></p> <p>1.1 - Two new campaigns ("ISLAS-2022", "Norway-2023") added.</p> <p>1.0 - Two new campaigns ("Jura-2023", "Norway-2016") added. Added references and list of campaigns.</p> <p>0.3 - a new campaign is added to the dataset ("Remoray-2022")</p> <p>0.2 - rename of variables. Variable precision (digits) standardized</p> <p>0.1 - first upload</p>
Harmonised LUCAS database classified by crop sequence type
<p>Assessing the benefits of crop diversification – a pillar of the agroecological transition – on a large scale requires a description of current crop sequences as a baseline, which is lacking at the scale of the European Union (EU). This work is based on the Harmonised LUCAS in-situ land cover and use database for field surveys from 2006 to 2018 in the European Union (doi: <a href="http://doi.org/10.2905/f85907ae-d123-471f-a44a-8cca993485a2">10.2905/f85907ae-d123-471f-a44a-8cca993485a2)</a> to fill this gap, We completed this dataset with a crop sequence type information for each point under non-perennial agricultural land cover in 2012, 2015 and 2018.</p> <p>The dataset lucas_classified.csv includes 31 159 points. Variables "point_id", "nuts0", "nuts2", "th_lat", "th_long", "LC1_2012", "LC1_2015", "LC1_2018" are inherited from the Harmonised LUCAS databse. Variables "cereals", "corn", "rapeseed", "sunflower", "pulses", "rootCrops", "forageLeg", "grassland" correspond to the temporal frequencies of respectively cereals, corn, rapeseed, sunflower, pulses, root crops, forage legumes and grassland within the 2012, 2015 and 2018 crop sequence for each point. Variable "crop_sequence_type" is the crop sequence type assigned to each point, among eight options: cereals, corn and cereals, forage legumes and cereals, pulses and cereals, rapeseed and cereals, root crops and cereals, sunflower and cereals, temporary grasslands.</p> <p>This dataset could be used to map current dominant crop sequences in the European Union, as illustrated in the map attached, and to assess the benefits of future crop diversification.</p> <p> </p>
Database of past progressive collapse of buildings
<p>Database of past progressive collapse of building structures, including contextual information, structural typology, structural geometry, failure information and consequence information.</p> <p>Supplement to the journal article entitled "Learning from the progressive collapse of buildings" (<a href="https://doi.org/10.1016/j.dibe.2023.100194">https://doi.org/10.1016/j.dibe.2023.100194</a>).</p>
LIPID MAPS® Structure Database (LMSD) formatted for MetFrag
<p>This repository contains the LIPID MAPS® Structure Database (<a href="https://www.lipidmaps.org/databases/lmsd/overview">LMSD</a>) formatted for use in <a href="https://msbi.ipb-halle.de/MetFrag/">MetFrag</a> (and other workflows).</p> <p><em>LIPID MAPS® Lipidomics Gateway is a free, comprehensive website for researchers interested in lipid biology. Use <a href="https://www.lipidmaps.org"> https://www.lipidmaps.org</a> to stay abreast of developments each month from across the field, and explore the rich information collections, tools and resources from the LIPID Metabolites And Pathways Strategy (LIPID MAPS®) Consortium. </em><br> </p> <p>The workflow used to create this file (by B. Talavera Andújar) can be found here: <a href="https://gitlab.lcsb.uni.lu/eci/simple-utilities/sdf2csv">https://gitlab.lcsb.uni.lu/eci/simple-utilities/sdf2csv</a></p> <p><strong>Reference:</strong> LMSD: LIPID MAPS® structure database, Sud M., Fahy E., Cotter D., Brown A., Dennis E., Glass C., Murphy R., Raetz C., Russell D., and Subramaniam S., Nucleic Acids Research, 2006, DOI: <a href="https://doi.org/10.1093/nar/gkl838"> 10.1093/nar/gkl838 </a></p>
The Blood Exposome Database
<p>The Blood Exposome Database catalogues the chemicals (endogenous and exogenous) that are expected and detected in human blood specimens. The database was created using a text mining approach using the <a href="https://pubchem.ncbi.nlm.nih.gov/">NCBI PubChem</a> , <a href="https://pubmed.ncbi.nlm.nih.gov/">NCBI PubMed</a> and <a href="https://www.ncbi.nlm.nih.gov/pmc/">NCBI PMC</a> databases. Chemicals that have been reported in the primary literature (original research articles) for blood specimens are included in the database. Additionally, data from biomonitoring surveys and metabolomics datasets (public available) for human blood specimens are also covered in the database. </p> <p>Citation: Barupal Dinesh, Fiehn Oliver. Generating the blood exposome database using a comprehensive text mining and database fusion approach. Environmental health perspectives. 2019 Sep 26;127(9):097008. <a href="https://ehp.niehs.nih.gov/doi/full/10.1289/EHP4713">EHP Link</a> </p>
A bite force database of 654 insect species
<p>The insect bite force database as described in Rühr et al. (<strong>accepted</strong>): A bite force database for 654 insect species. doi: <a href="https://doi.org/10.1038/s41597-023-02731-w">1038/s41597-023-02731-w</a>.</p><p>The code used to convert the raw measurements to the final database and to create all tables and figures of the original publication can be found on its <a href="https://github.com/Peter-T-Ruehr/InsectBiteForceDatabase">GitHub Page</a> (under release <a href="https://github.com/Peter-T-Ruehr/InsectBiteForceDatabase/releases/tag/v1.0.0">v1.0.0</a>).</p>
Database of Residual Stress Measurements on Hot-rolled Wide Flange Steel Cross Sections
<p><a href="https://zenodo.org/deposit/7677600#:~:text=Delete-,Data_info.csv,-md5%3A98a0f787ce2ea1b81d42ac898f6bb110">Data_info.csv</a>: Database of 'Residual Stress Measurements on Hot-rolled Wide Flange Steel Cross Sections' including cross-sectional and material characteristics as well as information relevant to ploting the residual stress distributions.</p> <p><a href="https://zenodo.org/deposit/7677600#:~:text=7%20kB-,Distributions.zip,-md5%3Af7f66a9ad607f27edde3dc7438b82ad2">Distributions.zip</a>: Residual stress distributions for the web and the flanges. To be unziped and positioned at the same location with the 'Data_info.csv', 'Processor.m' and 'QP_Coefficients' folder.</p> <p><a href="https://zenodo.org/deposit/7677600#:~:text=197%20kB-,Processor.m,-md5%3A29935e24d40cbd4398d260124ec71fa8">Processor.m</a>: MATLAB code that plots the residual stress distributions of a selected research work.</p> <p><a href="https://zenodo.org/deposit/7677600#:~:text=16%20kB-,QP_Coefficients.zip,-md5%3Aafd1a8771cf39c9c6584331d10030d96">QP_Coefficients.zip</a>: Coefficients of a proposed optimization method to fit the measured residual stresses in the web and the flanges. To be unziped and positioned at the same location with the 'Data_info.csv', 'Processor.m' and 'Distributions' folder.</p>
The Brazilian Soil Spectral Library (VIS-NIR-SWIR-MIR) Database: Open Access
<p><strong>Abstract:</strong></p> <p>NEW VERSION V.002 (Some Lat Long Coordinates added).</p> <p>Soil spectroscopy has emerged as a solution to the limitations associated with traditional soil surveying and analysis methods, addressing the challenges of time and financial resources. Analyzing the soil's spectral reflectance enables to observe the soil composition and simultaneously evaluate several attributes because the matter, when exposed to electromagnetic energy, leaves a "spectral signature" that makes such evaluations possible. The Soil Spectral Library (SSL) consolidates soil spectral patterns from a specific location, facilitating accurate modeling and reducing time, cost, chemical products, and waste in surveying and mapping processes. Therefore, an open access SSL benefits society by providing a fine collection of free data for multiple applications for both research and commercial use.</p> <p><strong>BSSL Description and Usefulness</strong></p> <p>The Brazilian Soil Spectral Library (BSSL), available at <a href="https://bibliotecaespectral.wixsite.com/english">https://bibliotecaespectral.wixsite.com/english</a>, is a comprehensive repository of soil spectral data. Coordinated by JAM Demattê and managed by the GeoCiS research group, the BSSL was initiated in 1995 and published by Demattê and collaborators in 2019. This initiative stands out due to its coverage of diverse soil types, given Brazil's significance in the agricultural and environmental domains and its status as the fifth largest territory in the world (IBGE, 2023). In addition, a Middle Infrared (MIR) dataset has been published (Mendes et al., 2022), part of which is included in this repository. The database covers 16,084 sites and includes harmonized physicochemical and spectral (Vis-NIR-SWIR and MIR range) soil data from various sources at 0-20 cm depth. All soil samples have Vis-NIR-SWIR data, but not all have MIR data.</p> <p>The BSSL provides open and free access to curated data for the scientific community and interested individuals. Unrestricted access to the BSSL supports researchers in validating their results by comparing measured data with predicted values. This initiative also facilitates the development of new models and the improvement of existing ones. Moreover, users can employ the library to test new models and extract information about previously unknown soil properties. With its extensive coverage of tropical soil classes, the BSSL is considered one of the most significant soil spectral libraries worldwide, with 42 institutions and 61 researchers participating. However, 47 collaborators from 29 institutions have authorized the data opening. Other researchers can also provide their data upon request through the coordinator of this initiative.</p> <p>The data from the BSSL project can also help wet labs to improve their analytical capabilities, contributing to developing hybrid wet soil laboratory techniques and digital soil maps while informing decision-makers in formulating conservation and land use policies. The soil's capacity for different land uses promotes soil health and sustainability.</p> <p><strong>Coverage</strong></p> <p>The BSSL data covers all regions of Brazil, including 26 states and the Federal District. It is in a <em>.xlsx</em> format and has a total size of 305 Mb. The table is structured in sheets with rows for observations, and columns, representing various soil attributes in the surface layer, from 0 to 20 cm depth. The database includes environmental and physicochemical properties (22 columns and 16,084 rows), Vis-NIR-SWIR spectral bands (2151 columns and 16,084 rows), and MIR channels (681 columns and 1783 rows). An ID unique column can merge the sheet for each attribute or spectral range.</p> <p><strong>Accessing original data source</strong></p> <p>Using these data requires their reference in any situation under copyright infringement penalty. Three mechanisms are available for users to reach the original and complete data contributors:</p> <p>a) Refer to sheet two for name and code-based searches;</p> <p>b) Visit the website <a href="https://bibliotecaespectral.wixsite.com/english/lista-de-cedentes">https://bibliotecaespectral.wixsite.com/english/lista-de-cedentes</a> or locate the contributors' list by Brazilian state;</p> <p>c) Visit the website of the Brazilian Soil Spectral Service – Braspecs <a href="http://www.besbbr.com.br/">http://www.besbbr.com.br/</a>, an online platform for soil analysis that uses part of the current SSL (Demattê et al., 2022) - It was developed and managed by GeoCiS. There, owners from all over the country can be found.</p> <p><strong>Proceeding to data analysis</strong></p> <p>We registered and organized the samples at the ESALQ/USP Soil Laboratory. Some samples arrived without preliminary data analyses, so we analyzed them for soil organic matter (SOM), granulometry, cation exchange capacity (CEC), pH in water, and the presence of Ca, Mg, and Na, following the recommendations of Donagemma et al. (2011).</p> <p>The GeoCiS research group performed spectral analyses following the procedures described by Bellinaso et al. (2010). Demattê et al. (2019) provide detailed methods for sampling, preparation, and soil analyses, including reflectance spectroscopy. Latitude and longitude data can be requested directly from the data owner. In summary, the following steps are involved in data acquisition.</p> <p>a) We subjected the soil samples to a preliminary treatment, which involved drying them in an oven at 45°C for 48 hours, grinding them, and sieving them through a 2mm mesh;</p> <p>b) We placed the samples in Petri dishes with a diameter of 9 cm and a height of 1.5 cm;</p> <p>c) We homogenized and flattened the surface of the samples to reduce the shading caused by larger particles or foreign bodies, making them ready for spectral readings;</p> <p>d) The spectral analyses took place in a darkened room to avoid interference from natural light. We used a computer to record the electromagnetic pulses through an optical fiber connected to the sensor, capturing the spectral response of the soil sample;</p> <p>e) We obtained reflectance data in the Visible-Near Infrared-Shortwave Infrared (Vis-NIR-SWIR) range using a FieldSpec 3 spectroradiometer (Analytical Spectral Devices, ASD, Boulder, CO), which operates in the spectral range from 350 to 2500 nm;</p> <p>f) The sensor had a spectral resolution of 3 nm from 350-700 nm and 10 nm from 700-2500 nm, automatically interpolated to 1 nm spectral resolution in the output data, resulting in 2151 channels (or bands); and</p> <p>g) We positioned the lamps at 90° from each other and 35 cm away from the sample, with a zenith angle of 30°.</p> <p>The sensor captured the light reflected through the fiber optic cable, which was positioned 8 cm from the sample's surface.</p> <p>We used two 50W halogen lamps as the power source for the artificial light. It's important to note that we took three readings for each sample at different positions by rotating the Petri dish by 90°.</p> <p>Each reading represents the average of 100 scans taken by the sensor. From these three readings, we calculated the final spectrum of the samples. Notably, the laboratory's equipment and procedures for soil sample spectral analyses followed the ASD's recommendations, particularly about sensor calibration using a white spectralon plate as a 100% reflectance standard.</p> <p>For the analysis in the Middle Infrared (MIR) spectral region, we followed the procedures outlined by Mendes et al. (2022). We milled the soil fraction smaller than 2 mm, sieved it to 0.149 mm, and scanned it using a Fourier Transform Infrared (FT-IR) alpha spectroradiometer (Bruker Optics Corporation, Billerica, MA 01821, USA) equipped with a DRIFT accessory.</p> <p>The spectroradiometer measured the diffuse reflectance using Fourier transformation in the spectral range from 4000 cm<sup>-1</sup> to 600 cm<sup>-1</sup>, with a resolution of 2 cm<sup>-1</sup>. We conducted these measurements in the Geotechnology Laboratory of the Department of Soil Science at Esalq-USP. We took the average of 32 successive readings to obtain a soil spectrum. Sensor calibration took place before each spectral acquisition of the sample set by standardizing it against the maximum reflectance of a gold plate.</p> <p> </p> <p><strong>Dataset characterization</strong></p> <p>The database, named BSSL_DB_Key_Soils, has five sheets containing the key soil attributes, Vis-NIR-SWIR and MIR datasets, descriptions of the contributors and the proximal sensing methods used for spectral soil analysis. The sheets can be linked by "ID_Unique" columns, which bring the corresponding rows according to the data type. Some cells are empty because collaborators have already provided data in this way. However, we have decided to keep them in the database because they have other soil key attributes. Every Column in the data sheets is described as follows:</p> <p> </p> <p><strong>Sheet 1. BSSL_Soil_Attributes_Dataset</strong></p> <p>Column 1. <strong>ID_unique</strong>: Sequential code assigned to every record;</p> <p>Column 2. <strong>Owner code</strong>: Acronym assigned to each contributor who allowed access to their proprietary data;</p> <p>Column 3. <strong>Vis_NIR_SWIR_availability</strong>: availability of spectral data in visible, near-infrared, and shortwave infrared ranges;</p> <p>Column 4. <strong>MIR_availability</strong>: availability of spectral data in the middle infrared range;</p> <p>Column 5. <strong>Sampling</strong>: type of soil sampling;</p> <p>Column 6. <strong>Depth_cm</strong>: soil surface layer depth in centimeters; </p> <p>Column 7. <strong>Lat</strong>: Latitude; </p> <p>Column 8. <strong>Lat</strong>: Longitude; </p> <p>Column 9. <strong>Region</strong>: Brazilian geographical region of samples' source;</p> <p>Column 10. <strong>Municipality</strong>: Brazilian municipality of samples' source;</p> <p>Column 11. <strong>State</strong>: Brazilian Federation Unit of samples' source;</p> <p>Column 12. <strong>Vegetation</strong>: type of vegetal covering;</p> <p>Column 13. <strong>Biome</strong>: groupings of ecosystems that share similar characteristics and span different regions;</p> <p>Column 14. <strong>Geology</strong>: type of rock matter from local soil sampling;</p> <p>Column 15. <strong>Sand_gkg</strong>: Content of the soil fraction with grain size between 2 and 0.053 mm, expressed in grams per kilogram;</p> <p>Column 16. <strong>Clay_gkg</strong>: Content of soil fraction with grain size smaller than 0.002 mm, expressed in grams per kilogram;</p> <p>Column 17. <strong>SOM_gkg</strong>: Soil organic matter content, expressed in grams per kilogram;</p> <p>Column 18. <strong>pH_H2O</strong>: Soil hydrogen ion potential measured in water;</p> <p>Column 19. <strong>Ca_mmolkg</strong>: Exchangeable calcium content in the soil, expressed in millimoles per kilogram;</p> <p>Column 20. <strong>Mg_mmolkg</strong>: Exchangeable magnesium content in the soil, expressed in millimoles per kilogram;</p> <p>Column 21. <strong>Na_mmolkg</strong>: Exchangeable sodium content in the soil, expressed in millimoles per kilogram; and</p> <p>Column 22. <strong>CEC_Ph7_mmolkg</strong>: Cation exchange capacity of the soil at neutral pH, expressed in millimoles per kilogram.</p> <p> </p> <p><strong>Sheet 2. BSSL_Vis_NIR_SWIR_Dataset</strong></p> <p>Column 1. <strong>ID_Unique</strong>: Sequential code assigned to every record;</p> <p>Column 2. <strong>Owner code</strong>: Acronym assigned to each contributor who allowed access to their proprietary data; and</p> <p>Column 3 – 2153. <strong>350 – 2500</strong>: Reflectance in 2151 spectral bands in nanometers from visible and near-infrared to shortwave infrared range (350 – 2500 nm).</p> <p> </p> <p><strong>Sheet 3. BSSL_MIR_Dataset</strong></p> <p>Column 1. <strong>ID_Unique:</strong> Sequential code assigned to every record;</p> <p>Column 2. <strong>Owner_code:</strong> Acronym assigned to each contributor who allowed access to their proprietary data; and</p> <p>Column 3 – 683. <strong>4000 – 600:</strong> Reflectance in 681 spectral bands in centimeters in the middle infrared range (4000 – 600 cm<sup>-1</sup>).</p> <p> </p> <p><strong>Sheet 4. Contributors</strong></p> <p>Column 1. <strong>Owner_code</strong>: Acronym assigned to each contributor who allowed access to their proprietary data, which identifies and links it to datasets;</p> <p>Column 2. <strong>Owner</strong>: Name of the collaborator who agreed to the availability of the data;</p> <p>Column 3. <strong>E-mail</strong>: Contact the e-mail of the owner for more information or a data request;</p> <p>Column 4. <strong>Institution</strong>: Contributor's affiliation;</p> <p>Column 5. <strong>Samples NIR</strong>: Number of Vis-NIR-SWIR samples sent to the BSSL collection;</p> <p>Column 6. <strong>Samples MIR</strong>: Number of MIR samples sent to the BSSL collection;</p> <p> </p> <p><strong>Sheet 5. Metadata</strong></p> <p>Column 1. <strong>Material and Methods</strong>: Description of procedures performed for soil data analyses</p> <p> </p> <p><strong>Expectation and Social Relevance</strong></p> <p>These data can impact various disciplines such as soil surveying, soil attribute mapping, soil analysis, soil mineralogy, soil management zones, precision agriculture, development of new datasets and scientific groups, and others. We expect this contribution to be valuable and useful to the soil research community in promoting this non-renewable natural resource's conservation and sustainable use.</p>
Database of New Phase Change Materials
<p>Database of structural and optical properties of Gallium sulfide (GaS), Antimony sulfide (Sb2S3), Gallium Selenide (GaSe), Selenium (Se), and Molybdenum oxide (MoOx) thin film phase change materials. Contains Raman spectra and room temperature optical constants from spectroscopic ellipsometry for amorphous and polycrystalline phases. Also includes temperature dependent optical constants measured in-situ during thermal annealing to induce crystallization. Thin films fabricated by chemical vapor deposition, exfoliation, and solution processing methods. Provided by the PHEMTRONICS project (Deliverable 2.7).</p>
Database of pyroclastic cover deposit thickness measurements (PT-Cam) in peri-volcanic areas of Campania (Italy)
<p>In an eruptive event, tephra deposits (i.e. ash and pumice) disperse in the atmosphere and deposit on the ground surface according to the speed and direction of the wind. Because the geotechnical and hydraulic properties of the unconsolidated pyroclastic fall deposits usually differ from the bedrock, their spatial thickness significantly influences geomorphological and hydrogeological processes such as landscape evolution, erosion, landslide, and hillslope hydrogeology.</p> <p>The PT-Cam database presents the thickness of tephra deposits (i.e. the unconsolidated materials over the bedrock) in Campania region (Italy), measured through in-situ investigations of some territories around the Somma-Vesuvius, Campi Flegrei, Roccamonfina, and Ischia volcanoes during the last decades. The measurements were conducted with probing tests, dynamic penetration tests, trenches, man-made pits, seismic surveys, and outcrops.</p> <p>Explanation for database attribute:</p> <ul> <li>CODE: identification code of the measurement;</li> <li>z: measured thickness expressed in cm;</li> <li>type_z: thickness type (i.e. if investigation has reached to the bedrock the type is “total” otherwise it is “partial”);</li> <li>type_investigation: method of in-situ investigation;</li> <li>locality: municipality to which the measurement point belongs;</li> <li>Longitude, Latitude: km coordinates in UTM WGS 84 system.</li> </ul>
The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.
<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Université Claude Bernard Lyon 1, VetAgro Sup, Research Team “Bacterial Opportunistic Pathogens and Environment” (BPOE), 69280 Marcy L’Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L’Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bonté et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer’s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O'LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project “Iouqmer” EST 2016/1/120, l'Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rhône, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: Céline COLINON, Romain MARTI, Emilie BOURGEOIS, Sébastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257–4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161–168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bonté, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153–164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>
Deliverable 2.3 Design Load Case Database for Code-to-Code Comparison
<p>In work package 2 of FLOATECH a detailed validation and verification of the capabilities of QBlade-Ocean was performed. Thereby, three wind turbine models mounted on floating substructures with differing characteristics serve as the means for the validation.This dataset contains Floating Offshore Wind Turbine (FOWT) calculations in various design situations, computed with three different codes. In more detail, three floating platform archetypes are used in the code-to-code comparison ongoing in work package 2: a semi-submersible-type floater and a spar-type floater as well as the Hexafloat® concept recently proposed by Saipem®. The three test-cases are the NREL 5MW RWT mounted on the DeepCwind semi-submersible platform, the DTU 10MW RWT mounted on the SOFTWIND spar-type platform and the DTU 10MW RWT mounted on the Hexafloat® platform. <br> More details regarding the dataset structure and the testcases can be found in the accompanying document. </p> <p> </p> <p><em>Changelog:</em> </p> <p>Version 5.0.0</p> <ul> <li>Changed structural damping ratio (increased) in OpenFAST results (SOFTWIND and Hexafloat models) to match structural damping ratios in QBlade</li> <li>Added wave time series at platform undisplaced position</li> <li>Added TurbSim input files for wind field generation. TurbSim v2.0.0 (https://www.nrel.gov/wind/nwtc/turbsim.html - National Renewable Energy Laboratory)</li> <li>Added detailed DLC and simulation information in "DLC&SimulationList.xlsx"</li> </ul> <p>Version 4.0.0</p> <ul> <li>Correction of bugs in QB 5MWOC4 and 10MWSOFT models: wind shear exponent in wind fields (0.11 -> 0.14)</li> <li>Update of hydrodynamic database used in OpenFAST and QBlade results of 10MWSOFT model </li> </ul> <p> </p> <pre><code>Data in version 5.0.0 is used in WES publication: Papi, F., Troise, G., Behrens de Luna, R., Saverin, J., Perez-Becker, S., Marten, D., Ducasse, M.-L., and Bianchini, A.: A Code-to-Code Comparison for Floating Offshore Wind Turbine Simulation in Realistic Environmental Conditions: Quantifying the Impact of Modeling Fidelity on Different Substructure Concepts, Wind Energ. Sci. Discuss. [preprint], https://doi.org/10.5194/wes-2023-107, in review, 2023.</code></pre> <p> </p>
CoCO2 global emission point source database
<p>This dataset contains a global emission catalogue of CO2 and co-emitted species (NOx, SO2, CO, CH4) from thermal power plants for the year 2018. The dataset contains annual emission information for individual thermal power plants at their exact geographical location. Each facility is linked to a specific temporal (i.e., monthly, day-of-the-week and hourly) and vertical distribution profile to derive spatial- and temporal-resolved emissions for modelling efforts. The dataset was produced as part of the CoCO2 project, which has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement No 958927.</p>
LAGOS-NE-LOCUS v1.01: A module for LAGOS-NE, a multi-scaled geospatial and temporal database of lake ecological context and water quality for thousands of U.S. Lakes: 1925-2013
This data package, LAGOS-NE-LOCUS v1.01, is 1 of 5 data packages associated with the LAGOS-NE database-- the LAke multi-scaled GeOSpatial and temporal database. Three of the data packages each contain different types of data for 51,101 lakes and reservoirs larger than 4 ha in 17 lake-rich U.S. states to support research on thousands of lakes. These three package are: (1) LAGOS-NE-LOCUS v1.01: lake location and physical characteristics for all lakes. (2) LAGOS-NEGEO v1.05: ecological context (i.e., the land use, geologic, climatic, and hydrologic setting of lakes) for all lakes. These geospatial data were created by processing national-scale and publicly-accessible datasets to quantify numerous metrics at multiple spatial resolutions. And, (3) LAGOS-NE-LIMNO v1.087.1: in-situ measurements of lake water quality from the past three decades for approximately 2,600-12,000 lakes, depending on the variable. This module was created by harmonizing 87 water quality datasets from federal, state, tribal, and non-profit agencies, university researchers, and citizen scientists. The other two data packages contain supporting data for the LAGOS-NE database: (4) LAGOS-NE-GIS v1.0: the GIS data layers for lakes, wetlands, and streams, as well as the spatial resolutions that were used to create the LAGOS-NE-GEO module. (5) LAGOS-NE-RAWDATA: the original 87 datasets of lake water quality prior to processing, the R code that converts the original data formats into LAGOS-NE data format, and the log file from this procedure to create LAGOS-NE. This latter data package supports the reproducibility of LAGOS-NE-LIMNO. The LAGOS-NE-LOCUS v1.01 module includes information on the physical location and features of all lakes > 4 ha. The information provided for this population of lakes includes: lake unique identifiers, lake area, perimeter, latitude and longitude, and the zone IDs that the lake is located within (e.g., state, county, the hydrologic unit at each level (4, 8, and 12). Citation for
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.