Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
30,948
datasets available to search
ShareScore release 0.9.0
Dataset results
30,948 results for “Profiler”
Data from 'Tracability of Forest Reproductive Material with the quality label 'Plant van Hier': A DNA database with genetic profiles of native autochthonous tree and shrub species of Flanders, Belgium'
<h2>Background</h2> <p>Indigenous trees and shrubs play an important role in multifunctional forest management. They form a significant part of the biodiversity in our forests. Forest reproductive material (FRM) of autochthonous Flemish origin is sold under the quality label ‘Plant van Hier’, a certification mark of the Agency for Nature and Forests. To ensure the provenance of the seedlings, we developed a DNA-database of genetic profiles of potential parent trees, using species-specific genetic markers. This database enables the traceability of FRM of the ‘Plant van Hier’ label throughout the entire production chain; from seed harvesting and cultivation to planting by the end user.</p> <p>This database contains the genetic profiles of almost all possible parent trees present within 27 Flemish autochthonous seed orchards of eight ecologically important tree and shrub species: <em>Carpinus betulus</em>, <em>Corylus avellana</em>, <em>Frangula alnus</em>, <em>Populus tremula</em>, <em>Sorbus aucuparia</em>, <em>Tilia cordata</em>, <em>Tilia platyphyllos,</em> and <em>Ulmus laevis</em>. The profiles were established using microsatellite markers (11 to 24 markers per species). New genetic markers were developed for <em>Carpinus betulus</em> and <em>Ulmus laevis</em>. PCR products were run on an ABI 3500 Genetic Analyser (Thermo Fisher Scientific).</p> <h2>Files</h2> <p>The files will be updated when new genotypes are added to the seed orchards. The current data files contain data from genotypes collected in the period 2018-2023. </p> <h3>Species_genotypes</h3> <p>These files contain the genetic fingerprints of the parent trees of autochthonous Flemish seed orchards. Missing data is indicated as ‘MD’. For <em>Carpinus betulus</em>, an octoploid species, the allelic phenotype is given instead of the genotype as the number of times that an allele occurs on a specific locus is not known.</p> <p>The next metadata is additionally given:<br>- Species: the Latin name of the species<br>- Seed_orchard: the name of the seed orchard in which the genotypes are located<br>- Code_seed_orchard: the code of the seed orchard in which the genotypes are located as given in the Register of Flemish Forest Reproductive Material (‘Register bosbouwkundig uitgangsmateriaal’; inbo.be)<br>- Genotype: the fieldname given to the genotype<br>- Origin: the location where the genotype was collected in Flanders, Belgium. Genotypes were collected from natural stands which are assumed to have an autochthonous origin. When the specific location is unknown, the location ‘Flanders’ is given. <br>- Year_sampled: the year in which the genotypes were sampled in the respective seed orchard for genetic analysis.</p> <h3>Species_binsets</h3> <p>These files contain the binsets and allele names that are used to score the alleles of the genotypes in the programme Geneious Prime 2019.3.2 (<a href="https://www.geneious.com">https://www.geneious.com</a>). For <em>Tilia platyphyllos </em>and <em>Tilia cordata</em>, the same binsets were used.</p>
Dataset of the paper "Control of electronic band profiles through depletion layer engineering in core-shell nanocrystals"
<p>This dataset provides the raw data of the paper "Control of electronic band profiles through depletion layer engineering in core-shell nanocrystals"</p>
TOMCAT CTM simulated ozone profiles using NRL2, SATIRE and SORCE solar fluxes
<p>Individual file contain TOMCAT CTM simulated ozone profiles from five model simulations analysed in the following publication. Briefly, </p> <p>vmro3_T2Mz_TOMCAT_A_NRL2_2005-2020.nc contain ozone profiles from the control simulation that uses ERA5 dynamical forcing fields and NRL V2 solar fluxes</p> <p>vmro3_T2Mz_TOMCAT_B_SATIRE_2005-2020.nc and vmro3_T2Mz_TOMCAT_C_SORCE_2005-2020.nc contain ozone profiles from a simulations that are similar to the control simulation but with SATIRE and SORCE solar fluxes</p> <p>vmro3_T2Mz_TOMCAT_D_SFix_2005-2020.nc has ozone profiles from simulation that is similar to the control simulation but with fixed solar fluxes, whereas vmro3_T2Mz_TOMCAT_E_DFix_2005-2020.nc also contain ozone profiles from a simulation where model uses annually repeating dynamical fields.</p> <p> </p> <p>Dhomse, S. S., Chipperfield, M. P., Feng, W., Hossaini, R., Mann, G. W., Santee, M. L., and Weber, M.: A Single-Peak-Structured Solar Cycle Signal in Stratospheric Ozone based on Microwave Limb Sounder Observations and Model Simulations, Atmos. Chem. Phys. Discuss. [preprint], https://doi.org/10.5194/acp-2021-663, in review, 2021.</p>
OMPS-NPP L2 LP USask Ozone (O3) Vertical Profile swath daily V1.1
<p>The USask OMPS-LP L2 2D Ozone v1.1 product provides ozone profile retrievals performed at the University of Saskatchewan for the central slit of the Ozone Mapping and Profiler Suite Limb Profiler (OMPS-LP) instrument on the Suomi-NPP satellite. The two-dimensional retrieval algorithm accounts for variation in the along orbital track dimension, retrieving an entire orbit simultaneously instead of treating each image independently. Ozone is retrieved from the thermal tropopause to 59 km on a 1 km grid with a vertical resolution of approximately 2 km.</p> <p>Each granule contains data from the daylight portion of each orbit measured for a full month. Spatial coverage is global (-82 to +82 degrees latitude), and there are about 14.5 orbits per day, each has typically 160 profiles with an along orbital track sampling of 125 km. The files are written using NetCDF4.</p>
Vertical profiles of urban wind speed, wind direction and turbulence measured by LiDAR on campus of University College Cork, Ireland
<p><strong>Vertical Profiles of Urban wind speed, wind direction and turbulence measured by LiDAR on campus of University College Cork, Ireland</strong></p> <p>=================================</p> <p>README version 1.3, 21/07/2022</p> <p>==================================</p> <p>Contact info:</p> <p>Paul Leahy, University College Cork</p> <p>paul.leahy@ucc.ie | +353 21 4902017</p> <p>================================</p> <p> </p> <p><strong>Contents</strong></p> <p><strong>1. Measurement location and time period</strong></p> <p><strong>2. What is measured (brief description)</strong></p> <p><strong>3. Instrumentation</strong></p> <p><strong>4. CSV file detailed descriptions</strong></p> <p>================================</p> <p> </p> <p><strong>1. Measurement location and time period: </strong></p> <p>North roof of Kane Building, University College Cork (UCC), Ireland.</p> <p>Lat 51 d 53 m 34 s N.</p> <p>Long 8 d 29 m 39 s W.</p> <p>Roof is c. 39 m above sea level, and c. 26 m above ground level (ground level reference point is the car park West of the UCC Kane Building).</p> <p>The measurements were taken over a time period of several months in the years 2013 / 2014.</p> <p>=================================</p> <p><strong>2. What is measured (brief description):</strong></p> <p>* LiDAR Wind speed (horizontal and vertical), wind direction, turbulence intensity at 5 altitudes; reference point (0 m) for these altitudes is the top of the LiDAR instrument c. 1.2 m above roof level.</p> <p>* Air temperature, atmospheric pressure, relative humidity.</p> <p>* Wind speed and direction from an ultrasonic anemometer mounted on top of the instrument (c. 1.2 m above roof level).</p> <p>* 10-minute average values (2 files) and high-resolution (c. 23 sec) data (1 file) are provided.</p> <p>See 'CSV file detailed description' below for detailed information.</p> <p>* Diagnostic information.</p> <p>=================================</p> <p><strong>2.1 Surrounding terrain:</strong></p> <p>Surrounding area is urban/suburban. The aspect is northerly.</p> <p>To the West: 2-5 storey buildings, open spaces, suburban.</p> <p>To the South: 2-3 storey buildings, open spaces, trees, river.</p> <p>To the East: 2-3 storey buildings, open spaces.</p> <p>To the North: A higher section of the Kane Building roof (47 m asl), 1-3 storey buildings, suburban.</p> <p>=================================</p> <p><strong>3. Instrumentation:</strong></p> <p>ZephIR 175 continuous wave wind profiling LiDAR with integrated sonic anemometer, temperature, humidity, air temperature pressure sensors and GPS.</p> <p>=================================</p> <p><strong>4. CSV files detailed description:</strong></p> <p><strong>4.1 Data on 10-minute averages:</strong></p> <p>Filename 05092013-03122013_10min_res.csv contains:</p> <p>10 minute averaged data from 05/09/2013 to 03/12/2013.</p> <p>Measurement altitudes: 148 m, 90 m, 69 m, 44 m, 19m above instrument level.</p> <p> </p> <p>Filename 03122013-07082014_10min_res.csv contains:</p> <p>10 minute averaged data from: 03/12/2013 to 07/08/2014.</p> <p>Measurement altitudes: 148 m, 90 m, 50 m, 35 m, 15 m above instrument level.</p> <p>Note: from 19/06/2014 onwards, LiDAR data missing (MET data continues).</p> <p> </p> <p>The first two rows contain header information.</p> <p>Row 1 contains location information (GPS record)) and the measurement altitudes for wind speeds.</p> <p>Sample GPS record: N51535775W8296590 = 51 d 53.5775 m North; 8 d 29.6590 m West.</p> <p>Row 2 contains the data column headers including units.</p> <p> </p> <p>Wind speeds at each altitude are recorded:</p> <p>No of Packets (= number of scan units averaged over) []</p> <p>Wind direction (mean) [deg]</p> <p>Horizontal wind speed (mean) & standard deviation [m/s]</p> <p>Vertical wind speed (mean) & standard deviation [m/s]</p> <p>Horizontal variance [m^2/s^2] </p> <p>Horizontal min [m/s] </p> <p>Horizontal max [m/s] </p> <p>TI (turbulence intensity) []</p> <p> </p> <p>Other meteorological data:</p> <p>Air temperature [<sup>o</sup>C]</p> <p>Pressure [mbar]</p> <p>Rel. Humidity [%]</p> <p>Rain indicator [unitless] Higher values indicate more rain during the averaging interval.</p> <p>Wind Speed [m/s] (column 'MET Wind Speed' measured at the top of the instrument by the ultrasonic anemometer)</p> <p>Wind direction [deg] (column 'MET Direction' measured at the top of the instrument by the ultrasonic anemometer).</p> <p> </p> <p>Other housekeeping and diagnostic data:</p> <p>Instrument tilt [deg]</p> <p>Instrument bearing [deg]</p> <p>GPS data [degrees N, degrees W]</p> <p>Battery voltage [V] </p> <p>Optics, electronics and battery temperature [<sup>o</sup>C]</p> <p> </p> <p>=====================================================</p> <p> </p> <p><strong>4.2 Data with high time resolution (~23 s):</strong></p> <p> </p> <p>Filename 05092013-11112013_23s_res.csv contains:</p> <p>High resolution data from 05/09/2013 to 11/11/2013</p> <p>Measurement altitudes: 148 m, 90 m, 69 m, 44 m, 19m.</p> <p> </p> <p>Note on time resolution:</p> <p>The time resolution of processed wind measurements is c. 3 seconds per wind level, and around 8 seconds to reset to the first level. A full wind profile measurement at 5 altitudes therefore takes around (5 x 3) + 8 = 23 s to complete.</p> <p>The raw scanning resolution of the instrument is higher than this, as each wind measurement is an average of several values.</p> <p> </p> <p>Row 1 contains location information (lat, long) and the vertical measurement levels for wind speeds.</p> <p>Row 2 contains the data column headers including units.</p> <p> </p> <p>Wind speeds at each altitude are recorded:</p> <p>No of Packets (= scan units averaged over) []</p> <p>Wind direction (mean) [deg]</p> <p>Horizontal wind speed (mean) & standard deviation [m/s]</p> <p>Vertical wind speed (mean) & standard deviation [m/s]</p> <p>Horizontal variance [m^2/s^2] not defined as measurement interval is too short.</p> <p>Horizontal min [m/s] not defined as measurement interval is too short. </p> <p>Horizontal max [m/s] not defined as measurement interval is too short. </p> <p>TI (turbulence intensity) [] not defined as measurement interval is too short.</p> <p> </p> <p>Other meteorological data:</p> <p>Air temperature [<sup>o</sup>C]</p> <p>Pressure [mbar]</p> <p>Rel. Humidity [%]</p> <p>Rain indicator [unitless] Higher values indicate more rain during the scanning interval.</p> <p>Wind Speed [m/s] (column 'MET Wind Speed' measured at the top of the instrument by the ultrasonic anemometer)</p> <p>Wind direction [deg] (column 'MET Direction' measured at the top of the instrument by the ultrasonic anemometer.</p> <p> </p> <p>Other housekeeping and diagnostic data:</p> <p>Instrument tilt [deg]</p> <p>Instrument bearing [deg]</p> <p>GPS data [degrees N, degrees W]</p> <p>Battery voltage [V] </p> <p>Optics, electronics and battery temperature [<sup>o</sup>C]</p> <p> </p> <p>=====================================================</p> <p><strong>4.3 Quality control indicators:</strong></p> <p> </p> <p>9998 atmospheric conditions which adversely affect LiDAR wind speed measurements e.g. fog</p> <p>9999 high quality wind speed measurement not possible e.g. very low wind speed or obscuration of optical path</p> <p>Status Flag 'Green' => good</p> <p>=======================================================</p> <p> </p>
DATASET: characterization of the seed coat extractable phenolic profile and color in 308 common bean lines of the Spanish Diversity Panel
<p>Characterizarion of the seed coat extractable phenolic profile and color in 308 common bean lines of the Spanish Diversity Panel</p>
500-meter grid of Derived Soil Profiles (DSP) for Italy - SuoliCella500
<p>National database of Italian Soil Typological Units (STU) and corresponding Derived Soil Profiles (DSP) obtained on a 500 meters grid (1,109,672 points) by neural network. The most probable WRB Reference Soil Group (RSG), WRB Qualifiers, and USDA textural soil types were mapped on the 500 meters grid, by neural network. 18,707 Observed soil profiles and the respective 33,014 Soil Horizons were grouped into 4,472 STUs based on the combinations of Soil Region, WRB Reference Soil Group (RSG), WRB Qualifiers, and USDA textural soil types obtained on the 500 meters grid. Statistics were calculated (Mean Value, Standard Deviation Value, and Numerosity) for soil rooting depth and for the most common analytical parameters of the soil horizons (Coarse fragment content fraction; pH in water; Carbon (C) - organic; Carbonate (CO3--) - Total; Clay, Sand, and Silt fraction; Granulometry; Textural soil types). The 500 meters grid adopts EPSG 23032 (ED50 UTM-32). A reference scale of 1:250.000 may be attributed to the 500-meters grid map, on the base of the numerosity of DSP produced for the whole italian territory.</p>
Moderate associations between the use of levonorgestrel-releasing intrauterine device and metabolomics profile; Supplementary Figures
<p>Supplementary Material for the article "Moderate associations between the use of levonorgestrel-releasing intrauterine device and metabolomics profile", in the Journal of Clinical Endocrinology and Metabolism</p>
Data for: Impact of SO2 injection profiles on simulated volcanic forcing for the Sarychev 2009 eruptions - investigating the importance of using high vertical resolution methods when compiling SO2 data
<p>The files are data assosicated with the study High-resolution stratospheric volcanic SO2 injections in WACCM. The files are associated with four differnt simulaions described in the paper: M16, S21-1D, S21-3D and No-Volc. The files with "input" in the name are the SO2 input files used in the WACCM (Whole Atmosphere Community Climate Model) simulations in the paper. The files with "monthly_averages" in the filenames are monthly averages of model output data the variables used in the paper. </p> <p>The CALIOP_monthly_averages.nc file is monthly average of the CALIOP (Cloud-Aerosol Lidar with Orthogonal Polarization) satellite data used in the study to evaluate the WACCM simulations. </p> <p> </p>
Shipboard Conductivity–Temperature–Depth (CTD) and dissolved oxygen profile data collected during hypoxia surveys along six hydrographic sampling lines within Olympic Coast National Marine Sanctuary, 2004–2015
<p>This data set includes Conductivity-Temperature-Depth (CTD) and dissolved oxygen profile data that were collected along Washington State’s outer coast within Olympic Coast National Marine Sanctuary (OCNMS). Measurements were made along six cross-shelf hydrographic sampling lines during a series of hypoxia survey cruises from 2004 – 2015. The 398 CTD profiles were acquired using Sea-Bird Scientific 19 SeaCAT or 19plus SeaCAT CTD profilers with associated SBE-43 (Sea-Bird Electronics) or Beckman or YSI-type (Yellow Springs Instruments) dissolved oxygen sensors. The data were processed via Sea-Bird Scientific’s SBE Data Processing application using six of the modules in the following order: Data Conversion, Filter, Align CTD, Loop Edit, Derive, and Bin Average. These processing steps and associated methods are the same as those used to process CTD data collected during OCNMS mooring maintenance cruises (<a href="https://www.sciencedirect.com/science/article/pii/S2352340924001422">Risien et al., 2024</a>) and along the Newport Hydrographic Line (<a href="https://www.sciencedirect.com/science/article/pii/S2352340922001342">Risien et al., 2022</a>) located off the central Oregon coast.</p> <table> <tbody> <tr> <td><strong>Station Name </strong></td> <td><strong>Latitude</strong></td> <td><strong>Longitude</strong></td> <td><strong>Water Depth (m, MLLW)</strong></td> </tr> <tr> <td><strong>Cape Alava (CA)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>CA010</td> <td>48.1661oN</td> <td>124.7540oW</td> <td>10</td> </tr> <tr> <td>CA020</td> <td>48.1661oN</td> <td>124.7598oW</td> <td>20</td> </tr> <tr> <td>CA030</td> <td>48.1659oN</td> <td>124.7783oW</td> <td>30</td> </tr> <tr> <td>CA040</td> <td>48.1659oN</td> <td>124.7852oW</td> <td>40</td> </tr> <tr> <td>CA045</td> <td>48.1659oN</td> <td>124.8335oW</td> <td>45</td> </tr> <tr> <td>CA050</td> <td>48.1658oN</td> <td>124.8578oW</td> <td>50</td> </tr> <tr> <td>CA060</td> <td>48.1659oN</td> <td>124.8843oW</td> <td>60</td> </tr> <tr> <td>CA070</td> <td>48.1655oN</td> <td>124.9011oW</td> <td>70</td> </tr> <tr> <td>CA080</td> <td>48.1657oN</td> <td>124.9141oW</td> <td>80</td> </tr> <tr> <td>CA090</td> <td>48.1659oN</td> <td>124.9247oW</td> <td>90</td> </tr> <tr> <td>CA100</td> <td>48.1658oN</td> <td>124.9319oW</td> <td>100</td> </tr> <tr> <td><strong>Teahwhit Head (TH)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>TH030</td> <td>47.8759oN</td> <td>124.6481oW</td> <td>30</td> </tr> <tr> <td>TH035</td> <td>47.8761oN</td> <td>124.7024oW</td> <td>35</td> </tr> <tr> <td>TH040</td> <td>47.8760oN</td> <td>124.7281oW</td> <td>40</td> </tr> <tr> <td>TH050</td> <td>47.8761oN</td> <td>124.7567oW</td> <td>50</td> </tr> <tr> <td>TH060</td> <td>47.8765oN</td> <td>124.7822oW</td> <td>60</td> </tr> <tr> <td>TH070</td> <td>47.8765oN</td> <td>124.8084oW</td> <td>70</td> </tr> <tr> <td>TH080</td> <td>47.8768oN</td> <td>124.8415oW</td> <td>80</td> </tr> <tr> <td>TH090</td> <td>47.8769oN</td> <td>124.8868oW</td> <td>90</td> </tr> <tr> <td>TH100</td> <td>47.8769oN</td> <td>124.9182oW</td> <td>100</td> </tr> <tr> <td><strong>Hoh Head (HH)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>HH025</td> <td>47.7688oN</td> <td>124.5605oW</td> <td>25</td> </tr> <tr> <td>HH042</td> <td>47.7688oN</td> <td>124.6428oW</td> <td>42</td> </tr> <tr> <td>HH065</td> <td>47.7688oN</td> <td>124.7401oW</td> <td>65</td> </tr> <tr> <td><strong>Raft River (RR)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>RR015</td> <td>47.4632oN</td> <td>124.3748oW</td> <td>15</td> </tr> <tr> <td>RR020</td> <td>47.4644oN</td> <td>124.4510oW</td> <td>20</td> </tr> <tr> <td>RR042</td> <td>47.4632oN</td> <td>124.5199oW</td> <td>42</td> </tr> <tr> <td>RR065</td> <td>47.4629oN</td> <td>124.6074oW</td> <td>65</td> </tr> <tr> <td><strong>Cape Elizabeth (CE)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>CE010</td> <td>47.3541oN</td> <td>124.3347oW</td> <td>10</td> </tr> <tr> <td>CE020</td> <td>47.354oN</td> <td>124.3608oW</td> <td>20</td> </tr> <tr> <td>CE030</td> <td>47.3538oN</td> <td>124.3913oW</td> <td>30</td> </tr> <tr> <td>CE040</td> <td>47.3534oN</td> <td>124.4678oW</td> <td>40</td> </tr> <tr> <td>CE050</td> <td>47.3532oN</td> <td>124.5064oW</td> <td>50</td> </tr> <tr> <td>CE060</td> <td>47.3529oN</td> <td>124.5510oW</td> <td>60</td> </tr> <tr> <td>CE070</td> <td>47.3528oN</td> <td>124.5823oW</td> <td>70</td> </tr> <tr> <td>CE080</td> <td>47.3527oN</td> <td>124.6158oW</td> <td>80</td> </tr> <tr> <td>CE090</td> <td>47.3526oN</td> <td>124.6491oW</td> <td>90</td> </tr> <tr> <td>CE100</td> <td>47.3522oN</td> <td>124.6754oW</td> <td>100</td> </tr> <tr> <td><strong>Moclips (MO)</strong></td> <td> </td> <td> </td> <td> </td> </tr> <tr> <td>MO010</td> <td>47.2214oN</td> <td>124.2394oW</td> <td>10</td> </tr> <tr> <td>MO015</td> <td>47.2214oN</td> <td>124.2599oW</td> <td>15</td> </tr> <tr> <td>MO020</td> <td>47.2214oN</td> <td>124.2791oW</td> <td>20</td> </tr> <tr> <td>MO030</td> <td>47.2195oN</td> <td>124.3347oW</td> <td>30</td> </tr> <tr> <td>MO042</td> <td>47.2195oN</td> <td>124.3958oW</td> <td>42</td> </tr> </tbody> </table>
Feature selection on microbial profiles of CRC samples with chopin2 (powered by hdlib)
<p>This Zenodo entry contains the result of the feature selection algorithm implemented through a backward variable elimination strategy in <a href="https://github.com/cumbof/chopin2" target="_blank" rel="noopener">chopin2</a> (powered by <a href="https://github.com/cumbof/hdlib" target="_blank" rel="noopener">hdlib</a>) applied on <a href="https://github.com/biobakery/MetaPhlAn" target="_blank" rel="noopener">MetaPhlAn3</a> microbial profiles of a public dataset of metagenomic stool samples collected from patients affected by the colorectal cancer (CRC) as well as from healthy individuals.</p> <p>Microbial profiles have been extracted through the <a href="https://bioconductor.org/packages/release/data/experiment/html/curatedMetagenomicData.html" target="_blank" rel="noopener">curatedMetagenomicData</a> package for R under the IDs <em>ThomasAM_2018a</em>, <em>ThomasAM_2018b</em>, and <em>ThomasAM_2019_a</em>.</p> <p>The feature selection algorithm is implemented as a backward variable elimination method, and it makes use of the vector-symbolic architecture described in <a href="https://doi.org/10.3390/a13090233" target="_blank" rel="noopener">Cumbo F 2020</a>.</p> <p>Deposited data is described below:</p> <ul> <li><em>datasets.tar.gz</em>: it contains the datasets used as input of <em>chopin2</em> as the result of merging the three datasets with relative abundances mentioned above, also stratified by age and sex (with prefix RA). The same datasets have been also binarized (with prefix BIN);</li> <li><em>hd-models.tar.gz</em>: it contains the output of the feature selection performed with <em>chopin2</em> (powered by <em>hdlib</em>) on the datasets with both relative abundance and binary profiles (RA and BIN);</li> <li><em>ml-models.tar.gz</em>: it contains the result of the feature selection produced with classical wrapper-based techniques (i.e., Random Forest, Decision Tree, Support Vector Machine, Logistic Regression, and Extreme Gradient Boosting) in addition to a Python 3.8 script to reproduce the results.</li> </ul> <p>Please note that the datasets <em>RA__ThomasAM__species.csv</em> and <em>BIN__ThomasAM__species.csv</em> are also included into the <em>datasets.tar.gz</em> archive.</p>
Dataset collected by the s-Nautilus profiler at "La Isleta" yacht club in the Mar Menor
<p>This dataset presents a collection of environmental measurements taken by the s-Nautilus profiler at "La Isleta" yacht club in the Mar Menor. The recorded variables include:</p> <ul> <li><strong>Time (yyyy-mm-dd hh:mm:ss)</strong>: The timestamp indicating the date and time of each measurement.</li> <li><strong>Depth (m)</strong>: The depth at which the measurements were taken.</li> <li><strong>Dissolved Oxygen (DO) </strong>: The concentration of dissolved oxygen, essential for assessing water quality and the health of the marine ecosystem.</li> <li><strong>Electrical Conductivity (EC)</strong>: A parameter indicating the ability of the water to conduct electricity, directly related to salinity and the presence of ions.</li> <li><strong>Temperature (T) (ºC)</strong>: The water temperature, a critical factor influencing many chemical and biological processes.</li> <li><strong>Battery Voltage (Batt3)</strong>: The voltage of the batteries powering the electronics of the profiler, providing insight into its autonomy.</li> </ul> <p>Measurements were taken every 6 hours during the ascent of the profiler.</p>
SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages
<p>Supplementary information relating to the manuscript titled "SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages" that has been published on the preprint server arXiv.</p> <ul> <li>PersistentInfectionScore.nb: Mathematica code used to process the data and generate Figure 2</li> <li>PersistentInfectionScore.pdf: pdf version of the above file</li> <li>Supplementary_tables_Harari_et_al_2022.xlsx: raw data from (<a href="https://www.nature.com/articles/s41591-022-01882-4#Sec19">Harari et al. 2022</a>) that was used to generate mutation distributions</li> </ul>
Multi-Profile Ultra High Definition (UHD) AVC and HEVC 4K DASH Datasets
<p>We present a Multi-Profile Ultra High Definition (¥emph{UHD}) DASH dataset composed of both AVC (H.264) and HEVC (H.265) video content, generated from three well known open-source 4K video clips. The representation rates and resolutions of our dataset range from 40Mbps in 4K down to 235kbps in 320x240, and are comparable to rates utilised by on demand services such as Netflix, Youtube and Amazon Prime. We provide our dataset for both real-time testbed evaluation and trace-based simulation. The real-time testbed content provides a means of evaluating DASH adaptation techniques on physical hardware, while our trace-based content offers simulation over frameworks such as ns-2 and ns-3. We also provide the original pre-DASH MP4 files and our associated DASH generation scripts, so as to provide researchers with a mechanism to create their own DASH profile content locally. Which improves the reproducibility of results and remove re-buffering issues caused by delay/jitter/losses in the Internet.<br> <br> The primary goal of our dataset is to provide the wide range of video content required for validating DASH Quality of Experience (QoE) delivery over networks, ranging from constrained cellular and satellite systems to future high speed architectures such as the proposed 5G mmwave technology.</p>
Frictionless Tabular Data Package for GC-MS Rose scent profile data for Data published in Nature genetics, June, 2018 & Science, July 2015
<p>This dataset, in the form of a Frictionless Tabular Data Package (<a href="https://frictionlessdata.io/specs/tabular-data-package/">https://frictionlessdata.io/specs/tabular-data-package/)</a>, holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxonomy Identifiers) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable <a href="https://github.com/ISA-tools/stato">STATO</a> terms. </p> <p>The data were extracted from:</p> <ul> <li>a supplementary material table, available from <a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip">https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip</a> and published alongside the Nature Genetics manuscript identified by the following doi: <a href="https://doi.org/10.1038/s41588-018-0110-3">https://doi.org/10.1038/s41588-018-0110-3</a>, published in June 2018</li> <li>a supplementary material table available as a pdf from "Biosynthesis of monoterpene scent compounds in roses" by Magnard et al, Science 03 Jul 2015 identified by the following doi: <a href="https://doi.org/10.1126/science.aab0696">https://doi.org/10.1126/science.aab0696</a></li> </ul> <p>This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR) and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.</p> <p>It is associated to the following project: <a href="https://github.com/proccaserra/rose2018ng-notebook">https://github.com/proccaserra/rose2018ng-notebook</a> with all the necessary information, executable code and tutorials in the form of Jupyter notebooks.</p> <p> </p>
PS116-Lidar_Data_Profiles
<p>The lidar profiles are retrieved with Klett or Raman method. Averaged periods are determined with taking into account of the continous cloud-free profiles.</p> <p>Each retrieving result consists of one *.txt and one corresponding *-info.txt file. The filename is structured as {instrument}_{date}_{starttime}-{endtime}-{smooth window}. (UTC is used as the time standard for all the analysis.)</p> <p>*.txt contains the backscatter (extinction) coefficient and some other related results. The *-info.txt contains the retrieving configuations, like retrieving method, reference height and so on.</p>
Comparative profiling of skeletal muscle models reveals heterogeneity of transcriptome and metabolism
<p>This dataset is a complement to the following publication: Ahmed M. Abdelmoez, Laura Sardón Puig, Jonathon AB. Smith, Brendan M. Gabriel, Mladen Savikj, Lucile Dollet, Alexander V. Chibalin, Anna Krook, Juleen R. Zierath, and Nicolas J. Pillon. <a href="https://doi.org/10.1152/ajpcell.00540.2019">Comparative profiling of skeletal muscle models reveals heterogeneity of transcriptome and metabolism. </a>Am J Physiol Cell Physiol. 2019 Dec 11.</p> <p>METHODS: Publicly available data from myotubes and skeletal muscle tissues were selected from the GEO database. Raw files were downloaded and robust multi array (RMA) normalization was performed in unison for all samples from the same platform. For each human ENSEMBL, the rat and mouse orthologs were found using the R package BioMart and the arrays were merged based on the human ENSEMBL annotation. The database was then aggregated according to the official human gene symbol. When multiple ENSEMBL were found for a single gene symbol, an average was calculated.</p>
GitHub Profiles (users/organisations) and Repositories (research/non-research) of Potsdam Researchers and Research Organisations: An annotated dataset of with howfairis and software quality variables.
<p>This dataset accompanies the paper <em>"Software FAIRness, Documentation and Development Practices in Potsdam Researchers' GitHub Repositories"</em> It includes 3 CSV files that contain data related to github profiles of users/organisations, their repositories annotated as research/non-research repositories and followed by FAIRness and other software qualtiy variables. The data were collected using <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP">SWORDS-template-UP</a> (v1.0.0) methods (collect_users, collect_repositories, collect_variables) which is extended version of <a href="https://github.com/UtrechtUniversity/SWORDS-template">SWORS-template</a> adopted according our needs and detailed in the paper.</p> <p><strong>GitHub (research) user/organisation profiles. ( <em>github_profiles.csv )</em></strong></p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>user_id</td> <td>GitHub username </td> </tr> <tr> <td>html_url </td> <td>URL of the GitHub profile </td> </tr> <tr> <td>type </td> <td>Type of profile (user or organization)</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the organization </td> </tr> </tbody> </table> <p><strong>GitHub repositories <em>(github_repositories.csv)</em></strong></p> <p>This file contains the repositories scraped from the GitHub profiles of research users and organizations.</p> <table> <tbody> <tr> <td><strong>Column name </strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>html_url </td> <td>URL link to the repository </td> </tr> <tr> <td>description</td> <td>GitHub project description </td> </tr> <tr> <td>project</td> <td>Specifies if the project is research or non-research</td> </tr> <tr> <td>language</td> <td>Programming language used in the project </td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the university, institution, or research organization</td> </tr> <tr> <td>research_group</td> <td>Acronym or name of the research group the repository belongs to</td> </tr> </tbody> </table> <p><strong>Research repositories filtered and annotated <em>(github_research_repositories_filtered_annotated.csv)</em></strong></p> <p>This file contains filtered and annotated information about research repositories.</p> <table> <tbody> <tr> <td><strong>Column Name </strong></td> <td><strong>Description </strong></td> <td><strong>Collection Method </strong></td> </tr> <tr> <td>html_url </td> <td>Repository URL </td> <td> </td> </tr> <tr> <td>howfairis_repository</td> <td>Indicates if the repository is public or private (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_license </td> <td>Indicates if the repository has a license (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_registry</td> <td>Indicates if the repository has implemented community registry (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_citation</td> <td>Indicates if the repository has a .cff file (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_checklist</td> <td>Indicates if the repository has implemented OpenSSF best practices badge (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>fair_score</td> <td>Score based on howfairis variables (0-5) </td> <td> </td> </tr> <tr> <td>dlr_soft_class</td> <td>Name of the university, company, research institute, or research organization</td> <td>(Manual) Annotated the repository based on <a href="https://core.ac.uk/reader/211557820">DLR software engineering guideline.</a> There are no specific definitions on metrics how to categorise them (github repositories) into application classes. Which were needed to do a comparitive analysis. </td> </tr> <tr> <td>installation_instruction</td> <td>Presence of installation instruction (True/False) </td> <td>(Manual) Checked the presense of Installation Instruction in the readme or in the project wiki pages. </td> </tr> <tr> <td>project_information </td> <td>Presence of basic project information in README (True/False) </td> <td>(Manual) Checked if the readme have basic information about the project. </td> </tr> <tr> <td>usage_guide</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td>(Manual) Checked the presense of Usage Guide in the readme or in the project wiki pages. For command line tools checked if they have help command which guides how to use the tool. </td> </tr> <tr> <td>test_folder</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td> <p>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/docs/collect_variables/scripts/soft_dev_pract/test_folder.py">test_folder.py</a>) Checks the folder names test/tests in the root directory of the repository.</p> </td> </tr> <tr> <td>requirements_explicit </td> <td>Explicit requirements for Python, R, C++ repositories (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/requirement_explicit.py">requirement_explicit.py</a>) Checks the files (requirements.txt, DESCRIPTION, CMakeLists.txt) in the root directory. </td> </tr> <tr> <td>continuous_integration</td> <td>Indicates if the repository uses continuous integration (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>ci_tool </td> <td>Name of the continuous integration tool used</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>add_lint_rule </td> <td>Indicates if additional linting rules are present (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (linters) Python, R, and C++.</td> </tr> <tr> <td>add_test_rule</td> <td>Indicates if additional testing rules are present (True/False) </td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (testing libraries) Python, R, and C++.</td> </tr> <tr> <td>comment_at_start</td> <td>Indicates the level of comments at the start of the program (most, more, some, less)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/comment_at_start.py">comment_at_start.py</a>) Checks the presence of brief comments at the start at source code files in GitHub repositories.</td> </tr> <tr> <td>language </td> <td>Programming language used in the repository </td> <td> </td> </tr> <tr> <td>type </td> <td>Specifies if the profile is a user or organization </td> <td>Github organisation or user profiles.</td> </tr> <tr> <td>organisation </td> <td>Name of the university, company, research institute, or research organization</td> <td>Oraganisation name (from where the user was found)</td> </tr> <tr> <td>research_group</td> <td>Name or acronym of the research group </td> <td> </td> </tr> </tbody> </table> <p> </p> <p>Data for publication - https://github.com/Software-Engineering-Group-UP/potsdam-research-repos</p>
ChSPD Chilean soil profile database V2
<p>ChSPD is a soil profile database for Chile. The data was compiled from different published and unpublished sources. This new soil database covers a wide range of ecosystems and climate conditions. It comprises 20 different soil physical, hydraulic, and chemical properties. Each soil property has its own number of observations, which is determined by the soil horizons surveyed and the measurements taken at each point. The ChSPD_V2 includes 19769 georeferenced records, which represent 14029 soil profiles. The properties with the most records are organic matter (15797 data points), texture distribution (clay, sand, and silt content, 4978 data points), bulk density (5088 data points), field capacity (2020), and permanent wilting point (2012). </p>
Backpain exercise therapy remodels human epigenetic profiles in buccal and human peripheral blood mononuclear cells: An exploratory study in young male participants
<pre><strong>###### Files description #####</strong><br> <strong>Notes</strong>. 1) "BT" refers to before therapy and "AT" to after therapy. 2) 0 refers to FALSE and 1 to TRUE for binary variables. The provided files have tab-separated columns except the .RDS which is and R output of the mixOmics DIABLO integration analysis. <strong># Questionnaire</strong> > participants_categories.tsv: per participant (rows), output of the clustering with the participant ("ID") category ("category") per class<br> ("class") > questionnaire_agility_metrics.tsv: questionnaire and agility metrics per participant (rows) for the participants ("ID") with at least one paired AT+BT data in one type of biological sample (indicated in the columns "swab", "PBMC", and "plasma") <strong># PTMs</strong> Samples´ names are encoded as PBMC_AT_8_batch1, i.e. cells origin_time upon therapy_ID_batch (we removed _batch column suffix for the <br>processed files). NA indicates an undetected intensity. > raw_PBMC_light_labelled_intensities.tsv: raw intensity of light/endogenous peptides (row) by precursor per sample (column) from PBMC > raw_swab_light_labelled_intensities.tsv: idem from buccal cells > raw_PBMC_heavy_labelled_intensities.tsv: raw intensity of light/endogenous peptides (row) by precursor per sample (column) from PBMC > raw_swab_heavy_labelled_intensities.tsv: idem from buccal cells > raw_PBMC_heavynormalized_intensities.tsv: raw intensity of light peptides normalized by heavy peptides intensity (row) by precursor per <br>sample (column) > raw_swab_heavynormalized_labelled_intensities.tsv: idem from buccal cells > processed_cleaned_PBMC_log2intensities.tsv: processed (heavy normalized, imputed, batch-corrected) intensity of peptides aggregated by modification (PTM, row) by precursor per sample (column) after log2-transformation. The relative abundances are computed from this file. Rows without me/ac suffix represents the amount of unmodified peptide for the considered site. > processed_cleaned_swab_log2intensities.tsv: idem from buccal cells > rel_abundance_PTM_PBMC.tsv: relative abundance computed per precursor, e.g. for a given sample, the H3_K4+H3_K4me1+H3_K4me2+H3_K4me3 <br>relative abundance values must sum to 100, with the relative abundance of H3_K4 representing the absence of modified K4. > rel_abundance_PTM_swab.tsv: idem from buccal cells > tests_from_rel_abundance_PTM_swab_PBMC.tsv: per type of samples ("Sample.origin", i.e.swab of PBMC) and per PTM (rows, "PTM"), report <br>the output of classic (p-values, adjusted with Benjamini-Hochberg (BH), or Benjamini-Yekutieli procedure (BY), from raw and arcsin square <br>root transformed percentage) and PLS-DA tests (VIP - Variable Importance score - and its 95% confidence interval). The percentage of change<br>of each PTM after therapy relative tobefore therapy is reported in "perc_change.AT.over.BT" column. The "is_candidate" indicates if the PTM has been considered as a hit in the swab or PBMC. <strong># Plasma</strong> Samples´ names are encoded as PLASMA_AT_8_batch1, i.e. cells origin_time upon therapy_ID_batch. NA indicates an undetected intensity. > raw_plasma_maxquant_log2ibaq_intensities.tsv: raw data from protein group MaxQuant file. The iBAQ columns are used in later steps. > processed_cleaned_plasma_log2intensities.tsv: processed (imputed, batch-corrected) intensity of protein groups after log2-transformation. > tests_from_intens_plasma.tsv: per protein group ("Proteins.ID"), report the output of classic (p-values, adjusted Benjamini-Hochberg (BH),<br>or Benjamini-Yekutieli procedure (BY), from log2-transformed intensities) and PLS-DA tests (VIP and its 95% confidence interval). The log2 <br>fold change after therapy relative to before therapy is reported in "log2FC.AT.over.BT" column. The "is_candidate" indicates if the protein group has been considered as a hit. <strong># Integration</strong> > circos_input: output of DIABLO analysis with correlation threshold set to 0.7. Use the readRDS R function to open.</pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.