Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,481 results for “data processing”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data Associated with Chemical Cartography with APOGEE: Two-process Parameters and Residual Abundances for 288,789 Stars from Data Release 17

<p>Stellar abundance measurements are subject to systematic errors that induce extra scatter and artificial correlations in elemental abundance patterns. &nbsp;We derive empirical calibration offsets to remove systematic trends with surface gravity log(g) in 17 elemental abundances of 288,789 evolved stars from the SDSS APOGEE survey. &nbsp;We fit these corrected abundances as the sum of a prompt process tracing core-collapse supernovae and a delayed process tracing Type Ia supernovae, thus recasting each star's measurements into the amplitudes A_cc and A_Ia and the element-by-element residuals from this two-parameter fit. Here we present the log(g)-calibrated abundances, fit parameters, process amplitudes, and element-by-element abundance residuals of 288,789 stars (310,427 spectra) in APOGEE DR17 that accompany <a href="https://arxiv.org/abs/2403.08067" target="_blank" rel="noopener">the paper</a>.</p> <p>calibration_values_final.dat contains all derived calibration offsets, including the grids of log(g) calibration offsets and zero-point offsets for two-process model analysis. The first five rows of this catalog are reproduced in Table 2 of the paper.</p> <p>logg_calib_example.ipynb is a Jupyter notebook containing Python code to load calibration_values_final.dat, extract the log(g) calibration offsets for specific element, and apply calibration offsets to 10 sample stars.</p> <p>2process_residual_abund_catalog_final.fits is the catalog of 310,427 APOGEE DR17 spectra (288,789 unique stars) containing calibrated abundances, two-process fit parameters, and abundance residuals. A full listing of columns in this catalog is given in Table 5 of the paper.</p> <p>catalog_examples.ipynb is a Jupyter notebook containing Python code to load 2process_residual_abund_catalog_final.fits, cross match with other catalogs (using AstroNN and the APOGEE DR17 Globular Cluster Value-Added Catalog as examples), and make some example plots utilizing the cross-matched data.</p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Processed data and code for manuscript "Non-negligible impact of Stokes drift and wave-driven Eulerian currents on simulated surface particle dispersal in the Mediterranean Sea"

<p>This repository contains the python code and processed data to reproduce analysis and figures from R&uuml;hs et al. (2024, Ocean Science): "Non-negligible impact of Stokes drift and wave-driven Eulerian currents on simulated surface particle dispersal in the Mediterranean Sea".</p> <p>To reproduce the whole analysis, including the calculations of the trajectories, the following needs to be downloaded/included into a local working directory:</p> <ul> <li>the content of this repository in respective sub-directories, i.e. code (created and maintained at <a href="https://github.com/sruehs/RuehsEtAl2024_ImpactWavesSurfaceDispersal">https://github.com/sruehs/RuehsEtAl2024_ImpactWavesSurfaceDispersal</a>), data-proc, figs</li> <li>the original surface velocity data, to be downloaded here:&nbsp;<a href="https://zenodo.org/records/10879702">https://zenodo.org/records/10879702</a>, in an additional sub-directory named data-orig</li> </ul> <p>Additionally, the OceanParcels package, available via <a href="https://github.com/OceanParcels/parcels">https://github.com/OceanParcels/parcels</a> or <a href="https://anaconda.org/conda-forge/parcels">https://anaconda.org/conda-forge/parcels</a> needs to be installed in the python working environment. Then, the scripts in the code directory can be executed to re-run the trajectory simulations and analysis. Alternatively, the output in forms of figures and processed data can be accesed directly in the respective sub-directories.</p>

openmit-licenseNov 2024View details →
zenodo48/100

Global Ocean Heat Content Anomalies and Ocean Heat Uptake based on mapping Argo data using local Gaussian processes

<p>Monthly Ocean Heat Content Anomalies (OHCA) in the top 2000 dbar of the ocean are calculated (during 2004-2024, equatorward of 65 degree latitude) subtracting the mean over the period 2004-2024 from the monthly time series of OHC. Yearly OHCA time series are then calculated that include 1. one point per year, i.e., from averaging Jan to Dec (see files ending in &ldquo;yearly.nc&rdquo;), and 2. two points per year, i.e., from averaging Jan to Dec and Jul to Jun, respectively&nbsp; (see files ending in &ldquo;yearly2.nc&rdquo;). OHC fields are mapped using locally stationary Gaussian processes (defined over space and time) with data-driven decorrelation scales (Kuusela and Stein, 2018). A linear time trend was included in the estimate of the mean field (along with spatial terms and harmonics for the annual cycle). Mapping is done separately for different vertical sections: 15-20 dbar, 15-300 dbar, 300-700 dbar, 700-1850 dbar, 1800-1850 dbar. The 15-20 dbar (1800-1850 dbar) section is used to estimate OHCA for 0-15 dbar (1850-2000 dbar), where observations are sparser. Different vertical sections are combined to estimate global OHCA time series for 0-2000 dbar, 0-700 dbar, 700-2000 dbar (as indicated in the file names). The attribute "area" is included in the netcdf files and it tells the corresponding surface area for the estimates. Regions of the ocean that are shallower than 300 m or are not sufficiently well sampled by the Argo array are not included. Maps of the ocean masks used for the different vertical sections can be found in the .png files (blue shading indicates the area used for the horizontal integral); the bathymetry mask by Roemmich and Gilson (included in the file RG_ArgoClim_Temperature_2019.nc at https://sio-argo.ucsd.edu/RG_Climatology.html) is also used to define the ocean mask. Ocean Heat Uptake is calculated from the monthly OHCA and then averaged as described above to produce yearly time series included in the files for the different layers.</p> <p>For the uncertainty at each time point, the standard deviation of each OHCA/OHU value in the time series is included. When plotting a time series, the user may consider, e.g., shading plus/minus 1* or 1.96*standard deviation (corresponding to a&nbsp; confidence level of 68% or 95% respectively). These standard deviations in the files are estimated using spatially and temporally dependent conditional simulations of monthly gridded anomalies. When combining different layers, the standard deviation of the sum is conservatively estimated as the sum of the standard deviations.&nbsp;</p> <p>Finally, OHCA/OHU trends are estimated via a least-squares fit and reported in the variable metadata with uncertainties (confidence level of 68%). Trend uncertainties are estimated by repeating the fit for each member of the conditional simulation ensemble described above.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo48/100

UK creators' earnings survey 2020 processed data

<p>We have started to retrospectively harmonize the Music Creators Earnigs survey for the the Digital Music Observatory. The survey&rsquo;s raw data is accessible on the website of the UKIPO <a href="https://www.gov.uk/government/publications/music-creators-earnings-in-the-digital-era">here</a>.</p> <p>Ex post harmonization will be limited, because of the following factors:</p> <ul> <li>The MCE survey did not use harmonized questions in many cases</li> <li>The MCE surveys answers do not cover a full range of possible answers</li> <li>The MCE survey does not appear to represent the UK artists and music professionals.</li> </ul> <p>Because of the bias of the survey, we did not include statistical indicators of the survey yet in our observatory, and we will make further processing steps in later versions of the data file.</p> <p>Nevertheless, because of the relatively large sample size (n=708) we believe that imporant comparisons can be made with our CEEMID surveys, and we can shed some light on the earnings distribution of UK artists, and the way they distribute and finance their recordings.</p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

Example of datasets processed to demonstrate a multisource data integration methodology

<p>This dataset contains the&nbsp;data processed to demonstrate the multi-source spatial data&nbsp;integration methodology proposed in the paper &quot;Multisource spatial data integration for use cases applications&quot;.</p> <p>It contains:</p> <p>- the building footprint extracted from the IFC model of a newly designed&nbsp;building in WKT format, by using the GeoBIM_Tool (<a href="https://github.com/twut/GEOBIM_Tool">https://github.com/twut/GEOBIM_Tool</a>);</p> <p>- the extrusion of the footprint until the measured height measured with the same GeoBIM_Tool;</p> <p>- a portion of the Rotterdam 3D city model generated with 3dfier and available at&nbsp;https://3d.bk.tudelft.nl/opendata/3dfier/, converted in CityJSON&nbsp;with the citygml-tools (https://www.cityjson.org/tutorials/conversion/),&nbsp;developed to convert data between CityGML and CityJSON.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Data for paper "Convolutional neural network-based statistical post-processing of ensemble precipitation forecasts"

<p>The forecasts and observation datasets are used in the paper &quot;Convolutional neural network-based statistical post-processing of ensemble precipitation forecasts&quot;.&nbsp;https://doi.org/10.1016/j.jhydrol.2021.127301</p> <p>The forecast&nbsp;data is a subset of the &quot;ensemble for machine learning dataset (ENS4ML)&quot; from ECMWF.&nbsp;</p> <p>The Python codes are stored in Github: https://github.com/wentao-bnu/LeNet_CSG_Precip</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Modified WRF/Chem source code, output data, and post-processing scripts for the GMD manuscript "Evaluation of WRF/Chem model (v3.9.1.1) real-time air quality forecasts over the Eastern Mediterranean"

<p>Here you will find the modified WRF/Chem code used in the simulations, the scripts used for post-processing and the model output data used in the manuscript.&nbsp;</p> <p>Two modifications have been made in&nbsp;module_aerosols_soa_vbs.F:</p> <ol> <li>ch_dust&nbsp;is set to1.0D-9*0.36</li> <li>The model is set not to initialize during restarts</li> </ol> <p>The model data directory includes:</p> <ol> <li>Two csv files (Winter and Summer) with the hourly concentrations of atmospheric pollutants&nbsp;at the locations of the ground stations. These data were used to produce Figures 4-8 in the manuscript as well as all the metrics.</li> <li>Two netcdf files&nbsp;(Winter and Summer) with the average ground concentrations of atmospheric pollutants over Cyprus. These data were use to produce Figure 3 in the manuscript.&nbsp;</li> </ol>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Pre-processed data of atlas in EUCP-WP2

<p>Outputs from the probabilistic projection methods developed or assessed in the European Climate Projection system (EUCP) Horizon2020 project. The data can be previewed through our <a href="\&quot;https://eucp-project.github.io/atlas/\&quot;">interactive atlas</a>.</p> <p>&nbsp;</p> <p>For more information, see the <a href="\&quot;https://eucp-project.github.io/atlas/about\&quot;">atlas about page</a>, or the corresponding <a href="\&quot;https://eucp-project.github.io/storyboards/atlas/1_intro\&quot;">storyboard</a>.</p> <p>&nbsp;</p> <p><strong>Preprocessed data of Atlas in EUCP-WP2</strong></p> <p>We provide some notebooks that check the original/raw data, fix/add the metadata using <a href="\&quot;https://cfconventions.org/Data/cf-conventions/cf-conventions-1.9/cf-conventions.html\&quot;">CF-conventions</a> and save data in a NetCDF format. See <a href="\&quot;https://github.com/eucp-project/atlas/blob/main/python/README.md\&quot;">https://github.com/eucp-project/atlas/blob/main/python/README.md</a>.</p> <p>For two of the methods, REA and ClimWIP, pre-calculated weights have also been included. Note that these weights are only valid in the context of this specific model ensemble. Therefore, the original (pre-processed) model data is published together with the weights.</p> <p>The pre-processed data follows the following standards:</p> <p><strong>coordinates</strong></p> <ul> <li>climatology_bounds (climatology_bounds) datetime64[ns] [&#39;2050-06-01&#39;, &#39;2050-09-01&#39;, &#39;2050-12-01&#39;, &#39;2051-03-01&#39;]</li> <li>time (time) (datetime64[ns]) [2050-07-16 2051-01-16] # &quot;JJA&quot;, &quot;DJF&quot;</li> <li>latitude (lat) (float64) [30, ..., 75]</li> <li>longitude (lon) (float64) [-10, ..., 40]</li> <li>percentile (percentile) (int64) [10, 25, 50, 75, 90]</li> </ul> <p><strong>variables</strong></p> <ul> <li>tas (time, latitude, longitude, percentile) (float64)</li> <li>pr (time, latitude, longitude, percentile) (float64)</li> </ul> <p><strong>attributes</strong></p> <p>The attributes of variables and coordinates are defined as:</p> <ul> <li>&quot;tas&quot;: {<br> &quot;description&quot;: &quot;Change in Air Temperature&quot;,<br> &quot;standard_name&quot;: &quot;Change in Air Temperature&quot;,<br> &quot;long_name&quot;: &quot;Change in Near-Surface Air Temperature&quot;,<br> &quot;units&quot;: &quot;K&quot;,<br> &quot;cell_methods&quot;: &quot;time: mean changes over 20 years 2041-2060 vs 1995-2014&quot;,<br> },</li> <li>&quot;pr&quot;: {<br> &quot;description&quot;: &quot;Relative precipitation&quot;,<br> &quot;standard_name&quot;: &quot;Relative precipitation&quot;,<br> &quot;long_name&quot;: &quot;Relative precipitation&quot;,<br> &quot;units&quot;: &quot;%&quot;,<br> &quot;cell_methods&quot;: &quot;time: mean changes over 20 years 2041-2060 vs 1995-2014&quot;,<br> },</li> <li>&quot;latitude&quot;: {&quot;units&quot;: &quot;degrees_north&quot;, &quot;long_name&quot;: &quot;latitude&quot;, &quot;axis&quot;: &quot;Y&quot;},</li> <li>&quot;longitude&quot;: {&quot;units&quot;: &quot;degrees_east&quot;, &quot;long_name&quot;: &quot;longitude&quot;, &quot;axis&quot;: &quot;X&quot;},</li> <li>&quot;time&quot;: {<br> &quot;climatology&quot;: &quot;climatology_bounds&quot;,<br> &quot;long_name&quot;: &quot;time&quot;,<br> &quot;axis&quot;: &quot;T&quot;,<br> &quot;climatology_bounds&quot;: [&quot;2050-6-1&quot;, &quot;2050-9-1&quot;, &quot;2050-12-1&quot;, &quot;2051-3-1&quot;],<br> &quot;description&quot;: &quot;mean changes over 20 years 2041-2060 vs 1995-2014. The mid point 2050 is chosen as the representative time.&quot;,<br> },</li> <li>&quot;percentile&quot;: {&quot;units&quot;: &quot;%&quot;, &quot;long_name&quot;: &quot;percentile&quot;, &quot;axis&quot;: &quot;Z&quot;},</li> </ul> <p>The attributes of the data is defined as:</p> <ul> <li>&quot;description&quot;: &quot;Contains modified <code>institute</code> <code>method</code> data used for Atlas in EUCP project.&quot;,</li> <li>&quot;history&quot;: &quot;original <code>institute</code> <code>method</code> data files ...&quot;,</li> </ul> <p><strong>output file names</strong></p> <p>output_file_name = <code>prefix_activity_institution-id_source_method_sub-method_cmor-var</code></p> <p>example: atlas_EUCP_CNRM_CMIP6_KCC_cons_tas.nc</p> <p><strong>Reference</strong>:</p> <p><a href="\&quot;https://eucp-project.github.io/atlas/about\&quot;">https://eucp-project.github.io/atlas/about</a></p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Multiplexed fluorescence imaging based on cycles, raw and processed data.

<p>This dataset was created from a larger acquisition in order to provide an example of reasonnable size, as a companion data set to the F1000Research paper preprint DOIXXX.</p> <ul> <li>The original raw data including metadata files are included in <strong>Microscope_Output.zip.</strong></li> <li><strong>Experiment.json</strong> and<strong> channelnames.txt </strong>are the ones generated by the acquisition software. They are the only files needed when starting from one of the processed data set below.</li> <li>The deconvolution obtained with the commercial software Microvolution is also provided in <strong>bu_deconvolution.zip.</strong> To start from Step 1(Extended Depth of Field) instead of Step 0 (deconvolution), unzip this file in your output directory and rename the folder bu_deconvolution to out.</li> <li>The extended field of view 2D images created from step 0 to step 2, provided for convenince in <strong>edfonly.zip</strong></li> <li>The final files generated by trhe Multiplex processor, including the segmentation mask , are provided in<strong> finaloutput.zip</strong>. These files can be used in a specific analysis software.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

EEG Data for: "Cortical oscillations and entrainment in speech processing during working memory load"

<p>This repository contains EEG and audio data used and described in:</p> <p><strong>Hjortkj&aelig;r, J, M&auml;rcher-R&oslash;rsted, J, Fuglsang, SA, Dau, T (2018). Cortical oscillations and entrainment in speech processing during working memory load. European Journal of Neuroscience.&nbsp;</strong><strong>doi</strong><strong>:10.1111/ejn.13855</strong></p> <p>Please cite this article when using the data</p> <p>&nbsp;</p> <p>The MAT-files contain the aligned EEG and audio data for each subject (N=22). The envelopes of the speech audio (without noise) have been extracted as described in the paper. Each file (data_N.mat) contains a Matlab struct in the format of the Fieldtrip toolbox containing the following fields:</p> <p>&nbsp;</p> <p>data.trial:&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;EEG and audio data for all 40 trials [channels x timepoints]</p> <ul> <li>channels 1-64: scalp EEG</li> <li>channel 65: left mastoid electrode</li> <li>channel 66: right mastoid electrode</li> <li>channel 67: horizontal EOG</li> <li>channel 68: vertical EOG for left eye</li> <li>channel 69: vertical EOG for right eye</li> <li>channel 70: audio envelopes</li> </ul> <p>data.trialinfo:&nbsp; &nbsp; &nbsp;Experimental condition in each trial</p> <ul> <li>1 = low noise, 1-back</li> <li>2 = low noise, 2-back</li> <li>3 = high noise, 1-back</li> <li>4 = high noise, 2-back</li> </ul> <p>data.time:&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample indices for each trial in seconds</p> <p>data.label:&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Name of each channel in data.trial</p> <p>data.fsample:&nbsp; &nbsp; EEG/audio sampling rate in Hz (128)</p>

opencc-by-sa-4.0Jan 2018View details →
zenodo48/100

Intermediate processing steps of quality-checking of Antarctic Circumnavigation Expedition (ACE) cruise track data.

<p><strong>Dataset abstract</strong></p> <p>The Antarctic Circumnavigation Expedition (ACE), undertaken in the austral summer of 2016/2017 recorded the cruise track using two independent geo-location instruments: one using GLobal NAvigation Satellite Systems (GLONASS; hereafter referred to as GLONASS) and another primarily using the Global Positioning System (GPS; hereafter referred to as the Trimble GPS). Daily log files were recorded in real-time from both instruments during the expedition and added to MySQL database tables. Following the expedition, quality-checking work has been undertaken to provide a one-second resolution set of positions for the cruise track. This dataset presents the intermediate files that were produced during the quality-checking, therefore it could be used to check the processing steps that have been undertaken, but should not be used as a final source of the cruise track data. Both the original raw data files and final quality-checked cruise track can be found in related datasets.</p> <p><strong>Dataset contents</strong></p> <ul> <li>ace_INSTRUMENT_YYYY-MM-DD.csv &ndash; daily files for each instrument that were output from the database, data file, comma-separated values</li> <li>ace_INSTRUMENT_concatenated_YYYY-MM.csv &ndash; input files concatenated by month and instrument, data file, comma-separated values</li> <li>flagging_data_ace_INSTRUMENT_YYYY-MM-DD.csv &ndash; daily output files for each instrument with flagged data points, data file, comma-separated values</li> <li>track_data_combined_overall_flags_YYYY-MM.csv &ndash; instrument data combined with overall data flag for each month, data file, comma-separated values</li> <li>track_data_prioritised_YYYY-MM.csv &ndash; prioritised data files with overall data flag for each month, data file, comma-separated values</li> <li>ace_INSTRUMENT_manual_position_errors.csv &ndash; files containing the manually-observed errors, metadata, comma-separated values</li> <li>in_port.csv - dates on when the ship was stationary in port, metadata, comma-separated values</li> <li>README.txt &ndash; metadata, text file</li> <li>data_file_header.txt &ndash; metadata, text file</li> </ul> <p><strong>Dataset license</strong></p> <p>This dataset containing intermediate processing files of the ACE cruise track is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>

opencc-by-4.0Oct 2019View details →
zenodo48/100

Processed GBM DEFND-seq Data

<p>This repository contains the fragments files for the gDNA component of two DEFND-seq libraries from each of two GBM patients.&nbsp; The correspond scRNA-seq data and additional metadata can be found on the Gene Expression Omnibus (GEO) under accession GSE224149. The two samples correspond to GEO entries GSM7817780 and GSM7016419.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Floating Car Data Collection for Processing and Benchmarking

<p>The dataset is outcome of a paper &quot;Floating Car Data Map-matching Utilizing the Dijkstra Algorithm&quot; accepted for 3rd International Conference on Data Management, Analytics &amp; Innovation held in Kuala Lumpur, Malaysia in 2019.</p> <p>The floating car data (FCD representing movement of cars with their position in time) is produced by the traffic simulator software (further referred to as Simulator) published in [1] and can be used as an input for data processing and benchmarking. The dataset contains FCD of various quality levels based on the routing graph of the Czech Republic derived from Open Street Map <a href="https://www.openstreetmap.org">openstreetmap.org</a>.<br> <br> Should the dataset be exploited in scientific or other way, any acknowledgement or references to our paper [1] and dataset are welcomed and highly appreciated.</p> <p><strong>Archive contents</strong></p> <p>The archive contains following folders.</p> <p><strong>city_oneway</strong> and <strong>city_roadtrip </strong>- FCD from the city of Brno, Czech Republic where FCD is based on Origin-Destination in case of oneway and Origin-Destination-Origin in case of a road trip</p> <p><strong>intercity_oneway </strong>and <strong>intercity_roadtrip </strong>- FCD from cities of Brno, Ostrava, Olomouc and Zlin, all Czech Republic where FCD is based on Origin-Destination in case of oneway and Origin-Destination-Origin in case of a road trip</p> <p><strong>Content explanation</strong></p> <p>All four of mentioned folders contain raw FCD as they come from our Simulator, post-processed FCD enriching Simulator FCD, and obfuscated raw FCD (of both low and high obfuscation level). In the both obfuscated data sets, each measured point was moved in a random direction a number of meters given by drawing a number from a Gaussian distribution. We utilized two Gaussian distributions, one for the roads outside the city (N(0,10) for the lower and N(0,20) for the higher obfuscation level) and one for the roads inside the city (N(0,15) and N(0,30) respectively). Then some predefined number of randomly chosen points were removed (3% in our case). This approach should roughly represent real conditions encountered by FCD data as described by El Abbous and Samanta [2].</p> <p>In case of post-processed road trip data, there is one extra dataset with &quot;cache&quot; suffix representing the very same dataset limited to a 5-minute session memoization. This folder also contains a picture of processed FCD represented on a map.</p> <p><strong>Data format</strong><br> Standard UTF-8 encoded CSV files, separated by a semicolon with the following columns:</p> <p><strong>RAW</strong></p> <p><em>Header</em></p> <p>session_id;timestamp;lat;lon;speed;bearing;segment_id</p> <p><em>Data</em></p> <p>session_id: (Type: unsigned INT) - session (car) identifier<br> timestamp: (Type: datetime) - timestamp in UTC<br> lat: (Type: unsigned long) - latitude as used in Google maps<br> lon: (Type: unsigned long) - longitude as used in Google maps<br> speed: (Type: unsigned INT) - actual speed in kmh<br> bearing: (Type: unsigned INT) - actual bearing in angles 0-360<br> segment_id: (Type: unsigned long) - unique edge identifier</p> <p><strong>POST-PROCESSED</strong></p> <p><em>Header</em></p> <p><br> gid;car_id;point_time;lat;lon;segment_id;speed_kmh;speed_avg_kmh;distance_delta_m;distance_total_m;speedup_ratio;duration;segment_changed;duration_segment;moved;duration_move;good;duration_good;bearing;interpolated</p> <p><em>Data</em></p> <p>gid: (Type: unsigned long) - global identifier of a record<br> car_id: (Type: unsigned INT) - session (car) identifier<br> point_time: (Type: datetime) - timestamp with timezone<br> lat: (Type: unsigned long) - latitude as used in Google maps<br> lon: (Type: unsigned long) - longitude as used in Google maps<br> segment_id: (Type: unsigned long) - unique edge identifier<br> speed: (Type: unsigned INT) - actual speed in kmh<br> speed_avg_kmh: (Type: unsigned long) - actual average speed of a car in kmh<br> distance_delta_m: (Type: unsigned long) - actual distance delta in metres<br> distance_total_m: (Type: unsigned long) - actual total distance of a car in metres<br> speedup_ratio: (Type: unsigned long) - actual speed-up ratio of a car<br> duration: (Type: time) - actual duration of a car<br> segment_changed: (Type: boolean) - signals if actual segment of a car differs from the previous one<br> duration_segment: (Type: time) - actual duration on a segment of a car<br> moved: (Type: boolean) - signals if actual position of a car differs from the previous one<br> duration_move:(Type: time) - actual duration of a car since moving<br> good: signals if actual record values satisfies all data constraints (all true as derived from Simulator)<br> duration_good: actual duration of a car since when all constraints conditions satisfied<br> bearing: (Type: unsigned INT) - actual bearing in angles 0-360<br> interpolated: (Type: boolean) - signals if actual segment identifier is calculated (all false as derived from Simulator)</p> <p><strong>References</strong><br> <br> [1] <em>V. Pto&scaron;ek, J. &Scaron;evč&iacute;k, J. Martinovič, K. Slaninov&aacute;, L. Rapant, and R. Cmar, </em><em>Real-time</em><em> traffic simulator for self-adaptive navigation system validation, Proceedings of EMSS-HMS: Modeling &amp; </em><em>Simulation</em><em> in Logistics, Traffic &amp; Transportation, 2018.</em></p> <p>[2] <em>A. El </em><em>Abbous</em><em> and N. Samanta. A </em><em>modeling</em><em> of GPS error </em><em>distri-butions</em><em>, In proceedings of 2017 European Navigation Conference (ENC), 2017.</em></p>

opencc-by-4.0Dec 2018View details →
zenodo48/100

Data and code from "Water availability is a stronger driver of soil microbial processing of organic nitrogen than tree species composition"

<p>#### Data description<br> Data from large scale, long-term tree diversity experiment in southwestern France (<a href="https://sites.google.com/view/orpheeexperiment/home">ORPHEE</a>), additionally manipulating water contraint. Variables presented are soil nitrogen cycling rates measured using isotope pool dilutions.</p> <p>Companion paper is found here:</p> <p>Maxwell TL, Augusto L, Tian Y, Wanek W &amp; Fanin N (2023). Water availability is a stronger driver of soil microbial processing of organic nitrogen than tree species composition. <em>European Journal of Soil Science</em>. <a href="https://doi.org/10.1111/ejss.13350">https://doi.org/10.1111/ejss.13350</a></p> <p>#### Metadata<br> Soil sampling: July 2020<br> Maxwell_ShortComm_Data.csv data description</p> <p>ID: unique identifier per sample<br> Block: numbered 1-6. Blocks 1,3,6 are control (unirrigated), Blocks, 2,4,5 are irrigated<br> Plot: numbered plot according to the ORPHEE design. Plot 1 = BP, Plot 5 = PP, Plot 9 = BP_PP<br> Espece: species ID. BP = pure birch (<em>Betula pendula</em>), PP = pure pine (<em>Pinus pinaster</em>), BP_PP (50% mixed birch-pine)<br> Rep: sample replicate, 3 replicates per plot<br> Sample name: long unique identifier per sample. Concatenation of Block, Plot, and Espece<br> PD: gross protein depolymerization rates (micrograms nitrogen per grams dry soil per day = &micro;g N g-1 d-1)<br> AAU: gross free amino acid uptake rates (&micro;g N g-1 d-1)<br> Cmicrobial_ug_g: microbial biomass carbon (&micro;g C g-1)<br> Nmicrobial_ug_g: microbial biomass nitrogen (&micro;g N g-1)<br> MRT_FAA_hrs: mean residence time of free amino acids (hours)<br> FAA_ugN_g: free amino acids (&micro;g N g-1)<br> Moisture_percent: soil moisture percent (%)<br> N_nonfumige_ug_g: nitrogen from non fumigated soils, i.e. extractable N (&micro;g N g-1)<br> C_nonfumige_ug_g: nitrogen from non fumigated soils, i.e. extractableC (&micro;g C g-1)</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Processing of MODIS-Aqua data with Self-Organizing Maps NeuroVaria method for the southern canary upwelling system

<p>Abstract</p> <p>This ocean color dataset is derived from MODIS_Aqua sensor measurements covering the Southern Canary upwelling system. The raw L1A measurements were downloaded from NASA&#39;s Ocean Color web site and then processed using the Ocean Biology Processing Group&#39;s (OBPG) Multi-Sensor Level-1 to Level-2 (MSL12) code. The l2gen program, based on its standard process, generates Level-2 parameters consisting of the top of atmosphere radiance, the radiance of each ocean and atmosphere component, the measurement angles, Level-2 flags, ... The top of atmosphere radiance is pre-corrected to keep only a dependence on the diffuse transmittance, the aerosol contribution and the water leaving radiance.</p> <p><br> The pre-corrected product and measurement angles are assimilated using the Self-Organizing Map<br> NeuroVaria (SOM-NV) code (Diouf et al., 2013). SOM-NV is an algorithm based on two statistical models<br> that classify a dataset into a map, and then use the information from that map to deliver atmospheric and oceanic parameters from the satellite observation.</p> <p>The parameters of interest are the remote sensing reflectance spectra (Rrs(&lambda;)) and the aerosol optical thickness (AOT) at 869 nm (aot_869). The Rrs at blue (443 and 488 nm) and green (547 nm) are used to calculate chlorophyll-a concentration from the OBPG OCx algorithm (chl_ocx, O&#39;Reilly et al., 1998; Mobley et al., 2016).</p> <p>These geophysical parameters are projected onto a fixed grid at 1/96&deg; resolution and archived in a daily netcdf format files. Each file contains five visible reflectances Rrs(&lambda;) (with &lambda; = 412, 443, 488, 531, and 547 nm), chl_ocx, aot_869, and latitude and longitude coordinates. These parameters are described in the files, along with the global attributes.</p> <p><br> The netcdf files are formatted as follows: SOM-NV-Ayyyydddhhmmss.nc; where yyyy = year; ddd = Julian<br> day; hh = hour; mm = minute; ss = second. The extension &quot;Ayyyydddhhmmss.nc&quot;, corresponds to the name<br> of the MODIS_aqua file of the day. When two input files exist for the same day, within 5 minutes, the two<br> scans are concatenated and the orbit keeps the name of the second file.<br> All files are compressed internally to a size of 4, to facilitate transfers.</p> <p>%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%</p> <p>R&eacute;sum&eacute;</p> <p>Ce jeu de donn&eacute;es de couleur de l&rsquo;eau est issu des mesures du capteur MODIS_Aqua sur la partie sud du syst&egrave;me d&rsquo;upwelling des Canaries. Les mesures brutes L1A ont &eacute;t&eacute; t&eacute;l&eacute;charg&eacute;es du site Ocean Color de la NASA, puis trait&eacute;es &agrave; l&rsquo;aide du code de traitement &laquo;&nbsp;Multi-Sensor Level-1 to Level-2 (MSL12)&nbsp;&raquo; du groupe Ocean Biology Processing Group (OBPG). La version standard du programme l2gen g&eacute;n&egrave;re les param&egrave;tres de niveau 2 constitu&eacute;s de la luminance totale mesur&eacute;e, de la luminance de chaque composante du syst&egrave;me oc&eacute;an-atmosph&egrave;re, des angles de mesures, des masques de niveau 2, &hellip;. La luminance totale est pr&eacute;-corrig&eacute;e pour ne garder qu&rsquo;une d&eacute;pendance &agrave; la transmittance diffuse, &agrave; la contribution des a&eacute;rosols et &agrave; la luminance marine.<br> <br> Le produit pr&eacute;-corrig&eacute; et les angles de mesure sont assimil&eacute;s &agrave; l&rsquo;aide du code Self-Organizing Map NeuroVaria (SOM-NV) de Diouf et al. (2013). SOM-NV est un algorithme bas&eacute; sur deux mod&egrave;les statistiques qui permettent de classer un ensemble de donn&eacute;es sur une carte, puis d&rsquo;utiliser les informations de cette carte pour restituer les param&egrave;tres atmosph&eacute;riques et oc&eacute;aniques de l&rsquo;observation satellite.<br> <br> Les param&egrave;tres restitu&eacute;s sont les spectres de r&eacute;flectance marine (Rrs(&lambda;)) et l&rsquo;&eacute;paisseur optique des a&eacute;rosols (AOT) &agrave; 869 nm (aot_869). Les Rrs au bleu (443 et 488 nm) et au vert (547 nm) servent &agrave; calculer la concentration en chlorophylle-a &agrave; partir de l&rsquo;algorithme OCx de OBPG (chl_ocx).<br> <br> Ces param&egrave;tres g&eacute;ophysiques sont projet&eacute;s sur une grille fixe &agrave; 1/96&deg; de r&eacute;solution et archiv&eacute;s au format de fichiers netcdf journaliers. Chaque fichier netcdf contient cinq r&eacute;flectances du visible Rrs(&lambda;) (avec &lambda; = 412, 443, 488, 531 et 547 nm), la chl_ocx, l&rsquo;aot_869, et les coordonn&eacute;es latitude et longitude. Ces param&egrave;tres sont d&eacute;crits dans les fichiers, ainsi que les attributs globaux.</p> <p><br> Les fichiers netcdf sont format&eacute;s comme suite : SOM-NV-Ayyyydddhhmmss.nc ; avec yyyy = ann&eacute;e ; ddd =<br> jour julien ; hh = heure ; mm = minute ; ss = seconde. L&#39;extension &quot;Ayyyydddhhmmss.nc&quot;, correspond au<br> nom du fichier MODIS_aqua du jour. Dans le cas o&ugrave; deux fichiers existent pour un m&ecirc;me jour, &agrave; 5 minutes<br> pr&egrave;s, les deux scans sont concat&eacute;n&eacute;s et l&#39;orbite garde le nom du deuxi&egrave;me fichier.<br> Tous les fichiers sont compress&eacute;s en interne &agrave; un niveau 4, pour faciliter le transfert.</p>

opencc-by-4.0May 2023View details →
edi48/100

CCE LTER process cruise, in the California Current region, event log records including date, time, position and activity for use in post-cruise data integration based on co-sampling indexes. From 2006 to 2019 CCE LTER used a locally developed event logging system. During P2107, CCE LTER started to utilize the R2R Event Logger on UNOL ships, 2006 - 2024 (ongoing).

The event logger program developed and maintained by the California Cooperative Oceanic Fisheries Investigations, SIO, program is used aboard CCE LTER process cruises to create indexes with temporal, spatial and activity information for post-cruise data integration. The event log is configured aboard the ship for the recording of sampling events by both ship crew personnel on the bridge, and research personnel in the lab. The event log is processed post-cruise to correct for various errors.

openCC0Aug 2025View details →
edi48/100

Software for processing data from a fast-responding RINKO EC oxygen/temperature sensor (JFE Advantech Co, Ltd)

This dataset describes how data from a fast-responding JFE Advantech RINKO EC ARO-EC-CM sensor connected to a Nortek Vector is processed to obtain accurate aquatic eddy covariance measurements. The code and documentation are stored in a .zip file. It consists of a manual, Fortran source code, a definition file and a complied executable suitable for running on Microsoft Windows. The software development was supported by NSF funding to PI Berg (OCE-1824144, OCE-2223204).

openCustomJul 2022View details →
zenodo44/100

Internet use: participating in social networks [percentage of individuals] processed Eurostat data [CEEMID indicator]

<p>The indicator&nbsp;&#39;<strong>Internet use: participating in social networks (creating user profile, posting messages or other contributions to facebook, twitter, etc.) [percentage of individuals]</strong>&#39; from the Eurostat statistical product&nbsp;<em>Individuals who used the internet, frequency of use and activities.</em></p> <p>- NUTS2013 regional codes are recoded to NUTS2016<br> - missing data is handled with last observation carry forward, next observation carry back, linear interpolation<br> -NUTS2 areas are imputed when only NUTS1 level data is available.&nbsp;<br> <br> The original dataset is available here:<br> <a href="https://appsso.eurostat.ec.europa.eu/nui/show.do?dataset=isoc_r_iuse_i&amp;lang=en">https://appsso.eurostat.ec.europa.eu/nui/show.do?dataset=isoc_r_iuse_i&amp;lang=en</a></p> <p>More about CEEMID: <a href="http://ceemid.eu">www.ceemid.eu</a><br> Get in touch: <a href="http://danielantal.eu/#contact">danielantal.eu/#contact</a></p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Processed data and analysis results for 104 RBPs

<p>This repository makes available the processed data and the results&nbsp;of our SURF paper.&nbsp;</p> <p>The paper presents the <strong>S</strong>tatistical <strong>U</strong>tility for <strong>R</strong>BP <strong>F</strong>unctions (SURF) for integrative analysis of RNA-seq and CLIP-seq data. The goal of SURF is to identify alternative splicing (AS), alternative transcription initiation (ATI), and alternative polyadenylation (APA) events regulated by individual RBPs and elucidate protein-RNA interactions governing these events. We applied&nbsp;the SURF pipeline to analyze 104 RBP data sets (from <a href="https://www.encodeproject.org">ENCODE</a>) and performed downstream&nbsp;analysis. Check out the browsable results from this <a href="http://www.statlab.wisc.edu/shiny/surf/">shiny</a>&nbsp;app!</p> <p>The current repository includes:</p> <ul> <li>meme.326.input.zip -- input of 326 MEME runs on SURF-inferred location features</li> <li>meme.326.output.zip -- output of 326 MEME runs on SURF-inferred location features</li> <li>surf_inferred_feature.gtf -- SURF-inferred location features for 52 RBPs</li> <li>gencode.v24.annotation.filtered.gtf -- filtered genome annotation used for ENCODE data analysis</li> <li>Homo_sapiens.GRCh37.71.primary_assembly.protein_coding.gtf -- &nbsp;genome annotation used for simulation study</li> <li>simulation_truth.txt -- truth parameters used for RNA-seq simulation</li> <li>[RBP].results.rds -- SURF output (each an&nbsp;R object) for 104&nbsp;RNA-binding proteins&nbsp;([RBP] the&nbsp;protein name).&nbsp;</li> </ul> <p>For reproducing&nbsp;the results,&nbsp;the source code is available at DOI:&nbsp;<a href="https://doi.org/10.5281/zenodo.3779853">10.5281/zenodo.3779853</a>.</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Grammatical category and the neural processing of phrases - EEG data

<p>EEG data presented in the paper</p> <p>Grammar, lexical category and the neural processing of phrases</p> <p>Amelia Burroughs, Nina Kazanina, Conor Houghton</p> <p>abstract:</p> <p><br> The interlocking roles of lexical, syntactic and semantic processing in language comprehension has been the subject of longstanding debate. Recently, the cortical response to a frequency-tagged linguistic stimulus has been shown to track the rate of phrase and sentence, as well as syllable, presentation. This could be interpreted as evidence for the hierarchical processing of speech, or as a response to the repetition of grammatical category. To examine the extent to which hierarchical structure plays a role in language processing we record EEG from human participants as they listen to isochronous streams of monosyllabic words. Comparing responses to sequences in which grammatical category is strictly alternating and chosen such that two-word phrases can be grammatically constructed -&nbsp;cold food loud room&nbsp;- or is absent - rough give ill tell - showed cortical entrainment at the two-word phrase rate was only present in the grammatical condition. Thus, grammatical category repetition alone does not yield entertainment at higher level than a word. On the other hand, cortical entrainment was reduced for the mixed-phrase condition that contained two-word phrases but no grammatical category repetition - that word send less - which is not what would be expected if the measured entrainment reflected purely abstract hierarchical syntactic units. Our results support a model in which word-level grammatical category information is required to build larger units.</p> <p>_______________________________</p> <p>all the corresponding code is at</p> <p>github.com/conorhoughton/NeuralProcessingOfPhrases</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record