Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

7 results for “parquet”

Learn how ShareScore rates datasets ↗
zenodo40/100

ENTSO-E Pan-European Climatic Database (PECD 2021.3) in Parquet format

<p><strong>ENTSO-E Pan-European Climatic Database (PECD 2021.3) in Parquet format</strong></p> <p><strong>TL;DR</strong>: this is a tidy and friendly version of a subset of the PECD 2021.3 data by ENTSO-E: hourly capacity factors for wind onshore, offshore, solar PV, hourly electricity demand, weekly inflow for reservoir and pumping and daily generation for run-of-river. All the data is provided for &gt;30&nbsp;climatic years (1982-2019 for wind and solar, 1982-2016 for demand, 1982-2017 for hydropower)&nbsp;and at national and sub-national (&gt;140 zones) level.</p> <p><strong>UPDATE (19/10/2022):&nbsp;</strong>updated the demand files due after fixing a bug in the processing code (the file for 2030 was the same for 2025) and solving an issue caused by a malformed header in the ENTSO-E excel files.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>ENTSO-E has released with the latest European Resource Adequacy Assessment (<a href="https://www.entsoe.eu/outlooks/eraa/">ERAA 2021</a>) all the inputs used in the study.<br> Those inputs include:<br> - Demand dataset: <a href="https://eepublicdownloads.azureedge.net/clean-documents/sdc-documents/ERAA/Demand%20Dataset.7z">https://eepublicdownloads.azureedge.net/clean-documents/sdc-documents/ERAA/Demand%20Dataset.7z</a><br> - Climate data: <a href="https://eepublicdownloads.entsoe.eu/clean-documents/sdc-documents/ERAA/Climate%20Data.7z">https://eepublicdownloads.entsoe.eu/clean-documents/sdc-documents/ERAA/Climate%20Data.7z</a></p> <p>The data files and the methodology are available on the <a href="https://www.entsoe.eu/outlooks/eraa/2021/eraa-downloads/">official webpage</a>.&nbsp;</p> <p>As done for the previous releases (see <a href="https://zenodo.org/record/3702418#.YbmhR23MKMo">https://zenodo.org/record/3702418#.YbmhR23MKMo</a> and <a href="https://zenodo.org/record/3985078#.Ybmhem3MKMo">https://zenodo.org/record/3985078#.Ybmhem3MKMo</a>), the original data - stored in large Excel spreadsheets - have been tidied and formatted in open and friendly formats (CSV for the small tables and Parquet for the large files)</p> <p>Furthermore, we have carried out a simple country-aggregation for the original data - that uses instead &gt;140 zones.</p> <p><strong>DISCLAIMER</strong>: <em>the content of this dataset has been created with the greatest possible care. However, we invite to use the original data for critical applications and studies.&nbsp;</em></p> <p><strong>Description</strong></p> <p>This dataset includes the following files:</p> <p>- <em>capacities-national-estimates.csv</em>: installed capacity in MW per zone, technology and the two scenarios (2025 and 2030). The files include also the total capacity for each technology per country (sum of all the zones within a country)<br> - <em>PECD-2021.3-wide-LFSolarPV-2025</em> and <em>PECD-2021.3-wide-LFSolarPV-2030</em>: tables in Parquet format storing in each row the capacity factor for solar PV for a hour of the year and all the climatic years (1982-2019) for a specific zone. The two files contain the capacity factors for the scenarios &quot;National Estimates 2025&quot; and &quot;National Estimates 2030&quot;<br> - <em>PECD-2021.3-wide-Onshore-2025</em> and<em> PECD-2021.3-wide-Onshore-2030</em>: same as above but for wind onshore<br> - <em>PECD-2021.3-wide-Offshore-2025</em> and <em>PECD-2021.3-wide-Offshore-2030</em>: same as above but for wind offshore<br> - <em>PECD-wide-demand_national_estimates-2025</em> and<em> PECD-wide-demand_national_estimates-2030</em>: hourly electricity demand for all the climatic years for a specific zone. The two files contain the load for the scenarios &quot;National Estimates 2025&quot; and &quot;National Estimates 2030&quot;&nbsp;<br> - <em>PECD-2021.3-country-LFSolarPV-2025</em> and <em>PECD-2021.3-country-LFSolarPV-2030</em>: tables in Parquet format storing in each row the capacity factor for country/climatic year and hour of the year. The two files contain the capacity factors for the scenarios &quot;National Estimates 2025&quot; and &quot;National Estimates 2030&quot;<br> -<em> PECD-2021.3-country-Onshore-2025</em> and<em> PECD-2021.3-country-Onshore-2030</em>: same as above but for wind onshore<br> -<em> PECD-2021.3-country-Offshore-2025</em> and <em>PECD-2021.3-country-Offshore-2030</em>: same as above but for wind offshore<br> - <em>PECD-country-demand_national_estimates-2025</em> and <em>PECD-country-demand_national_estimates-2030</em>: same as above but for electricity demand<br> - <em>PECD_EERA2021_reservoir_pumping.zip</em>: archive with four files per each scenario: 1. table.csv with generation and storage capacities per zone/technology, 2. zone weekly inflow (GWh), 3. table.csv with generation and storage per country/technology and 4. country weekly inflow (GWh)<br> - <em>PECD_EERA2021_ROR.zip</em>: as for the previous file but the inflow is daily<br> - <em>plots.zip</em>: archive with 182 png figures with the weekly climatology for all the variables (daily for the electricity demand)</p> <p><strong>Note</strong></p> <p>I would like to thank Laurens Stoop for sharing the onshore wind data for the scenario 2030, that was corrupted in the original archive.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Snapshot data in Parquet format from ENTSO-E Transparency Platform

<p>This is a snapshot of the following three datasets downloaded from the <a href="https://transparency.entsoe.eu/content/static_content/Static%20content/knowledge%20base/SFTP-Transparency_Docs.html">ENTSO-E Transparency Platform SFTP</a>:</p> <ol> <li>ActualTotalLoad (6.1.A)</li> <li>AggregatedGenerationPerType (16.1.B C)</li> <li>DayAheadPrices (12.1.D)</li> </ol> <p>The data has been downloaded and converted from CSV to Parquet using the package {arrow} in R.</p> <p><strong>WARNING</strong>: as <a href="https://github.com/energy-modelling-toolkit/entsoe-tp-survival-kit">I have explained here</a>, the data in the Transparency Platform changes every day. This means that the data contained in these files may be different from the data currently shown on the website or available via the APIs.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

CSO Networked Gas Consumption by Dublin Postal District for Residential & Non-Residential Sector 2011-2019 in Parquet & CSV Form

<p>Metered network gas data on a Dublin postal district level from 2010-2019, translated from a HTML to parquet files.&nbsp;</p> <p>Original dataset can be found at&nbsp;https://www.cso.ie/en/releasesandpublications/er/ngc/networkedgasconsumption2019/&nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

ENTSO-E ERAA 2023 files in Parquet format

<p>demand time-series, PECD (renewables) and hydropower data from the ERAA 2023 (<a href="https://www.entsoe.eu/outlooks/eraa/2023/eraa-downloads/">ERAA Downloads | ENTSO-E &ndash; ERAA 2023 (entsoe.eu)</a>).&nbsp;</p> <p>Original data (in CSV and Excel format) is converted in Parquet in a format more suitable for modelling and analysis.</p> <p>This dataset includes:</p> <ul> <li>Demand time-series for all the modelled zones for the years 2025, 2028, 2030, 2033</li> <li>Capacity factors for all the modelled zones for the following technologies: wind onshore, wind offshore, solar PV rooftop, solar PV utility-scale, CSP with storage and without storage. All the capacity factors are available for the year 2025, 2028, 2030, 2033. All the files are available in the zipped file PECD.zip</li> <li>Hydropower data (installed capacity and time-series of daily/weekly inflow, min/max generating power, min/max reservoir levels) for run-of-river, pondage, reservoir-based plants, pumped-storage open- and closed-loop.</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo28/100

ParlerPosts_Parquet_NoScores

<p>Useful dataset for GNN applications with parler post data.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo24/100

Identifying galaxies, quasars and stars with machine learning: a new catalogue of classifications for 111 million SDSS sources without spectra - parquet format

<p>This is the same as the published data available under&nbsp;10.5281/zenodo.3768398, but in the format of parquet files. This means you can access it using Dask for convenience when using cloud compute facilities.&nbsp;</p> <p>Abstract: We used 3.1 million spectroscopically labelled sources from the Sloan Digital Sky Survey (SDSS) to train an optimised random forest classifier using photometry from the SDSS and the Widefield Infrared Survey Explorer (WISE). We applied this machine learning model to 111 million previously unlabelled sources from the SDSS photometric catalogue which did not have existing spectroscopic observations. Our new catalogue contains 50.4 million galaxies, 2.1 million quasars, and 58.8 million stars. We provide individual classification probabilities for each source, with 6.7 million galaxies (13%), 0.33 million quasars (15%), and 41.3 million stars (70%) having classification probabilities greater than 0.99; and 35.1 million galaxies (70%), 0.72 million quasars (34%), and 54.7 million stars (93%) having classification probabilities greater than 0.9. Precision, Recall, and F1 score were determined as a function of selected features and magnitude error. We investigate the effect of class imbalance on our machine learning model and discuss the implications of transfer learning for populations of sources at fainter magnitudes than the training set. We used a non-linear dimension reduction technique (Uniform Manifold Approximation and Projection: UMAP) in unsupervised, semi-supervised, and fully-supervised schemes to visualise the separation of galaxies, quasars, and stars in a two-dimensional space. When applying this algorithm to the 111 million sources without spectra, it is in strong agreement with the class labels applied by our random forest model.</p> <p>When using this dataset, please reference our paper via the journal (<a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a>) and this DOI (10.5281/zenodo.4060257). If you make use of our scripts please reference our Github repository DOI (10.5281/zenodo.3855160).</p> <p>File descriptions:</p> <p>All of these files are Pandas Dataframes, saved as uncompressed parquet files for ease of access when using cloud compute such as Dask. df_spec_classprobs.parquet&nbsp;contains the spectroscopically observed sources used for training and testing. This has been cleaned, and has the results of the random forest classifier added as additional columns (sources used for training have NaNs in the class_pred column). SDSS-ML-all.parquet contains the 111 million photometrically observed sources, with our class labels and probabilities added.</p>

opencc-by-4.0Sep 2020View details →
zenodo24/100

Pan-European Climate Database 4.1 - wind onshore data Pan-European Onshore Zones - Parquet format

<p>Data from <a href="https://cds.climate.copernicus.eu/datasets/sis-energy-pecd?tab=documentation">Climate and energy related variables from the Pan-European Climate Database derived from reanalysis and climate projections</a></p> <p>The python script to download using the CDS API are included.</p> <p>This is a Parquet (long and tidy) version of the original CSV files.&nbsp;</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record