Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,505

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,505 results for “Generation”

Learn how ShareScore rates datasets ↗
zenodo44/100

Weekly plots of Great Britain's half-hourly electrical system weather dependent generation, net imports and overall demand from 2008-11-10

<p>Plots that show the electrical system transition of Great Britain, they were created to form the individual frames for a video of the transition.</p>

opencc-zeroOct 2024View details →
zenodo44/100

Dataset of "Advancing Logic Circuits with Halide Perovskite Memristors for Next-Generation Digital Systems"

<p><span>This dataset supports the article "Advancing Logic Circuits with Halide Perovskite Memristors for Next-Generation Digital Systems" &nbsp; </span></p> <p>&nbsp;</p> <p><span>Raw data for the article "<span>Advancing Logic Circuits with Halide Perovskite Memristors for Next-Generation Digital Systems</span>". For further details see the readme.txt file.</span></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Data used to generate figures for Trossman et al. for Phil. Trans. A in 2024

<p>These are .mat files for the first six figures of a manuscript intended for a special edition of Philosophical Transactions A and .data/.meta files in MITgcm format that can be read using rdmds.m for the seventh figure of the same manuscript.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Verification of library complexity in the HEK-Cas9 sublibraries - sequence data of the generated sublibraries A and B

<p>Sequence data of the generated HEK-Cas9 sublibraries A and B, linked to the manuscript 10.1128/mbio.01925-24: The <em>Bordetella</em> effector protein BteA induces host cell death by disruption of calcium homeostasis by Martin Zmuda, Eliska Sedlackova, Barbora Pravdova, Monika Cizkova, Marketa Dalecka, Ondrej Cerny, Tania Romero Allsop, Tomas Grousl, Ivana Malcova, and Jana Kamanova</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Generation of transcriptional novelty by transposable element insertions in Arabidopsis, Genome Sequencing and eccDNA Data

<p><strong>Raw Illumina sequencing data from the Manuscript entitled &quot;Generation of transcriptional novelty by transposable element insertions in Arabidopsis&quot;</strong></p> <p><strong>A. Illumina genome sequencing reads of Arabidopsis control and hcLines that contain novel transposable element insertions.</strong></p> <p>To identify the genomic position of the new <em>ONSEN</em> insertions, the extracted DNA of the 11 selected lines (nine lines with new insertions and two control lines) was sent to BGI, Hong-Kong for Illumina paired-end 150 bp sequencing, aiming for a minimum of 20X sequencing coverage. Quality control of the raw reads was done using FastQC (Andrews S. (2010). FastQC: a quality control tool for high throughput sequence data. Available online at: <a href="http://www.bioinformatics.babraham.ac.uk/projects/fastqc">http://www.bioinformatics.babraham.ac.uk/projects/fastqc</a>) and trimming/clipping was done using Trimmomatic with parameters ILLUMINACLIP: TruSeq3:2:30:10 LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 and MINLEN:36. Quality of the reads was deemed excellent and no further actions were taken.</p> <p>Samples identifications: genome_hcLineX with &quot;_1&quot; indicating the forward and &quot;_2&quot; the reverse reads.</p> <p><strong>B. Illumina eccDNA sequencing&nbsp;of Arabidopsis control and hcLines following stress treatments</strong></p> <p>Extrachromosomal circular DNA was prepared and sequenced as follows:&nbsp;twenty plants from each petri dish were pooled separately and DNA was extracted using the CTAB method (<a href="https://dx.doi.org/10.17504/protocols.io.quidwue">dx.doi.org/10.17504/protocols.io.quidwue</a>). Following the mobilome-seq method described in (Lanciano et al., 2017), for all samples, we digested linear DNA from 2 &micro;g of total DNA for 17 hours at 37<sup>o</sup>C using 10 U of PlasmidSafe (<em>LubioScience cat# E3101K</em>), followed by enzyme denaturation (30 mins at 70<sup>o</sup>C). Digested DNA was precipitated with isopropanol supplemented with 1 &micro;g of GlycoBlue coprecipitant (<em>Fisher Scientific cat# 10391565</em>). Circular DNA was then amplified through rolling circle amplification (RCA) with the Illustra TempliPhi kit (<em>GE Healthcare cat# 25-6400-10</em>), following the manufacturer recommendation and leaving the reaction for 16h at 30<sup>o</sup>C. DNA was once again precipitated with isopropanol and sent for Illumina paired end 150 bp sequencing at BGI, Hong Kong.&nbsp;</p> <p>Samples identification:&nbsp;</p> <p>eccDNA_A.thaliana_ctrl:&nbsp;control reads</p> <p>eccDNA_A.thaliana_HS: heat stressed plants reads</p> <p>eccDNA_A.thaliana_AZ_HS: reads of&nbsp;alpha-amanitin, zebularine and heat-stressed plants</p> <p>&quot;R1&quot; indicates forward and &quot;R2&quot; reverse reads.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Electricity demand data and solar generation data from Plymouth. UK

<p>This dataset was used in the Western Power Distribution Presumed Open Data competition in 2021.</p> <p>The data is provided under the&nbsp;Western Power Distribution Open Data Licence.</p> <p>There are five files:</p> <p>pv_train_set4.csv - contains solar panel data - an irradiance, power output and solar panel temperature for each half-hour from 3rd November 2017 through 2nd July 2020.</p> <p>weather_train_set4.csv - contains hourly reanalysis temperature and solar radiation at 6 weather stations near the solar panels (near Plymouth, UK).</p> <p>demand_train_set4.csv - contains half-hourly electricity demand data from a substation near Plymouth, UK, running from&nbsp;3rd November 2017 through 2nd July 2020.</p> <p>pv_test_set4.csv - contains an extra week of data to&nbsp;pv_train_set4.csv, running from 3rd July 2020 through 9th July 2020.</p> <p>demand_test_set4.csv - contains an extra week of data to&nbsp;demand_train_set4.csv,&nbsp;running from 3rd July 2020 through 9th July 2020.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Hybrid LCA database generated using ecoinvent and EXIOBASE

<p>Hybrid LCA database generated using ecoinvent and EXIOBASE, i.e., each process of the original ecoinvent database is added new direct inputs (coming from EXIOBASE) deemed missing (e.g., services). Each process of the resulting hybrid database is thus not (or at least less) truncated and the calculated lifecycle emissions/impacts should therefore be closer to reality.</p> <p>For license reasons, only the added inputs for each process of ecoinvent are provided (and not all the inputs).</p> <p><em>Why are there two versions for hybrid-ecoinvent3.5?</em></p> <p>One of the version corresponds to ecoinvent hybridized with the normal version of EXIOBASE and the other is hybridized with a capital-endogenized version of EXIOBASE.</p> <p><em>What does capital endogenization do?</em></p> <p>It matches capital goods formation to the value chains of products where they are required. In a more LCA way of speaking, EXIOBASE in its normal version does not allocate capital use to value chains. It&#39;s like if ecoinvent processes had no inputs of buildings, etc. in their unit process inventory. For more detail on this, refer to (S&ouml;dersten et al., 2019) or (Miller et al., 2019).</p> <p><em>So which version do I use?</em></p> <p>Using the version &quot;with capitals&quot; gives a more comprehensive coverage. Using the &quot;without capitals&quot; version means that if a process of ecoinvent misses inputs of capital goods (e.g., a process does not include the company laptops of the employees), it won&#39;t be added. It comes with its fair share of assumptions and uncertainties however.</p> <p><em>Why is it only available for hybrid-ecoinvent3.5?</em></p> <p>The work used for capital endogenization is not available for exiobase3.8.1.</p> <p><em>How do I use the dataset?</em></p> <p>First, to use it, you will need both the corresponding ecoinvent [cut-off] and EXIOBASE [product x product] versions. For the reference year of EXIOBASE to-be-used, take 2011 if using the hybrid-ecoinvent3.5 and 2019 for hybrid-ecoinvent3.6 and 3.7.1.</p> <p>In the four datasets of this package, only added inputs are given (i.e. inputs from EXIOBASE added to ecoinvent processes). Ecoinvent and EXIOBASE processes/sectors are not included, for copyright issues. You thus need both ecoinvent and EXIOBASE to calculate life cycle emissions/impacts.</p> <p>Module to get ecoinvent in a Python format: https://github.com/majeau-bettez/ecospold2matrix (make sure to take the most up-to-date branch)</p> <p>Module to get EXIOBASE in a Python format: https://github.com/konstantinstadler/pymrio (can also be installed with pip)</p> <p>If you want to use the &quot;with capitals&quot; version of the hybrid database, you also need to use the capital endogenized version of EXIOBASE, available here: https://zenodo.org/record/3874309. Choose the pxp version of the year you plan to study (which should match with the year of the EXIOBASE version). You then need to normalize the capital matrix (i.e., divide by the total output x of EXIOBASE). Then, you simply add the <em>normalized</em> capital matrix (K) to the technology matrix (A) of EXIOBASE (see equation below).</p> <p>Once you have all the data needed, you just need to apply a slightly modified version of the Leontief equation:</p> <p><span class="math-tex">\(\begin{equation} \textbf{q}^{hyb} = \begin{bmatrix} \textbf{C}^{lca}\cdot\textbf{S}^{lca} &amp; \textbf{C}^{io}\cdot\textbf{S}^{io} \end{bmatrix} \cdot \left( \textbf{I} - \begin{bmatrix} \textbf{A}^{lca} &amp; \textbf{C}^{d} \\ \textbf{C}^{u} &amp; \textbf{A}^{io}+\textbf{K}^{io} \end{bmatrix} \right) ^{-1} \cdot \left( \begin{bmatrix} \textbf{y}^{lca} \\ 0 \end{bmatrix} \right) \end{equation}\)</span></p> <p>q<sup>hyb</sup> gives the hybridized impact, i.e., the impacts of each process including the impacts generated by their new inputs.</p> <p>C<sup>lca</sup> and C<sup>io</sup> are the respective characterization matrices for ecoinvent and EXIOBASE.</p> <p>S<sup>lca</sup> and S<sup>io</sup> are the respective environmental extension matrices (or elementary flows in LCA terms) for ecoinvent and EXIOBASE.</p> <p>I is the identity matrix.</p> <p>A<sup>lca</sup> and A<sup>io</sup> are the respective technology matrices for ecoinvent and EXIOBASE (the ones loaded with ecospold2matrix and pymrio).</p> <p>K<sup>io</sup> is the capital matrix. If you do not use the endogenized version, do not include this matrix in the calculation.</p> <p>C<sup>u</sup> (or upstream cut-offs) is the matrix that you get in this dataset.</p> <p>C<sup>d</sup> (or downstream cut-offs) is simply a matrix of zeros in the case of this application.</p> <p>Finally you define your final demand (or functional unit/set of functional units for LCA) as y<sup>lca</sup>.</p> <p><em>Can I use it with different versions/reference years of EXIOBASE?</em></p> <p>Technically speaking, yes it will work, because the temporal aspect does not intervene in the determination of the hybrid database presented here. However, keep in mind that there might be some inconsistencies. For example, you would need to multiply each of the inputs of the datasets by a factor to account for inflation. Prices of ecoinvent (which were used to compile the hybrid databases, for all versions presented here) are defined in &euro;2005.</p> <p><em>What are the weird suite of numbers in the columns?</em></p> <p>Ecoinvent processes are identified through unique identifiers (uuids) to which metadata (i.e., name, location, price, etc.) can be retraced with the appropriate metadata files in each dataset package.</p> <p><em>Why is the equation (I-A)<sup>-1</sup> and not A<sup>-1</sup> like in LCA?</em></p> <p>IO and LCA have the same computational background. In LCA however, the convention is to represents outputs and inputs in the technology matrix. That&#39;s why there is a diagonal of 1s (the outputs, i.e. functional units) and negative values elsewhere (inputs). In IO, the technology matrix does not include outputs and only registers inputs as positive values. In the end, it is just a convention difference. If we call T the technology matrix of LCA and A the technology matrix of IO we have T = I-A. When you load ecoinvent using ecospold2matrix, the resulting version of ecoinvent will already be in IO convention and you won&#39;t have to bother with it.</p> <p><em>Pymrio does not provide a characterization matrix for EXIOBASE, what do I do?</em></p> <p>You can find an up-to-date characterization matrix (with Impact World+) for environmental extensions of EXIOBASE here: https://zenodo.org/record/3890339</p> <p>If you want to match characterization across both EXIOBASE and ecoinvent (which you should do), here you can find a characterization matrix with Impact World+ for ecoinvent: https://zenodo.org/record/3890367</p> <p><em>It&#39;s too complicated...</em></p> <p>The custom software that was used to develop these datasets already deals with some of the steps described. Go check it out: https://github.com/MaximeAgez/pylcaio. You can also generate your own hybrid version of ecoinvent using this software (you can play with some parameters like correction for double counting, inflation rate, change price data to be used, etc.).<strong> As of pylcaio v2.1, the resulting hybrid database (generated directly by pylcaio) can be exported to and manipulated in brightway2.</strong></p> <p><em>Where can I get more information?</em></p> <p>The whole methodology is detailed in (Agez et al., 2021).</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

A global inventory of solar photovoltaic generating units - dataset

<p>This is the data repository accompanying Kruitwagen, L., Story, K., Friedrich, J., Byers, L., Skillman, S., &amp; Hepburn, C. (2021) A global inventory of photovoltaic solar generating units, <strong>Nature, </strong><em>forthcoming</em>. This repository contains the training, cross-validation, test, and predicted data set as described in the publication. The contents of this repository are briefly summarised here, see the publication for further details.</p> <p><strong>Repository contents:</strong></p> <p><em>trn_tiles.geojson:&nbsp;</em>18,570&nbsp;rectangular areas-of-interest used for sampling training patch data.</p> <p><em>trn_polygons.geojson:&nbsp;</em>36,882&nbsp;polygons obtained from OSM in 2017&nbsp;used to label training patches.</p> <p><em>cv_tiles.geojson:&nbsp;</em>560&nbsp;rectangular areas-of-interest used for sampling cross-validation data seeded from <a href="https://www.wri.org/research/global-database-power-plants">WRI GPPDB</a></p> <p><em>cv_polygons.geojson:&nbsp;</em>6,281 polygons corresponding to all PV solar generating units present in cv_tiles.geojson at the end of 2018.</p> <p><em>test_tiles.geojson:&nbsp;</em>122 rectangular regions-of-interest used for building the test set.</p> <p><em>test_polygons.geojson: </em>7,263 polygons corresponding to all utility-scale (&gt;10kW) solar generating units present in test_tiles.geojson at the end of 2018.</p> <p><em>predicted_polygons.geojson:&nbsp;</em>68,661 polygons corresponding to predicted polygons in global deployment, capturing the status of deployed photovoltaic solar energy generating capacity at the end of 2018.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Assembled transcriptomes of ovary, testis, and brain (male and female) of Amphibolurus muricatus (jacky dragon) generated using Trinity v2.11.0

<p><strong><em>A. muricatus</em> transcriptome assemblies generated using Trinity v2.11.0&nbsp;(Haas et al. 2013; Grabherr et al. 2011; Henschel et al. 2012)</strong><br> &bull; Amphibolurus-muricatus_brain.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> brain (male and female).<br> &bull; Amphibolurus-muricatus_combined.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> ovary, testis, and brain (male and female).<br> &bull; Amphibolurus-muricatus_female_brain.fa.tar.gz: Trinity assembly of female <em>A. muricatus</em> brain.<br> &bull; Amphibolurus-muricatus_male_brain.fa.tar.gz: Trinity assembly of male <em>A. muricatus</em> brain.<br> &bull; Amphibolurus-muricatus_ovary.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> ovary.<br> &bull; Amphibolurus-muricatus_testis.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> testis.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <ul> <li>Grabherr, M.G., B.J. Haas, M. Yassour, J.Z. Levin, D.A. Thompson et al., 2011 Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29 (7):644-652.</li> <li>Haas, B.J., A. Papanicolaou, M. Yassour, M. Grabherr, P.D. Blood et al., 2013 De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat Protoc 8 (8):1494-1512.</li> <li>Henschel, R., M. Lieber, L.-S. Wu, P.M. Nista, B.J. Haas et al., 2012 Trinity RNA-Seq assembler performance optimization, pp. 45 in Proceedings of the 1st Conference of the Extreme Science and Engineering Discovery Environment: Bridging from the eXtreme to the campus and beyond. Association for Computing Machinery, Chicag, IL, USA.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Dispersal of alien species in relation to the historic development of hydropower generation and navigation

<p>Dataset on dispersal of alien species in relation to the historic development of hydropower generation and Navigation along the River Danube.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Surfalex HF formability study - Workflow 6 - Generate random volume element

<p>This MatFlow workflow is the sixth in a set of eight workflows developed during our formability study of the Surfalex HF (AA6016A) material. In this workflow, we generate a comparison volume element from a random texture and equiaxed microstructure. This RVE is used in a comparison of the simulated Lankford coefficients between the Surfalex model RVE and this &quot;random&quot; RVE.</p> <p>This workflow can be downloaded and explored in a Jupyter notebook, as explained in the <a href="https://github.com/LightForm-group/surfalex_data_explorer">GitHub repository here</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Surfalex HF formability study - Workflow 1 - Generate volume element

<p>This MatFlow workflow is the first in a set of eight workflows developed during our formability study of the Surfalex HF (AA6016A) material. In this first workflow, we generated a representative volume element (RVE) for the Surfalex HF material. To do this, we sampled 2000 orientations from a CTF file generated from EBSD measurements on the sheet RD-TD plane. The MTEX toolbox was used to sample the texture. The grain morphology was approximated using a Voronoi tessellation that was subsequently stretched by a factor of 1.5 in the RD direction, to mimic the slight grain elongation that was observed. The pre-processing tools in the DAMASK package were used to generated the RVE.</p> <p>This workflow can be downloaded and explored in a Jupyter notebook, as explained in the <a href="https://github.com/LightForm-group/surfalex_data_explorer">GitHub repository here</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Precipitation oxygen isoscape for mainland China from 1870 to 2017 generated based on data fusion and bias correction of iGCMs simulations

<p>The dataset includes the stable oxygen isotope of precipitation for the mainland of China over the 1870-2017 period, at a spatial resolution of 50-60 km and a monthly temporal resolution. In order to make&nbsp;full use of observations to integrate the advantages of various iGCMs, the combination of data fusion and bias correction methods are used.&nbsp;Some physical-based ancillary data are introduced in the fusion methods, including elevation and meteorological data, to enrich the climate and terrain information in the process of data fusion.&nbsp;Specifically,</p><p>(1) for the 1979-2001 period, nine simulations from six iGCMs (CAM2, GISS E, HadAM3, IsoGSM2, LMDZ4, and MIROC32) and ancillary data are fused with observations by using the CNN fusion method;</p><p>(2) for the 2002-2007 period, seven simulations from four iGCMs (GISS E, IsoGSM2, LMDZ4, and MIROC32) and ancillary data are fused by using the CNN fusion method;</p><p>(3) for the 1969-1978 period, four simulations from three iGCMs (CAM2, GISS E, and HadAM3) and ancillary data are fused by using the CNN fusion method;</p><p>(4) for the 1958-1968 and 2008-2017 periods, two iGCM simulations (CAM2 and HadAM3 for 1958-1968 and IsoGSM2 and LMDZ4 zoomed for 2008-2017) are corrected by using two BCMs, and ensemble mean (mean of four simulations) is then calculated;</p><p>(5) for the 1870-1957 period, one iGCM simulation (HadAM3) is corrected by using two BCMs, and the ensemble mean (mean of two simulations) is then calculated.</p><p>Compared with the existing iGCMs, the isoscape has high quality and stability for a large region in China at the monthly scale.&nbsp;However, it should be noted that the isoscape may be more reliable for the common periods of most iGCMs (1969-2007), but mediocre for other periods.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Dataset of 20 energy prosumers with flexibility data, distributed generation and energy storage

<p>The dataset has 20 prosumers, each with three&nbsp;appliances to provide flexibility for DR events, two PV generation resources, and an energy storage system.&nbsp;The values represent a day using 15 minutes reading periods. All the values are expressed in W, and the matrixes were created as [&nbsp;time_period x info].</p> <p>&nbsp;</p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

kac_drumset: A Dataset Generator for Arbitrarily Shaped Drums

<p>This publication documents the various datasets generated using the kac_drumset codebase. The aims of kac_drumset is to provide a robust framework for the generation and analysis of arbitrarily shaped drums. The source code for this project is available here:&nbsp;<a href="https://github.com/lewiswolf/kac_drumset">https://github.com/lewiswolf/kac_drumset</a>.</p> <p><strong>Background</strong></p> <p>Arbitrarily shaped drums are a strange family of percussion instruments and a wholly meta-physical construction in this contemporary setting. These percussive instruments possess a number of interesting musical characteristics resulting from their particular geometric designs. As it currently stands, these instruments remain largely unexplored throughout musical practice, as they were originally devised as a collection of hypothetical mathematical objects. These datasets serve to sonify these objects so as to explore these conceptual constructions in the audio domain.</p> <p><strong>Usage</strong></p> <p>To use these datasets, first install kac_drumset:</p> <pre><code class="language-bash">pip install "git+https://github.com/lewiswolf/kac_drumset.git#egg=kac_drumset"</code></pre> <p>And then in python:</p> <pre><code class="language-python">from kac_drumset import ( # methods loadDataset, transformDataset, # classes TorchDataset, ) dataset: TorchDataset = transformDataset( # load a dataset (any folder containing a metadata.json) loadDataset('absolute/path/to/data'), # alter the dataset representation, either as an end2end, fft or mel. {'output_type': 'end2end'}, ) # use the dataset for i in range(dataset.__len__()): x, y = dataset.__getitem__(i) ...</code></pre> <p>For more details on using kac_drumset, see <a href="https://github.com/lewiswolf/kac_drumset/blob/master/readme.md">the project&#39;s documentation</a>.</p> <p><strong>2000 Convex Polygonal Drums of Varying Size</strong></p> <p>Each sample in this dataset corresponds to a randomly generated convex polygon. The audio for each sample was generated using a two-dimensional physical model of a drum. Each sample is one&nbsp;second long and decays linearly.</p> <p>Contained in this dataset&nbsp;are ten different sizes of drums - 0.05, 0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.5, 0.6 - each of which is a measure of the longest vertex of each drum in meters. There are 40 different drums sampled for each size. Each drum is sampled five&nbsp;times, first by being struck in the geometric centroid, and then by being struck four more times in random locations. This dataset is labelled with the vertices of each polygon, normalised to the unit interval, and the strike location of each sample.</p> <p>The audio is sampled at 48khz, and the default representation is raw audio. Each sample is stored in the metadata.json, alongside being made available audibly as a 24-bit .wav and graphically as a .png.</p> <p><strong>5000 Circular Drums of Varying Size</strong></p> <p>Each sample in this dataset corresponds to a randomly generated circular drum. The audio for each sample was generated using additive synthesis, inferred using a closed form solution to the two dimensional wave equation. Each sample is one&nbsp;second long and decays exponentially.</p> <p>Contained in this dataset&nbsp;are 1000 different drums, each determined by a randomly generated size (0.1, 2.0) in meters. Each drum is sampled five&nbsp;times, first being struck in the geometric centroid, and then by being struck four more times in random locations. This dataset is labelled with the size of each drum and the strike location of each sample.</p> <p>The audio is sampled at 48khz, and the default representation is raw audio. Each sample is stored in the metadata.json, alongside being made available audibly as a 24-bit .wav and graphically as a .png.</p> <p><strong>5000 Rectangular Drums of Varying Dimension</strong></p> <p>Each sample in this dataset corresponds to a randomly generated rectangular drum. The audio for each sample was generated using additive synthesis, inferred using a closed form solution to the two dimensional wave equation. Each sample is one&nbsp;second long and decays exponentially.</p> <p>Contained in this dataset&nbsp;are 1000 different drums, each determined by a randomly generated size (0.1, 2.0) in meters and aspect ratio (0.25, 4.0). Each drum is sampled five&nbsp;times, first being struck in the geometric centroid, and then by being struck four more times in random locations. This dataset is labelled with the size and aspect ratio of each drum, and the strike location of each sample.</p> <p>The audio is sampled at 48khz, and the default representation is raw audio. Each sample is stored in the metadata.json, alongside being made available audibly as a 24-bit .wav and graphically as a .png.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Classifying the generation and formation channels of dynamically-formed gravitational-wave events

<p>This dataset contains all the simulations of dynamically-formed binaries performed with the software <a href="https://github.com/Kkritos/Rapster">rapster</a>, together with trained&nbsp;machine-learning classification models from (Antonelli, Kritos&nbsp;et al, in prep.), see <a href="https://github.com/aantonelli94/TheBHClassifier">the public codes online</a>.</p> <p>All items starting with &quot;mergers_*&quot; are simulations of clusters&nbsp;and they follow&nbsp;the structure reported in the documentation of&nbsp;<a href="https://github.com/Kkritos/Rapster">rapster</a>. The simulations differ in the choice of the hyperparameters for the distribution of the cluster mass, half-mass radius and initial spin distribution for the binaries.</p> <p>All items starting from &quot;RFClassifier_*&quot; are machine-learning classification models&nbsp;that use a Random Forest Classifier and that are trained with the simulations above. The models ending with &quot;*_gen&quot; predict the generation of the black holes, those with &quot;*_form&quot; predict their&nbsp;formation channels.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Global gross primary production (GPP) product generated by data fusion based on random forest

<p>Improving the ability of gross primary production (GPP) estimates to capture extreme climate perturbations and reduce the uncertainty of GPP response processes to extreme climate is a new challenge. Based on the random forest algorithm, we integrated the multimodel GPP simulation results published by the Multiscale Synthesis and Terrestrial Model Intercomparison Project, the FLUXNET flux-site-observed GPP, the standardized precipitation index (SPI) and the standardized temperature index (STI) to generate a set of global GPP time-series data products from 2001 to 2010. The new GPP product was named DFRF-GPP, referring to the GPP generated by data fusion based on random forest. DFRF-GPP is highly reliable and can be used as a valuable data source for various applications, especially in high-temperature and drought-related studies.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

University campus buildings electrical and thermal demand and generation.

<p>UVTgv has agreed to share data for both generation and consumption profile of their buildings. The load data is provided as one .csv per building and year with 8760 rows each one representing one hourly consumption of the building. The load datasets comprise 2019 and 2020 load data for electricity and heat demand, for the following buildings: ABR, C, ICSTM. The generation profiles provided are for 3 PV installations (generation profile), one solar thermal plant (capacity factor), and a mini-wind turbine (generation profile).</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Auroville consumption and generation

<p>The datasets contain hourly capacity factors for a PV plant with 394 kWp capacity and a windpower facility with 800 kWp capacity. The demand profile provided contains houlrly consumption data (kWh) of Auroville. All these datasets have been used to generate the results that can be found in D6.1 and D6.2.&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Additional results for article "A new approach for the generation of real-time GNSS low-latitude ionospheric scintillation maps"

<p>Complete set of interpolation error and correlation metrics for the approaches GDA, IDW, RBF and GPR for all the 12 pre-processing options using the SSS cross-validation scheme for the 10-hour dataset (40 maps with 16-minute interval for each approach and each pre-processing options) - file &ldquo;Complete table of interpolation errors and correlation.csv&rdquo;.</p> <p>Complete set of scintillation maps for the approaches GDA, IDW, RBF and GPR for all the 12 pre-processing options covering the 10-hour dataset (40 maps with 16-minute interval for each approach and each pre-processing options) - file &ldquo;Scintillation maps for the 10-hour dataset.zip&rdquo;.</p> <p>Comparison plots of the scintillation maps generated by the approaches GDA, IDW, RBF and GPR with the pre-processing options SAR, SMR and VQI for each of the &nbsp;40 intervals of time of 16 minutes covering the 10-hour dataset - file &ldquo;Set of maps for all 4 approaches with the SAR, SMR and VQI sets of options.zip&rdquo;.</p> <p>Sequence of scintillation maps for the 8-hour dataset generated by the GPR(VQI) approach for the three time resolutions (1, 2, and 16-minute) - file &ldquo;Scintillation maps for the 8-hour comparison dataset.zip&rdquo;.</p> <p>Animations corresponding to the sequence of maps generated by the GPR(VQI) approach for the 8-hour dataset, and for the three time resolutions (1, 2, and 16-minute) - file &ldquo;Animations of GPR(VQI) maps for the 8-hour comparison dataset.zip&rdquo;.</p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record