Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,045
datasets available to search
ShareScore release 0.9.0
Dataset results
1,045 results for “Generated Data”
Data for paper Economic disparity among generations under Paris Agreement
<p>The datasets here are used to replicate the results Economic disparity among generations under Paris Agreement in paper, using the code available at https://github.com/climate-change-ucsb/generation-disparity</p> <p>The file age structure.csv is used to generate the Figure 2e in the paper, and is derived from SSP database. </p> <p>The ZIP file data.zip is the data used to calculate the benefits of climate change mitigation, which is needed for replicating the results in https://github.com/climate-change-ucsb/generation-disparity. The original source of this data is on https://github.com/country-level-scc/cscc-paper-2018/tree/master/data.</p> <p>The ZIP file gadm28_levels.shp.zip is the shapefile fo country boundaries, and is available at https://gadm.org/.</p> <p>cytemp is the data file generated by using the code and data in https://github.com/wmadavis/BDD2018 using the 02DataProcessing.R scripts (https://github.com/wmadavis/BDD2018/blob/master/scripts/02DataProcessing.R).</p> <p>Population structure.xlsx is the data used to derive the age structure.csv, and can be accessed from UN population prospect.</p> <p>Life expectancy at birth.xls is available at world bank open data.</p> <p>temp.csv is the temperature change from 2020 to 2100 in 169 countries.</p>
PSYCHE-D: predicting change in depression severity using person-generated health data (DATASET)
<p>This dataset is made available under <a href="https://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)</a>. See LICENSE.pdf for details.</p> <p><strong>Dataset description</strong></p> <p>Parquet file, with:</p> <ul> <li>35694 rows</li> <li>154 columns</li> </ul> <p>The file is indexed on [<em>participant</em>]_[<em>month</em>], such that 34_12 means month 12 from participant 34. All participant IDs have been replaced with randomly generated integers and the conversion table deleted.</p> <p>Column names and explanations are included as a separate tab-delimited file. Detailed descriptions of feature engineering are available from the linked publications.</p> <p>File contains aggregated, derived feature matrix describing person-generated health data (PGHD) captured as part of the DiSCover Project (<a href="https://clinicaltrials.gov/ct2/show/NCT03421223">https://clinicaltrials.gov/ct2/show/NCT03421223</a>). This matrix focuses on individual changes in depression status over time, as measured by PHQ-9.</p> <p>The DiSCover Project is a 1-year long longitudinal study consisting of 10,036 individuals in the United States, who wore consumer-grade wearable devices throughout the study and completed monthly surveys about their mental health and/or lifestyle changes, between January 2018 and January 2020.</p> <p>The data subset used in this work comprises the following:</p> <ul> <li>Wearable PGHD: step and sleep data from the participants’ consumer-grade wearable devices (Fitbit) worn throughout the study</li> <li>Screener survey: prior to the study, participants self-reported socio-demographic information, as well as comorbidities</li> <li>Lifestyle and medication changes (LMC) survey: every month, participants were requested to complete a brief survey reporting changes in their lifestyle and medication over the past month</li> <li>Patient Health Questionnaire (PHQ-9) score: every 3 months, participants were requested to complete the PHQ-9, a 9-item questionnaire that has proven to be reliable and valid to measure depression severity</li> </ul> <p>From these input sources we define a range of input features, both static (defined once, remain constant for all samples from a given participant throughout the study, e.g. demographic features) and dynamic (varying with time for a given participant, e.g. behavioral features derived from consumer-grade wearables).</p> <p>The dataset contains a total of 35,694 rows for each month of data collection from the participants. We can generate 3-month long, non-overlapping, independent samples to capture changes in depression status over time with PGHD. We use the notation ‘SM0’ (sample month 0), ‘SM1’, ‘SM2’ and ‘SM3’ to refer to relative time points within each sample. Each 3-month sample consists of: PHQ-9 survey responses at SM0 and SM3, one set of screener survey responses, LMC survey responses at SM3 (as well as SM1, SM2, if available), and wearable PGHD for SM3 (and SM1, SM2, if available). The wearable PGHD includes data collected from 8 to 14 days prior to the PHQ-9 label generation date at SM3. Doing this generates a total of 10,866 samples from 4,036 unique participants.</p>
Energy consumption and PV generation data of 15 prosumers (15 minute resolution)
<p><strong>Energy consumption and PV generation data of 15 prosumers (15 minute resolution)</strong></p> <p>Sérgio Ramos, João Soares, Zahra Foroozandeh, Inês Tavares, Zita Vale</p> <p><strong>Paper title: (All papers)</strong></p> <p>Type: Energy consumption and PV generation data</p> <p>Duration: Year 2019 (15 minute – 35 040 periods)</p> <p>Resolution: 15 minutes</p> <p>Application: Paper submitted on</p> <p>Sheets description:</p> <ul> <li>Total PV production: Contains the generation of the PV panels;</li> <li>Common services: Contains information of the energy consumption of the common services of the building;</li> <li>Consumer 1-15: Contains the information of the energy consumption of each consumer.</li> </ul>
Data for Bernabeu-Herrero et al, Mutations causing premature termination codons discriminate and generate cellular and clinical variability in HHT
<p>This dataset is for the 2024 manuscript<strong>: </strong></p> <p><strong>Bernabéu-Herrero ME, Patel D, Bielowka A, Zhu J, Jain K, Mackay IS, Chaves Guerrero P, Emanuelli G, Jovine L, Noseda M, Marciniak SJ, Aldred MA, Shovlin CL. </strong></p> <p><strong>Mutations causing premature termination codons discriminate and generate cellular and clinical variability in HHT. </strong></p> <p><strong>Blood. 2024 May 30;143(22):2314-2331. </strong></p> <p><strong>doi: 10.1182/blood.2023021777. PMID: 38457357; PMCID: PMC11181359.</strong></p> <p>It was originally uploaded in 2021 ahead of an earlier manuscript submission<br>- see https://www.biorxiv.org/content/10.1101/2021.12.05.471269v1</p>
SEED-G: Simulated EEG Data Generator for testing connectivity algorithms
<p>SEED-G toolbox was developed in MATLAB environment (tested on version R2017a and R2020b) and released on the GitHub page <a href="https://github.com/aanzolin/SEED-G-toolbox">https://github.com/aanzolin/SEED-G-toolbox</a> (accessed date 12 April 2021). It is organized in the following subfolders:</p> <ul> <li> <p><strong>main</strong>: it is the core of the toolbox and contains all the functions for the generation of EEG data according to a predefined ground-truth network.</p> </li> <li> <p><strong>dependencies</strong>: containing parts of other toolboxes required to successfully run SEED-G functions. The links to the full packages can be found in the documentation on the GitHub page. The additional packages are Brain Connectivity Toolbox (BCT) [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B40-sensors-21-03632">40</a>], FieldTrip [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B41-sensors-21-03632">41</a>], Multivariate Granger Causality Toolbox (MVGC) [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B24-sensors-21-03632">24</a>], and AsympPDC Package (PDC_AsympSt) [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B42-sensors-21-03632">42</a>,<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B43-sensors-21-03632">43</a>]. Additionally, the implemented forward model is solved according to the New York Head (NYH) model, whose parameters are contained in the structure available on the ICBM-NY platform [<a href="https://www.mdpi.com/1424-8220/21/11/3632/htm#B28-sensors-21-03632">28</a>].</p> </li> <li> <p><strong>real data</strong>: containing real EEG data acquired from one healthy subject during resting state at scalp level (‘EEG_real_sources.mat’) and its reconstructed version in source domain (‘sLOR_cortical_sources.mat’). These signals can be employed to extract the AR components to be included in the model to generate data with the same spectral properties of the real ones.</p> </li> <li> <p><strong>demo</strong>: containing examples of MATLAB scripts to be used to learn the different functionalities of the toolbox. For example, the code ‘run_generation.m’ allows to specify the directory containing the real sources and each specific input of the function ‘simulatedData_generation.m’.</p> </li> <li> <p><strong>auxiliary functions</strong>: containing either original MATLAB functions or modified version of free available functions.</p> </li> </ul>
Data for "Using an Uncertainty Quantification Framework to Calibrate the Runoff Generation Scheme in E3SM Land Model V1"
<p>The domain file and surface data file that used to run ELMv1, and processed ISIMP2a runoff data that used in <a href="https://gmd.copernicus.org/preprints/gmd-2021-401/">https://gmd.copernicus.org/preprints/gmd-2021-401/</a></p> <p>ELM_runoff_parameter_post.nc contains the ELM runoff generation relevant parameter posteriors at a global half degree spatial resolution.</p>
Illumina next generation ddRAD sequencing SNP data from: Contrasting genetic diversity and structure between endemic and widespread damselfishes are related to differing adaptive strategies
<p class="MsoNormal"><strong><u><span>Aim:</span></u></strong><span> Discerning when, where, and how processes of isolation lead to differing biogeography is especially complex for marine species with similar ecological niches and within the same geographic location. We assessed population genetics of congeneric and ecologically similar damselfishes within their overlapping distributions and across potential barriers to geneflow.</span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Taxon:</span></u></strong><span> <em>Dascyllus marginatus </em>(endemic) and <em>Dascyllus abudafur </em>(widespread)<em>.</em></span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Location:</span></u></strong><span> Coral reefs from the Red Sea, Djibouti, Yemen, Oman, and Madagascar. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Methods:</span></u></strong><span> We used RADseq derived SNPs to investigate key differences in population genetics between both species and discuss barriers shaping genetic differentiation (neutral vs. selective) and biogeography. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Results:</span></u></strong><strong><span> </span></strong><em><span>Dascyllus marginatus </span></em><span>inhabited the Red Sea, the coasts of Yemen (including Socotra), and the Gulf of Oman. <em>Dascyllus abudafur</em> species was present from the Red Sea to Madagascar but was absent from Yemen and Oman. Populations of <em>D. marginatus </em>had an order of magnitude higher genetic differentiation compared to <em>D. abudafur</em>, as well as several outlier loci (suggesting selective pressure), which were absent in <em>D. abudafur</em> despite equal sampling locations. In both species, specimens from the Red Sea and Djibouti formed one genetic cluster separated from all other locations. </span></p> <p class="MsoNormal"> </p> <p class="MsoNormal"><strong><u><span>Main conclusions:</span></u></strong><span> The stronger genetic structure at smaller geographic scale of the endemic species seems associated to faster adaptation to environmental differences; whereas the widespread species only experienced reduced geneflow and neutral differentiation at much larger geographic scales. Restrictive transitions (between the Gulf of Aqaba and the Red Sea or the Red Sea and the Gulf of Aden) did not affect the genetic architecture of either species, while the environmental shift within the Red Sea (at 22°N/20°N) affected the endemic but not the widespread species. Samples from continental Yemen revealed that a genetic break in the Gulf of Aden likely reflects historical colonization processes and not contemporary environmental regimes.</span></p>
Data for "Density staircases generated by symmetric instability in a cross-equatorial deep western boundary current"
<p>Data associated with the git repository <a href="https://GitHub.com/fraserwg/dwbc-proj">dwbc-proj</a>.</p>
Photovoltaic generation data, for 3 years, regarding the 2022-3 Competition on solar generation forecasting
<p>These data were released under the 2022-3 Competition on solar generation forecasting.</p> <p>Please check our competitions: <a href="http://www.gecad.isep.ipp.pt/smartgridcompetitions">www.gecad.isep.ipp.pt/smartgridcompetitions</a></p> <p> </p> <p>The data set comprises the power generated by photovoltaic panels and data collected from a near weather station. The data was collected in 5 minutes periods. Data comprises:</p> <ul> <li>Hour</li> <li>Starting minute (inclusive) </li> <li>Ending minute (exclusive) </li> <li>Generated power (kW)</li> <li>Temperature (ºC) </li> <li>Dewpoint (ºC)</li> <li>Pressure (hPa) </li> <li>Wind Direction (Degrees)</li> <li>Wind Speed (KM/h)</li> <li>Wind Speed Gust (KM/h)</li> <li>Humidity (%)</li> <li>Hourly Precipitation (mm)</li> <li>Daily rain (mm)</li> <li>Solar Radiation (Watts/m2)</li> </ul> <p><br>The data set represents raw data without any treatment, this means that it is possible to find errors. Data can have missing data or missing reading periods, and a fixed zero (0) value, indicating a failure in the system readings.</p> <p> </p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>
Taxonomy of Knowledge Types for Synthetic Data Generation
<p>The full taxonomy of knowledge types for synthetic data generation in production.</p> <p>For more information, see <a href="https://doi.org/10.54941/ahfe1002915">IHSI 2023 conference paper</a>.</p>
Raw data for the article "Substrate-Controlled C-H or C-C Alkynylation of Cyclopropanes: Generation of Aryl Radical Cations by Direct Light Activation of Hypervalent Iodine Reagents "
<p>Raw computational, NMR, IR and MS data for the article "Substrate-Controlled C-H or C-C Alkynylation of Cyclopropanes: Generation of Aryl Radical Cations by Direct Light Activation of Hypervalent Iodine Reagents " published in Chemical Science, DOI: </p> <p><a href="https://doi.org/10.1039/D2SC04344K">https://doi.org/10.1039/D2SC04344K</a></p> <p>The number of the folders either correspond to compounds numbers in the article or the name of the folder is self-describing. All details concerning conditions and equipment for measurements can be found in the supporting information of the article.</p>
Generation of combined daily satellite-based precipitation products over Bolivia - Generated Precipitation Data
<p><strong>A journal paper published in Remote Sensing details the method to generate the data.</strong></p> <p>Saavedra, O.; Ureña, J. Generation of Combined Daily Satellite-Based Precipitation Products over Bolivia. <em>Remote Sens.</em> <strong>2022</strong>, <em>14</em>, 4195. https://doi.org/10.3390/rs14174195</p>
Data for HydrAMP - a deep generative model for antimicrobial peptide discovery
<ul> <li>data- training data for peptides < 25 AA (16.8 MB)</li> <li>models - checkpoints of HydrAMP, PepCVAE, and Basic models for every training epoch (466 MB)</li> <li>results - dumped generation results for every model. Required for running comparison notebooks (832 MB)</li> <li>wheels - custom TensorFlow packages (1 GB)</li> </ul> <p> </p>
Data from: Kinematic and hydrodynamic analyses of turning manoeuvres in penguins: Body banking and wing upstroke generate the centripetal force
<p>Penguins perform lift-based swimming by flapping their wings. Previous kinematic and hydrodynamic studies have revealed the basics of wing motion and force generation in penguins. Although these studies have focused on steady forward swimming, the mechanism of turning manoeuvres is not well understood. In this study, we examined the horizontal turning of penguins via 3D motion analysis and quasi-steady hydrodynamic analysis. Free swimming of gentoo penguins (<em>Pygoscelis papua</em>) at an aquarium was recorded, and body and wing kinematics were analysed. In addition, quasi-steady calculations of the forces generated by the wings were performed. Among the selected horizontal swimming manoeuvres, turning was distinguished from straight swimming by the body trajectory for each wingbeat. During the turns, the penguins maintained outward banking through a wingbeat cycle and utilized a ventral force during the upstroke as a centripetal force to turn. Within a single wingbeat during the turns, changes in the body heading and bearing also mainly occurred during the upstroke, while the subsequent downstroke accelerated the body forward. We also found contralateral differences in the wing motion; i.e., the inside wing of the turn became more elevated and pronated. Quasi-steady calculations of the wing force confirmed that the asymmetry of the wing motion contributes to the generation of the centripetal force during the upstroke and the forward force during the downstroke. The results of this study demonstrate that the hydrodynamic force of flapping wings, in conjunction with body banking, is actively involved in the mechanism of turning manoeuvres in penguins.</p>
Data from: Direct generation of spatially entangled qudits using quantum nonlinear holography - variances
<p>Nonlinear holography shapes the amplitude and phase of generated new harmonics using nonlinear processes. Classical nonlinear holography influenced many fields in optics, from information storage, de-multiplexing of spatial information and all-optical control of accelerating beams. Here, we extend the concept of nonlinear holography to the quantum regime. We directly shape the spatial quantum correlations of entangled photon pairs in two-dimensional patterned nonlinear photonic crystals using spontaneous parametric down conversion, without any pump shaping. The generated signal-idler pair obeys a parity conservation law that is governed by the nonlinear crystal. Furthermore, the quantum states exhibit quantum correlations and violate the Clauser-Horne-Shimony-Holt inequality, thus enabling entanglement-based quantum key distribution. Our demonstration paves the way for controllable on-chip quantum optics schemes utilizing the high-dimensional spatial degree of freedom.</p>
Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning: Raw Data and Generation Script
<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature. The dataset contains both the original raw data from the articles and the enriched ones. We also provide the scripts used to curate and enrich the raw data.</p>
The influence of electric circuit parameters on NOx generation by transient spark discharge _ data
<p>dataset for</p> <p>The influence of electric circuit parameters on NOx generationby transient spark discharge</p> <p> </p> <p>Abstract</p> <p>Nitrogen fixation, production of NO and NO<sub>2</sub> from N<sub>2</sub> and O<sub>2</sub> in air, has been investigated with<br> transient spark self-pulsing DC discharges. NO production is boosted by the addition of capacitors<br> and an inductor to the electrical circuit which drives the discharge. The quantity of NO produced<br> per joule of electrical input energy is doubled, though the quantity of NO<sub>2</sub> produced drops. The<br> yield of NO is also increased because the modified circuit enables higher discharge currents to be<br> used. NO concentrations as high as 2000 ppm were obtained with input energy densities of around<br> 300 J per litre of input gas, whilst NO<sub>2</sub> concentrations were around 150 ppm. This simple<br> modification of the driving circuit may have potential for optimizing the plasma chemistry with<br> other input gas mixtures and for scaling up nitrogen fixation from air.</p>
Raw data for the Github repository "Collection of scripts to download & process hydropower generation data in Argentina, Bolivia, Brazil, Uruguay"
<p>This is the raw data downloaded from the websites of the following system operators:</p> <p> - CNDC, Comité Nacional de Despacho de Carga (Bolivia): https://www.cndc.bo/home/index.php<br> - ONS, Operador Nacional do Sistema Elétrico (Brazil): https://www.ons.org.br/<br> - UTE, Usinas y Trasmisiones Eléctricas (Uruguay): https://www.ute.com.uy/<br> - CAMMESA, Compañía Administradora del Mercado Mayorista Eléctrico Sociedad Anónima (Argentina): https://cammesaweb.cammesa.com</p> <p>Hydropower generation data is extracted using the R scripts available here: https://github.com/matteodefelice/hydro-sam</p>
Generated data for "Limited ventilation of the central Baltic Sea due to elevated oxygen consumption" paper
<p>This data are essential for reproducing the figures from Naumov et al. "Limited ventilation of the central Baltic Sea due to elevated oxygen consumption" paper. Each archive is named after one of the ten figures and includes the data necessary for that specific figure. Some data are used in more than one figure. In some cases, performing a particular type of analysis with the given data (linear regression, for instance) is necessary to fully reproduce the figure.</p>
Python code generating the data of figures 2, 3, 4, 5 and 6 of the manuscript: The evolution of cooperation in the unidirectional linear division of labour of finite roles
<p>The evolution of cooperation is an unsolved mystery, which we see in many social and biological systems. In the study titled "The evolution of cooperation in the unidirectional linear division of labour of finite roles", we investigate under which sanction systems and how the evolution of cooperation happens in the linear division of labour. </p> <p>This python code has been used to produce the results of Figures 2, 3, 4, 5, and 6 of the manuscript. This code shows the evolution of cooperation among the population of various different groups which have different roles to play in the linear division of labour, on the basis of numerical analysis of a partial differential equation system, which originates from the replicator equations used in the evolutionary game theory. We find the locally stable equilibria using this code, which shows the ultimate results of the dynamics in the system under given parameters. Figures 3, 5, and 6 are direct products of the code, showing the dynamics of a system, and figures 2 and 4 are the end results of those dynamics. </p> <p>We found that in a social dilemma situation, cooperation never evolves in the system without punishment. However, with sanction systems by introducing a suitable amount of punishment, while having a suitable findability of the defector, and a suitable initial population structure, cooperation can evolve. These results can be found with this code. We have no legal or ethical concerns regarding this data as this is a numerical analysis based on theoretical equations. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.