Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,505
datasets available to search
ShareScore release 0.7.1
Dataset results
7,505 results for “Generation”
Data generated for the publication of Xie,S., Valente,L., &Etienne,R.S
<p>This repository shows the data for the publication of Xie,S., Valente,L., &Etienne,R.S. Can we ignore trait-dependent colonization and diversification in island biogeography?</p> <p>All files were obtained via computation at University of Groningen Peregrine High Performance Computing Cluster (HPCC).<br> We use the R package DAISIErobustness and the R package DAISIE to generate the data which was analyzed in the paper. The code for these packages is version controlled on GitHub and is freely available in open-source repositories. See the Related Identifiers section for links to relevant archived versions of both these packages.</p>
Diversity-Driven Unit Test Generation (Data Set)
<p>The goal of automated unit test generation tools is to create a set of test cases for the software under test that achieve the highest possible coverage for the selected test quality criteria. The most effective approaches for achieving this goal at the present time use meta-heuristic optimization algorithms to search for new test cases using fitness functions defined on existing sets of test<br> cases and the system under test. Regardless of how their search algorithms are controlled, however, all existing approaches focus on the analysis of exactly one implementation, the software under test, to drive their search processes, which is a limitation on the information they have available. In this paper we investigate whether the practical effectiveness of white box unit test generation tools can be increased by giving them access to multiple, diverse implementations of the functionality under test harvested from widely available Open Source software repositories. After presenting a basic implementation of such an approach, DivGen (Diversity-driven Generation), on top of the leading test generation tool for Java (EvoSuite), we assess the performance of DivGen compared to EvoSuite when applied in its traditional, mono-implementation oriented mode (MonoGen). The results show that while DivGen outperforms MonoGen in 33% of the sampled classes for mutation coverage (+16% higher on average), MonoGen outperforms<br> DivGen in 12.4% of the classes for branch coverage (+10% higher average).</p>
Dataset of "Exposure to airborne SARS-CoV-2 in four hospital wards and ICUs of Cyprus. A detailed study accounting for day-to-day operations and aerosol generating procedures."
<p>The authors highly appreciate being contacted if the data is to be used for any purpose.</p> <p>The following data set was used in the study entitled "<strong>Exposure to airborne </strong><strong>SARS-CoV-2 in four hospital wards and ICUs of Cyprus. A detailed study accounting for day-to-day operations and aerosol generating procedures.</strong>" and published in <em>Heliyon</em> Journal.</p> <p>This study characterized the transmission dynamics of airborne SARS-CoV-2 in normal and intensive care units. The data were collected over the period of 2020. In total, 165 and 62 air and environmental samples, respectively, were collected in four COVID-19 wards and ICUs in Cyprus and analyzed by RT-PCR. The comparison between RT-PCR and an alternative method for SARS-CoV-2 detection in air that provides comparable results but is less cumbersome and time demanding, is also given in the tab "Comparison with BELD".</p> <p>The data from sampling airborne SARS-CoV-2 using a MOUDI impactor are not included in this document but can be found in the supplement of the relevant publication.</p> <p>Please refer to the manuscript and its supplementary material for more information about how the data was collected. </p> <p> </p>
Two metabolomics data sets (mouse kidney, mouse plasma), generated for the publication Bignon et al., 2023: "Multiomics reveals multilevel control of renal and systemic metabolism by the renal tubular circadian clock".
<p><strong>Publication: </strong>Bignon Y, Wigger L, Ansermet C, Weger BD, Lagarrigue S, Centeno G, Durussel F, Götz L, Ibberson M, Pradervand S, Quadroni M, Weger M, Amati F, Gachon F, Firsov D. Multiomics reveals multilevel control of renal and systemic metabolism by the renal tubular circadian clock. J Clin Invest. 2023 Mar 2:e167133. doi: 10.1172/JCI167133. Epub ahead of print. PMID: 36862511.</p> <p> </p> <p><strong>Abstract: </strong> Circadian rhythmicity in renal function suggests rhythmic adaptations in renal metabolism. To decipher the role of the circadian clock in renal metabolism, we studied diurnal changes in renal metabolic pathways using integrated transcriptomic, proteomic, and metabolomic analysis performed on control mice and mice with inducible deletion of the circadian clock regulator Bmal1 in the renal tubule (cKOt). With this unique resource, we demonstrated that ~30% RNAs, ~20% proteins and ~20% metabolites are rhythmic in kidneys of control mice. Several key metabolic pathways including NAD+ biosynthesis, fatty acid transport, carnitine shuttle,and b-oxidation displayed impairments in kidneys of cKOt, resulting in a perturbed mitochondrial activity. Carnitine reabsorption from the primary urine was one of the most impacted processes with a ~50% reduction in plasma carnitine levels and a parallel systemic decrease in tissues carnitine content. This suggests that the circadian clock in the renal tubule controls both kidney and systemic physiology.</p> <p> </p> <p><strong>This record contains two separate mass-spectrometry metabolomics data sets associated with this study:</strong></p> <ol> <li>Metabolic profile of renal tubules, MS/MS data, Metabolon, Morrisville, NC (N=60)</li> <li>Metabolic profile of blood plasma, MS/MS data, Biocrates, Innsbruck, Austria (N=60)</li> </ol> <p>For each data set, original data as received from the platforms and processed data as used in the data analysis are provided. Preprocessing of kidney data included removal of metabolites with more than 80% missing data values, median normalization, imputation and glog2 transformation. Preprocessing of plasma data included filtering of metabolites with any missing data and log2 transformation. Details of data processing are available in the STAR*methods of the publication.</p> <p> </p> <p><strong>Data sets in other repositories associated with the same study:</strong></p> <p>Additional data sets (transcriptomics, proteomics) pertaining to the same study have been deposited in public repositories:</p> <ul> <li>Gene Expression Omnibus (NCBI GEO), GSE216252</li> <li>PRIDE Archive (EMBL-EBI), PXD036803</li> </ul> <p> </p>
Datasets generated by rurAllure project - promotion of rural museums and heritage sites in the vicinity of European pilgrimage routes
<p>These datasets have been generated as part of rurAllure project (funded by the European Union’s Horizon 2020 Research and Innovation programme under grant agreement no 101004887). Main goal of rurAllure is the promotion of rural museums and heritage sites in the vicinity of European pilgrimage routes: https://rurallure.eu/project/about/</p>
Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer
<p>Dataset for our paper "Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer"</p> <p>(<a href="https://github.com/zfj1998/M3NSCT5">zfj1998/M3NSCT5: the code base for our paper "Diverse Title Generation for Stack Overflow Posts with Multiple Sampling Enhanced Transformer" (github.com)</a>)</p> <p>Including three files representing the train/val/test datasets. Each file contains all the collected data covering eight programming languages.</p>
Towards the Future Generation of Railway Localization Exploiting RTK and GNSS
<p>This repository contains the datasets acquired by ETH-PBL in conjunction with Unibo and SADEL during two days of testing in October 2022 near Modena, Italy.</p> <p>The data were acquired using two sensor nodes developed by ETH Zurich running a <a href="https://www.st.com/en/microcontrollers-microprocessors/stm32l452ce.html">STM32L452CEU6</a> MCU.<br> Each node collected data on the motion of the train using an <a href="https://www.st.com/en/mems-and-sensors/asm330lhh.html">ST ASM330LHH</a> automotive grade IMU as well as a <a href="https://www.u-blox.com/en/product/zed-f9p-module">u-blox ZED-F9P</a> GNSS module fed with live RTCM-data from a closeby RTK base station provided by SADEL. The base station utilized another ZED-F9P GNSS module connected to a Raspberry Pi which transmitted the generated RTCM correction packages over a raw TCP socket.<br> The data was then received using a <a href="https://www.u-blox.com/en/product/sara-r4-series">u-blox SARA-R4</a> cellular network module.</p> <p>The track was chosen as it exposes a variety of interesting GNSS environments. Encountered environments are ranging from urban over suburban to open field environments as well as one tunnel. Due to this composition, the availability of cellular connection and thus RTK correction data was patchy but mostly stable.</p> <p>The two sensor nodes were fixed to the Train Chassis, one centered in the train and the other positioned on the left side in driving orientation. Node 1 was placed on the floor in front of the driver's seat and positioned to be aligned with the center of the train in the lateral direction. A TOPGNSS TOP106 L1/L2 multi-band antenna was placed below the rear-facing windscreen also aligned with the same axis. Node 2 was mounted on a window on the left side of the train when facing in the direction of travel. This is approximately 1m above the floor and 1.4m left to the lateral center of the train. An ANN-MB00 L1/L2 antenna was attached to the outside frame of the train above the window.</p> <p>This dataset is linked with the GitHub repository at <a href="https://github.com/ETH-PBL/Railway-Precise-Localization">Railway-Precise-Localization</a> where the data format description and the pre-processing scripts are provided.</p>
Generated Wikidata Subset for Taxons based on dump: 20201102-all
<p>Source file: GeneTaxon_wikidata-20201102-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: GeneTaxon_wikidata-20190121-all
<p>Source file: GeneTaxon_wikidata-20190121-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: GeneTaxon_wikidata-20180115-all
<p>Source file: GeneTaxon_wikidata-20180115-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: 20170821-all
<p>Source file: GeneTaxon_wikidata-20170821-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: wikidata-20150601-all
<p>Source file: GeneTaxon_wikidata-20150601-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: 20160613-all
<p>Source file: GeneTaxon_wikidata-20160613-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 2
<p>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al.</p> <p>The code to use with these data and reproduce the manuscript results is available at https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. param_fixing.zip - self-explanatory (Figure 4 & 5); contains an explanatory note for this part (experiment_details.txt), and the file containing Km values fetched from the BRENDA database (Km_database.csv).</p> <p>2. scripts.zip - scripts to generate figure 2-5 on toy data</p>
Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 1
<p><strong>Supplementary files containing datasets needed to reproduce the results of the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" by S. Choudhury et al (https://doi.org/10.1101/2023.02.21.529387).</strong></p> <p>The code to use with these data and reproduce the manuscript results is available at https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. models.zip - contains thermodynamically curated steady-state and nonlinear kinetic models of <em>E. coli </em>metabolism used in this study. Also contains the samples of steady-state metabolite concentrations and metabolic fluxes used in the study presented in Figure 3 (steady-state samples used for preparing Figures 2 and 4).</p> <p>2. renaissance_incidence_results.zip - self-explanatory (Figure 2a and 2b)</p> <p>3. ODE_solutions.zip - self-explanatory (Figure 2c)</p> <p>4. bioreactor_simulations1-3.zip - self-explanatory (Figure 2d)</p> <p>5. steady_state_analysis.zip - RENAISSANCE results obtained for each of the steady states (Figure 3a)</p> <p>6. subspace_analysis.zip - RENAISSANCE results presented in Figure 3b-g</p> <p><strong>The remaining datasets are published in the following links</strong></p> <p><em> - https://doi.org/10.5281/zenodo.7930084</em></p> <p><em> - https://doi.org/10.5281/zenodo.10391802</em></p>
Generated Wikidata Subset for Taxons based on dump: 20220630-all
<p>Source file: GeneTaxon_wikidata-20220630-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Generated Wikidata Subset for Taxons based on dump: 20210531-all
<p>Source file: GeneTaxon_wikidata-20210531-all.ttl.gz</p> <p>ShEx: https://github.com/kg-subsetting/paper-wikidata-subsetting-2023/blob/master/flexibility-experiments/genes%2Btaxons/GeneTaxon.shex</p> <p>More information: https://www.semantic-web-journal.net/content/wikidata-subsetting-approaches-tools-and-evaluation</p>
Air mass trajectory and connectivity data generated with tropolink (Richard et al., 2023)
<p>Archive containing trajectory and connectivity data generated with tropolink for the preparation of the manuscript Richard et al. (2023, <a href="https://doi.org/10.1029/2023GH000885">https://doi.org/10.1029/2023GH000885</a>), as well as the corresponding specifications (node coordinates, dates and other tropolink options). The archive contains specifications, trajectories and connectivities for the three applications presented in the manuscript:</p><p>- the study of airborne connectivity between areas of production of sugar beet, with starting altitude equal to 250m, 500m and 750m above ground level;</p><p>- the study of airborne connectivity between potyvirus populations;</p><p>- the study of invasion risk of Spodoptera frugiperda in Europe, North Africa and western Asia;</p><p> </p><p>Web application tropolink: https://tropolink.fr/</p><p>Associated gitlab: https://forgemia.inra.fr/tropo-group</p><p>Accompanying wiki: https://forgemia.inra.fr/tropo-group/tropolink/-/wikis</p><p>R code for analyzing tropolink output: https://forgemia.inra.fr/tropo-group/tropolink/-/wikis/Examples</p><p>Richard H., Martinetti D., Lercier D., Fouillat Y., Hadi B., Elkahky M., Ding J., Michel L., Morris C.E., Berthier K., Maupas F., <br>Soubeyrand S. (2023). Computing geographical networks generated by air-mass movement. GeoHealth 7:e2023GH000885. <a href="https://doi.org/10.1029/2023GH000885">https://doi.org/10.1029/2023GH000885</a>.</p>
Pertubation Profiles Dataset used for "Convection-generated gravity waves in the tropical lower stratosphere from Aeolus wind profiling and ERA5 reanalysis"
<p>These are the perturbation profiles, from 5km to 29.5km, with a 500m grid. In the study, we picked up the data between tropopause-1km to 22km, which was then squared, smoothed, and averaged into one value. We used a 14 points moving average for the smoothing.</p> <p>The data is from 2018-09 to 2022-09, based on the Aeolus L2B Rayleigh clear wind, using only quality flag 1 data.</p> <p>Please email me at mathieu.ratynski@estaca.eu if you're interested in the 100m resolution version, used in the final version of the manuscript.</p>
cMSSM parameter space points generated with SPheno and micrOMEGAS
<p>These two datasets were produced to be used in two lectures on Machine Learning for SUSY Model Building taught in <a href="https://indico.cern.ch/event/1214657/">pre-SUSY 2023 summer school</a> in Southampton. The code used to generate and to analyse these data can be found <a href="https://gitlab.com/miguel.romao/ml-for-model-building-susy-2023">here</a>.</p> <p>The datasets are as following:</p> <ul> <li>1 million points generated using SPheno only (so no Dark Matter relic density) for the cMSSM with the physical parameters randomly sampled from the table bellow. The columns are <ul> <li>'m0', 'm12', 'A0', 'tanb': the four physical parameters of the theory</li> <li>'idx': an utility identifier used during generation, can/should be ignored</li> <li>The flattened SPheno outputs. These are obtained by reading the resulting slha spectrum file outputted by SPheno and flatten the blocks. For example from the 'MINPAR' block, the key-value pairs are given by the columns 'MINPAR_1', 'MINPAR_2', 'MINPAR_3', 'MINPAR_4', 'MINPAR_5', and likewise for all blocks in the slha file.</li> </ul> </li> <li>10 thousand points generated using SPheno, and which spectrum outputs was then fed to micrOMEGAS (MSSM model configured to accept low-scale slha files as input), with the physical parameters randomly sampled from the same table bellow. The columns are: <ul> <li>The same as above, in addition to</li> <li> 'Omega', 'dm_spin', 'dm_mass' obtained from the micrOMEGAS output, representing Dark Matter relic density, Dark Matter spin, Dark Matter mass, respectively.</li> </ul> </li> </ul> <p>The full list of columns can be seen in `column_names.txt` file.</p> <p>Versions:</p> <ul> <li>SPheno 4.0.5, with a patch to output a warning when the LSP is charged. This version can be found <a href="https://gitlab.com/lip_ml/blackboxbsm">here</a>.</li> <li>micrOMEGAS 5.3.41, with the MSSM model adapted for low-scale slha inputs.</li> </ul> <p>The datasets are provided in <a href="https://parquet.apache.org/">Apache `parquet`</a> format. In order to read them using `pandas`, an installation with the optional flag `[parquet]` should be used. Alternatively, one can use <a href="https://arrow.apache.org/docs/python/index.html">`pyarrow`</a>.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.