Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,505

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,505 results for “Generation”

Learn how ShareScore rates datasets ↗
zenodo56/100

Tropical Pacific SST and wind anomalies generated by a Nonlinear Inverse Model

<p>Tropical Pacific (40S-40N; 120E-50W) sea surface temperature (SST), zonal wind (U) and meridional wind (V) anomalies generated by the Nonlinear Inverse Model described in Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5). The data consists in 99 realizations (<a href="../api/records/10411023/draft/files/NLIM_output_085.nc/content" target="_blank" rel="noopener noreferrer">NLIM_output_XXX.nc</a>) of 1,000yrs each emulating SST, U, and V monthly anomalies conditions during 1980-2020 (<a href="../api/records/10411023/draft/files/Monthly_obs_1980_2020.nc/content" target="_blank" rel="noopener noreferrer">Monthly_obs_1980_2020.nc</a>) given in a 2.5deg-2.5deg grid. For observations, we used the NOAA Extended Reconstruction SST v5 reanalysis (SST; Huang et al., 2017) and NCEP-NCAR reanalysis (winds; Kalnay et al., 1996) The observed anomalies are calculated as described in Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5).</p> <p>Given that the stochastic forcing considered is white in time and space (https://doi.org/10.1038/s41612-024-00675-5; Methods, section "Offline simulation of SSH_{12}, PC2, and spatial patterns fron nonlinear inverse model output"), the spatial patterns and lead-lag relationships are better identified using composites. A modification of the methodology that allows for spatially coherent stochastic forcing will be implemented in a future article.</p> <p>When using the data please cite https://doi.org/10.5281/zenodo.10411023 (the data) and Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5; for the methodology).&nbsp;</p> <p>Any question, please contact Cristian Martinez-Villalobos at his email cristian.martinez.v@uai.cl</p> <p>References</p> <p>Martinez-Villalobos, C., Dewitte, B., Garreaud, R.D.&nbsp;<em>et al.</em>&nbsp;Extreme coastal El Ni&ntilde;o events are tightly linked to the development of the Pacific Meridional Modes.&nbsp;<em>npj Clim Atmos Sci</em>&nbsp;<strong>7</strong>, 123 (2024). https://doi.org/10.1038/s41612-024-00675-5</p> <p>Huang, B. et al. Extended Reconstructed Sea Surface Temperature, Version 5 (ERSSTv5): Upgrades, Validations, and Intercomparisons. Journal of Climate 30, 8179&ndash;8205 (2017).</p> <p>Kalnay, E. et al. The NCEP/NCAR 40-Year Reanalysis Project. Bulletin of the American Meteorological Society 77, 437&ndash;471 (1996).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
OpenNeuro52/100

Reconstructing Faces from fMRI Patterns using Deep Generative Neural Networks.

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo52/100

Data archive for "Stochastic Super-Resolution for Downscaling Time-Evolving Atmospheric Fields with a Generative Adversarial Network"

<p>This datasets supports the paper &quot;Stochastic Super-Resolution for Downscaling Time-Evolving Atmospheric Fields with a Generative Adversarial Network&quot; submitted to IEEE Transactions in Geoscience and Remote Sensing. A preprint of the paper can be found here: <a href="https://arxiv.org/abs/2005.10374">https://arxiv.org/abs/2005.10374</a>. The code that uses these data is available at <a href="https://github.com/jleinonen/downscaling-rnn-gan">https://github.com/jleinonen/downscaling-rnn-gan</a>.</p> <p>The file &quot;goes-samples-2019-128x128.nc&quot; contains the training dataset called &quot;GOES-COT&quot; in the paper, consisting of cloud optical depth measurements from the GOES-16 satellite. The files &quot;gen_weights*.nc&quot; contain the generator weights saved at different time steps during training for the two different datasets described in the paper.<br> &nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo52/100

Generative convective parametrization of a dry atmospheric boundary layer

<p>The repository contains simulation snapshots of a dry convective boundary layer (CBL). The snapshots comprise horizontal snapshots of vertical velocity (w) and buoyancy (b) field at three heights, namely z/h(t) = 0.2, 0.5, 1.0. Further, the Python scripts for the Generative Adversarial Network (GAN) are also provided, as well as the DNS renormalization procedure.</p>

opencc-by-4.0Aug 2023View details →
zenodo52/100

Dataset for "The influence of the amount of recycled material on the microstructure and properties of the second generation of single-domain YBCO bulks"

<p>The development of a recycling process for various REBCO materials is crucial considering both environmental sustainability and economic efficiency, particularly in light of the upcoming large-scale applications. In this paper, a novel general recycling process based on chemical dissolution was employed to grow REBCO bulks; recycled material obtained by recycling defective YBCO single-domain bulks was added (15 wt. %, 30 wt. % and 45 wt. %) to raw materials to prepare recycled YBCO precursor powder. Subsequently, recycled single-domain YBCO bulks were produced using Top-Seeded Melt Growth. The waste recycling related to of single-domain bulks growth was chosen, as it represents the most challenging form of waste in the context of REBCO superconductor production. The properties and microstructure of recycled bulks were further analyzed to determine the influence of the amount of recycled material used and compared to commercially produced bulks. Single-domain YBCO bulks were grown successfully from the recycled precursor powder. Furthermore, it was found that their properties could be tuned by varying the amount of the added recycled powder, allowing the use of vast amounts of REBCO waste for the preparation of bulks, when achieving the best possible properties is not essential for a given application. Given that the underlying recycling process is designed to work for all REBCO systems and any form of waste, it has significant implications for the sustainability and cost-effectiveness of REBCO superconductor production.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Dataset - Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks

<p>This data is complementary to the paper by Leijnse et al. 2022 &quot;Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks&quot;&nbsp;<br> https://doi.org/10.5194/nhess-2021-181</p> <p>This data is made available in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE</p> <p>For questions about the data ask: tim.leijnse@deltares.nl</p> <p>For more information about the tool to generate the used synthetic tracks TCWiSE see:&nbsp;<a href="https://www.deltares.nl/en/software/tcwise/">https://www.deltares.nl/en/software/tcwise/</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo52/100

Subjective human thresholds over computer generated images

<p>Realistic image computation mimics the natural process of acquiring pictures by simulating the physical interactions of light between all the objects, lights and cameras lying within a modelled 3D scene. This process is known as global illumination and was formalised by Kajiya with the following rendering Equation:<br> <span class="math-tex">\(\begin{equation} \label{eq:rendering_equation} L_o(x, \omega_o) = {L_e(x, \omega_o)} + \int_{\Omega}^{} {L_i(x, \omega_i)} \cdot f_r(x, \omega_i \rightarrow \omega_o) \cdot \cos \theta_i d\omega_i \end{equation}\)</span></p> <p>where:</p> <ul> <li>&nbsp;<span class="math-tex">\(L_o(x, \omega_o)\)</span> is the luminance traveling from point&nbsp;<span class="math-tex">\(x\)</span> in direction <span class="math-tex">\(\omega_o\)</span>;</li> <li><span class="math-tex">\(L_e(x, \omega_o)\)</span> is point&nbsp;<span class="math-tex">\(x\)</span> emitted luminance (it is null if point x does not lie on a ligth source surface);</li> <li>the integral represents the set of luminances <span class="math-tex">\(L_i\)</span>incident in <span class="math-tex">\(x \)</span> from the hemisphere of the directions <span class="math-tex">\(\Omega\)</span> and reflected in the direction <span class="math-tex">\(\omega_o\)</span>. The reflected luminances are weighted by the materials reflecting properties (bidirectionnal reflectance function <span class="math-tex">\(f_r(x, \omega_i \rightarrow \omega_o)\)</span>) and the cosinus of the incident angle.</li> </ul> <p>This equation cannot be analytically solved and Monte Carlo approaches are generally used to estimate the value of the pixels of the final image.</p> <p>This proposed dataset is composed of 80 points of view of photo realistics images with different level of samples (following the Monte Carlo approach) for each. Each image is 800 x 800 pixels in size. The most noisy image is of 20 samples and the reference one (the most converged image obtained) is of 10000 samples. The <a href="https://www.pbrt.org/index.html">pbrt</a> rendering engine (version 3) was used to generate these images.</p> <p>By exploiting these levels of samples obtained and therefore of noise perceptible in the images, average subjective human thresholds were collected. For this purpose, the images were divided into 16 areas of 200 x 200 pixels in size for each point of view.</p> <p>The proposed image database is composed of the following files:</p> <ul> <li><strong>human-thresholds.csv</strong> : the set of human subjective thresholds obtained on 40 points of view. A line is composed of the name of the point of view followed by all the thresholds obtained for each of the 16 zones;</li> <li><strong>SIN3D_dataset.tar.gz</strong> : is an archive containing all the images from 20 to 10000 samples in steps of 20 samples for each point of view (i.e. 500 images per point of view). Each folder in the archive corresponds to a point of view.</li> </ul> <p><em>This image database has been exploited in order to propose an objective model for noise detection in photo-realistic computer-generated images (article referenced to this image database).</em></p> <p><strong>Note:</strong> Some of the proposed scenes come from:</p> <ul> <li><a href="https://pbrt.org/scenes-v3">https://pbrt.org/scenes-v3</a></li> <li><a href="https://benedikt-bitterli.me/resources/">https://benedikt-bitterli.me/resources/</a></li> </ul> <p><strong>Funding:</strong> This research was funded by ANR support: project ANR-17-CE38-0009.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2021View details →
zenodo52/100

MUHAI Benchmark : Task 1 (Short story generation with Knowledge Graphs)

<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 1 (Short story generation with Knowledge Graphs and Language Models)</strong>&nbsp;</p> <p>The dataset&nbsp;can be used to test understandability of text generated through the&nbsp;combination of&nbsp;knowledge graphs and language models without using knowledge graph embeddings.<br> <br> The task here is to generate 5-sentence stories from a set of <em>subject-predicate-object</em>&nbsp;triples that are extracted&nbsp;from&nbsp;a knowledge graph. Two steps need to be performed:</p> <p>1. Language model fine-tuning (SVO triple extraction + model fine-tuning)<br> 2. Story generation (knowledge enrichment + text generation)&nbsp;<br> <br> The submission includes the following data:</p> <ol> <li>Original ROC stories corpus (100 stories)</li> <li>ROC stories encoded&nbsp;with relevant triples&nbsp;(extracted through SpaCy, 2 versions, with and without coreference resolution)</li> <li>Stories generated by the pre-trained&nbsp;model (GPT2-simple)</li> <li>Stories generated by the fine-tuned model (DICE + ConceptNet + DBpedia&nbsp;)</li> <li>Stories generated by the fine-tuned model (DICE + ConceptNet + DBpedia + WordNet&nbsp;)</li> <li>Stories generated by the GPT-2-keyword-generation (an open-source software that uses&nbsp;GPT-2 to generate text pertaining to the specified keywords)</li> <li>Model results</li> <li>Evaluation metrics description</li> <li>User-evaluation questionnaire&nbsp;</li> </ol> <p>Code :&nbsp;https://github.com/kmitd/muhai-dice_story</p>

opencc-by-4.0Sep 2022View details →
zenodo52/100

Dataset of "Photoelectrochemical generation of H2O2 using hematite (α-Fe2O3) and gas diffusion electrode (GDE)"

<p>In contrast to the industrial-scale production of H2O2 the electrochemical or photoelectrochemical synthesis is environmentally friendly. In the present work,&nbsp;<br>the photoelectrochemical generation of H2O2 was studied by combining the hematite (&alpha;-Fe2O3/FTO/glass) photoanode and gas diffusion electrode (GDE) modified by&nbsp;<br>incorporation of tin (II) phthalocyanine (SnPc) in its hydrophilic layer. The experiments were carried out in a photoelectrochemical cell with two compartments&nbsp;<br>separated by a proton exchange membrane under applied bias and AM1.5 irradiation (100 mW/cm2). The generated amount of H2O2 was determined by chemical analysis&nbsp;<br>(visible light spectrophotometry) of the electrolyte. As a tool to determine the efficiency of such a process, the Faradaic efficiency (FE) was calculated. The&nbsp;<br>best configuration used air as an inlet gas for GDE and phosphate buffer (pH 6.4) as an electrolyte in the cathodic compartment. The combination of hematite and&nbsp;<br>GDE (with SnPc) was the most effective in H2O2 photoelectrochemical generation. The highest value of FE was 52.4 % for GDE (O2 reduction to H2O2) and 0.4 % for&nbsp;<br>hematite photoanode (H2O oxidation to H2O2).</p>

opencc-by-4.0Aug 2024View details →
zenodo52/100

Four lipidomics datasets (mouse liver, mouse pancreatic islets, mouse soleus muscle and mouse visceral adipose tissue), generated for the publication Mehl et al., "A multiorgan map of metabolic, signalling, and inflammatory pathways that coordinately control fasting glycemia in mice"

<p>Mehl, Thorens et al present a multiomics study aimiing to<span>&nbsp;identify the pathways that are coordinately regulated in pancreatic </span><span>b</span><span>-cells, muscle, liver, and fat to control fasting glycemia we fed C57Bl/6, DBA/2 and Balb/c mice a regular chow or a high fat diet for 3, 10 and 30 days. We measured fasted glycemia, insulinemia and whole-body insulin resistance. Transcriptomic and lipidomic analysis were used in a data fusion approach to identify organ-specific pathways related to the glycemic levels across all conditions investigated. In pancreatic islets, constant insulinemia despite higher glycemic levels were associated with reduced expression of mRNAs encoding hormone and neurotransmitter receptors as well as OXPHOS, cadherins, integrins and gap junction proteins. Higher glycemia and whole-body insulin resistance were associated, in muscle, with reduced expression of mRNAs encoding insulin signaling proteins and enzymes of the glycolysis, Krebs&rsquo; cycle and OXPHOS pathways, as well as endocytosis and exocytosis proteins; in hepatocytes, with lower expression of mRNAs of the insulin signaling pathway, of branched chain amino acid catabolism and of OXPHOS; in adipose tissue, with increased expression of mRNAs of innate immunity and lipid catabolism. These data provide a map of the pathways that are coordinately recruited in the investigated tissues to control fasting glycemia and a resource for further studies of interorgan communication in glucose homeostasis. </span></p>

opencc-by-4.0Sep 2024View details →
edi52/100

Inter- and intra-annual temperature and precipitation variability (1950-2022) across the ranges of non-migratory birds and their association with generation length

While environmental variability is theorized to impact the life history characteristics of organisms, these hypotheses have not been thoroughly tested with empirical data. To fill this gap, we synthesized a global data set of environmental variability metrics and life history characteristics across the ranges of 7,477 non-migratory, non-marine avian species. These data are derived from the ERA5 climate reanalysis, AVONET, BirdTree, and BirdLife databases as well as previously published research. By extracting environmental variability values across individual species' ranges, this data set allows users to evaluate avian species' pace of life in response to environmental change.

openCC (other)Jan 2025View details →
OpenNeuro48/100

BOLD Verb Generation

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo48/100

Molecular datasets from "SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design"

<p>Herein find the molecular datasets from &quot;<a href="https://chemrxiv.org/articles/SMILES-Based_Deep_Generative_Scaffold_Decorator_for_De-Novo_Drug_Design/11638383">SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design</a>&quot;. These were generated with&nbsp;SMILES-based scaffold decorator generative models&nbsp;trained with two training sets (DRD2 and ChEMBL). These generative models require a partially-built molecule (scaffold) as input and output several possible completions for each scaffold. Each dataset corresponds to a model trained with the&nbsp; ChEMBL or DRD2&nbsp;sets, wither multi-step (ms) or single-step (ss) and the provenance of the scaffolds (validation set, or non-dataset).</p> <p>The molecules generated are annotated with a set of descriptors. The DRD2 datasets have the predicted probability of each molecule to be active&nbsp;on DRD2 (p)&nbsp;obtained from a Random Forest model. The ChEMBL model&#39;s descriptors are related to the synthesizability of the molecules (see manuscript). Also, the datasets decorated from validation set scaffolds are annotated whether they are part of the validation set (in_validation).</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Data from: "Deep Generative Modeling of Periodic Variable Stars Using Physical Parameters"

<p>This dataset was used for the training of a conditioned Variational Autoencoder that generates physically informed light curves of periodic variable stars. The light curves correspond to data obtained from The Optical Gravitational Lensing Experiment (<a href="https://ui.adsabs.harvard.edu/abs/1992AcA....42..253U/abstract">OGLE</a>), while ancillary information was obtained from the Gaia Data Release 2 (<a href="https://ui.adsabs.harvard.edu/link_gateway/2016A&amp;A...595A...1G/doi:10.1051/0004-6361/201629272">GAIA DR2</a>). This repository contains the preprocessed OGLE light curves and the GAIA measurements corresponding to each cross-matched source. We also provided a subsample of cross-matched sources that were carefully validated following several steps described in the companion article (paper reference).</p> <p>This dataset is realized in tandem with the corresponding&nbsp;<a href="https://github.com/jorgemarpa/PELS-VAE">GitHub</a>&nbsp;and&nbsp;<a href="https://arxiv.org/abs/2005.07773">article</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Data release for paper "Towards the routine use of subdominant harmonics in gravitational-wave inference: re-analysis of GW190412 with generation X waveform models"

<p>This data release for the paper &quot;Towards the routine use of subdominant harmonics in gravitational-wave inference: re-analysis of GW190412 with generation X waveform models&quot; [<a href="https://arxiv.org/abs/2010.05830">arXiv:2010.2010.05830</a>] contains posterior samples for the GW190412 binary black hole merger event obtained from public GWOSC data with the parallel bilby Bayesian inference package, dynesty nested sampler and a set of waveforms from the &quot;generation X&quot; of phenomenological waveform models: IMRPhenomXAS, IMRPhenomXHM, IMRPhenomXP, IMRPhenomXPHM, IMRPhenomT and IMRPhenomTHM. The provided file is a &quot;meta file&quot; that can be read with the <a href="https://lscsoft.docs.ligo.org/pesummary/">PESummary</a> python package. The posterior samples included correspond to runs [2,6,10,12,14,26] in Table III of the paper (standard settings for each waveform, standar priors and sampler settings of Nlive=2048 and Nact=10 or 50). If you make use of these samples, please cite both this data release and the paper.</p>

opencc-by-4.0Oct 2020View details →
zenodo48/100

TMY hourly generation profiles for Insolight hybrid Si/III-V planar micro-tracking modules in Madrid

<p>Hourly energy density (1 m<sup>2</sup>)&nbsp;generation profiles for Insolight hybrid Si/III-V planar micro-tracking modules installed in Madrid (40.5&deg;N, -3.75&deg;E), synthetically generated using <a href="https://github.com/isi-ies-group/cpvlib">CPVLIB library</a> (based on <a href="https://pvlib-python.readthedocs.io/en/stable/">PVLIB Python</a>) and ERA5 typical meteorological year. Performance model parameters were empirically fitted using several outdoor monitoring campaigns and indoor characterization at the <a href="https://www.ies.upm.es/Investigacion/Research_Lines/Concentrator_photovoltaics/CPV_characterization">collimated-light solar simulator</a> available at IES-UPM.</p> <p><strong>Location</strong>:&nbsp;40.5&deg;N, -3.75&deg;E</p> <p><strong>Format</strong>: CSV (separator: semicolon);&nbsp;headers in first row.</p> <p><strong>Parameters </strong>(ordered from first column):&nbsp;</p> <ul> <li>Time: YYYY-MM-DD HH:MM:SS+TimeZoneOffset</li> <li>Latitude: latitude of the installation in&nbsp;&deg;N</li> <li>Longitude: longitude of the installation in &deg;E</li> <li>Wind speed [m/s]: average wind speed</li> <li>Tair [&deg;C]: average ambient temperature</li> <li>precipitable_water [mm]: average precipitable water in the atmosphere</li> <li>GHI [Wh/m2]: global horizontal irradiation</li> <li>DHI [Wh/m2]: diffuse horizontal irradiation</li> <li>DNI [Wh/m2]: direct (beam) normal irradiation</li> <li>CPV submodule [kWh/m2]: energy generated per m<sup>2</sup>&nbsp;by the III-V CPV submodule</li> <li>Flat-plate submodule [kWh/m2]: energy generated per m<sup>2</sup> by the Si flat-plate submodule</li> <li>Hybrid [kWh/m2]: energy generated per m<sup>2</sup> by the whole Insolight hybrid module</li> </ul>

opencc-by-4.0Jan 2021View details →
zenodo48/100

Compact continuum source finding for next generation radio surveys

<p>This is a data set that accompanies the paper "Compact continuum source finding for next generation radio surveys" (2012MNRAS.422.1812H)</p> <p>The image files and source catalogues contained here were used to test the completeness and false detection rate of a number of source finding algorithms including: Aegean, Selavy, Sfind, SExtractor, and IMSAD. These data can be used to assess the performance of future source finding codes, and to verify the the performance of code during development.</p>

opencc-by-4.0Feb 2012View details →
zenodo48/100

Pythia Generated Jet Images with Alternative Rotation Scheme for Location Aware Generative Adversarial Network Training

<p>Dataset containing 300k jet images that can be used to train Location Aware Generative Adversarial Networks (LAGAN) for High Energy Physics, such as the one in [arXiv:1701.05927].</p> <p><strong>Format</strong>:</p> <p>HDF5 file with the following fields:</p> <ul> <li>'image' : array of dim (300000, 25, 25), contains the pixel intensities of each 25x25 image</li> <li>'signal' : binary array to identify signal (1, i.e. W boson) vs background (0, i.e. QCD)</li> <li>'jet_eta': eta coordinate per jet</li> <li>'jet_phi': phi coordinate per jet</li> <li>'jet_mass': mass per jet</li> <li>'jet_pt': transverse momentum per jet</li> <li>'jet_delta_R': distance between leading and subleading subjets if 2 subjets present, else 0</li> <li>'tau_1', 'tau_2', 'tau_3': substructure variables per jet (a.k.a. n-subjettiness, where n=1, 2, 3)</li> <li>'tau_21': tau<sub>2</sub>/tau<sub>1</sub> per jet</li> <li>'tau_32': tau<sub>3</sub>/tau<sub>2</sub> per jet</li> </ul> <p><strong>Details</strong>:</p> <ul> <li>Simulated using Pythia 8.219 at √ s = 14 TeV</li> <li>Image pre-processing using method from in L. de Oliveira et al., <em>Jet-Images -- Deep Learning Edition </em>[arXiv:1511.05190]</li> <li>scikit-image==0.10.0 implementation of cubic spline rotation with fewer low energy artifacts than scikit-image&gt;=0.12.0</li> <li>Finite calorimeter granularity simulated with 0.1×0.1 grid in η and φ, with η × φ ∈ [−1.25, 1.25] × [−1.25, 1.25]</li> <li>Jet clustering with anti-k<sub>t</sub> algorithm with a radius R = 1.0 using FastJet 3.2.1; constituent re-clustering into R = 0.3 k<sub>t</sub> subjets</li> <li>Intensity of pixel = p<sub>T</sub> of cell</li> <li>60 GeV &lt; m<sup>jet</sup> &lt; 100 GeV</li> <li>250 GeV &lt; p<sub>T</sub><sup>jet</sup> &lt; 300 GeV</li> <li>Sparse images (~10% NNZ)</li> </ul> <p>Full dataset description in [arXiv:1701.05927].</p>

opencc-by-4.0Feb 2017View details →
zenodo48/100

Webis Generated Native Ads 2024

<p>Version of the&nbsp;<a href="https://zenodo.org/records/10802427">Webis Generated Native Ads 2024</a> dataset prepared for Sub-Task 2 of the <a href="https://touche.webis.de/clef25/touche25-web/advertisement-detection.html">Advertisement in Retrieval-Augmented Generation</a> task at Touch&eacute; 2025.</p> <p>The dataset contains the same data but split into JSONL-files and with separate files for responses/sentence pairs and labels.</p> <h2>Citation</h2> <pre>@InProceedings{schmidt:2024,<br> author = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; {Sebastian Schmidt and Ines Zelch and Janek Bevendorff and Benno Stein and Matthias Hagen and Martin Potthast},<br> booktitle = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;{WWW '24: Proceedings of the ACM Web Conference 2024},<br> doi = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;{10.1145/3589335.3651489},<br>&nbsp; publisher = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;{ACM},<br> site = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; {Singapore, Singapore},<br> title = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;{{Detecting Generated Native Ads in Conversational Search}},<br> year = &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2024<br>}</pre>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Dataset for the publication entitled "An exact system of generation for face-milled hypoid gears with uniform depth taper: application to hypoid gear drives with high gear ratio"

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record