Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

73

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

73 results for “artificial dataset”

Learn how ShareScore rates datasets ↗
zenodo56/100

Improving Artificial Teachers by Considering How People Learn and Forget: Dataset

<p>This dataset contains the results of the experiment described in&nbsp;<a href="https://dl.acm.org/doi/10.1145/3397481.3450696">Nioche et al. (2021)</a>.&nbsp;</p> <p>This&nbsp;dataset contains 4&nbsp;data files:</p> <ul> <li><em>data.csv</em>: the main data file.</li> <li><em>stimuli.csv:</em> the description/listing of the stimuli.</li> <li><em>demographic_info.csv</em>: the demographic information about the users.</li> <li><em>data_incl_preliminary_exp.csv</em>: an additional data file that includes the user of the preliminary experiments</li> </ul> <p>The main data file contains the logs of&nbsp;53 different users using a self-teaching application for one week. The goal of the users&nbsp;was to learn the English meaning of Japanese kanji. Each user completed between 1370 trials and 1608 trials. Each user saw between 85 and 204 characters.&nbsp;</p> <p>Two additional files are also joint to the data files:</p> <ul> <li><em>info.ipynb</em>: A Jupyter notebook that provides&nbsp;information about each data file, a few descriptive plots,&nbsp;and an example of data manipulation.</li> <li><em>info.pdf: </em>A pdf rendering of the notebook.</li> </ul> <p>If you use this dataset, please refer to it by citing&nbsp;<a href="https://dl.acm.org/doi/10.1145/3397481.3450696">Nioche et al. (2021)</a>.</p>

opencc-by-4.0Apr 2021View details →
zenodo48/100

A living catalogue of artificial intelligence datasets and benchmarks for medical decision making

<p>We provide&nbsp;a comprehensive curated catalogue of&nbsp;<strong>artificial intelligence datasets</strong> and <strong>benchmarks for medical decision making</strong>. At the time of first release (April 2021), the dataset contains more than 400&nbsp;biomedical and clinical datasets&nbsp;of which 252 are publicly available or available upon request.</p> <p>The dataset was compiled based on a systematic literature review covering both biomedical and computer science literature and&nbsp;grey literature data sources. All datasets were manually systematized and annotated for meta-information, such as:</p> <ul> <li>Availability and licensing information</li> <li>Type of source data</li> <li>Links to source publications, main references or dataset repositories</li> </ul> <p>Benchmark dataset were additionally annotated for the following information:</p> <ul> <li>Associated task</li> <li>Performance metrics commonly used for evaluation</li> <li>Clinical relevance</li> <li>The availability of data splits</li> </ul> <p>In addition to the versioned TSV file on Zenodo, the dataset can also be explored live via&nbsp;<a href="https://docs.google.com/spreadsheets/d/1QjUxxnZ3tuyW5dj6nkt_o5yJcWUZec4ttfJxO8Zlty4/edit?usp=sharing">this Google Spreadsheet</a>.&nbsp;The dataset is intended as a living, extendable resource. Edit suggestions and additions are encouraged and can be submitted via the comment function of the Google sheet.</p> <p>&nbsp;</p> <p><strong>File descriptions</strong></p> <p><em>annotated-datasets.tsv</em> -- contains the annotated datasets</p> <p><em>arXiv-literature-export.tsv</em> -- contains the original literature record export from arXiv</p> <p><em>pubmed-literature-export.tsv</em> -- contains the original literature record export from PubMed</p> <p><em>README.md</em> -- contains a detailed description of all annotation fields</p>

opencc-by-sa-4.0Apr 2021View details →
zenodo44/100

Artificial Neural Networks-generated Dataset: pH, Total Alkalinity, and Hydrogen Ion Concentration in Ría de Vigo (NW Spain), 1995–2020

<p>This dataset comprises input data from INTECMAR and the predicted outcomes. The variables and their units are as follows:</p> <p>station: 'Station ID [1-6]'</p> <p>year: 'Year [1995-2020]'</p> <p>month: 'Month [1-12]'</p> <p>day: 'Day'</p> <p>latitude: 'Latitude (decimal degrees)'</p> <p>longitude: 'Longitude (decimal degrees)'</p> <p>depth: 'Depth (meters)'</p> <p>temperature: 'Temperature (degrees Celsius)'</p> <p>salinity: 'Salinity (psu)'</p> <p>phosphate: 'Phosphate (umol/kg)'</p> <p>nitrate: 'Nitrate (umol/kg)'</p> <p>silicate: 'Silicate (umol/kg)'</p> <p>cweek: 'Cosine week'</p> <p>sweek: 'Sine week'</p> <p>TA: 'Total Alkalinity predicted (umol/kg)'</p> <p>NTA: 'Normalized Total Alkalinity (umol/kg)'</p> <p>NAT_st: 'Normalized per station Total Alkalinity (umol/kg)'</p> <p>NTA_gl: 'Normalized globally Total Alkalinity (umol/kg)'</p> <p>pHTS_insitu: 'pH insitu (pH units)'</p> <p>HT: 'Hydrogen ion concentration predicted (nmol/kg)'</p> <p>&nbsp;</p> <p>The authors gratefully acknowledge the financial support by the Programa de axudas &aacute; etapa predoutoral da Xunta de Galicia (Axencia Galega de Innovaci&oacute;n) (Grant n&ordm; IN606A-2022/025). F.F.P. and A.V. were supported by REDEIRA (TED2021-132188B-I00) project, funded by MCIN/AEI/10.13039/501100011033. The authors also express their gratitude to the Instituto Tecnol&oacute;xico para o Control do Medio Mari&ntilde;o de Galicia (INTECMAR), for the analyses and production of the database used to make predictions.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Artificial fingerprints engraved through block-copolymers as nanoscale physical unclonable functions for authentication and identification - Dataset

<p>This is the dataset of "Artificial fingerprints engraved through block-copolymers as nanoscale physical unclonable functions for authentication and identification" by Irdi Murataj, Chiara Magosso, Stefano Carignano, Matteo Fretto, Federico Ferrarese Lupi, and Gianluca Milano, Nature Communications (2024), DOI: 10.1038/s41467-024-54492-8</p> <p>Part of this was funded by the project MEMQuD, code 20FUN06. The project has received funding from the EMPIR program co-financed by the Participating States and from the European Union's Horizon 2020 research and innovation program.</p> <p>Part of this work was supported by the European project OpMetBat, code 21GRD01. The project has received funding from the European Partnership on Metrology, cofinanced from the the European Union's Horizon Europe Research and Innovation Programme, and by Participating States.</p> <p>Part of this work was supported by the European Union - Next Generation EU under the National Recovery and Resilience Plan (NRRP), Mission 04 Component 2 Investment 3.1 | Project Code: IR0000027 - CUP: B33C22000710006 - iENTRANCE@ENL: Infrastructure for Energy TRAnsition aNd Circular Economy @EuroNanoLab.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Dataset for: Pre-pandemic artificial MERS analog of polyfunctional SARS-CoV-2 S1/S2 furin cleavage site domain is unique among spike proteins of genus Betacoronavirus

<table> <tbody> <tr> <th>&nbsp;</th> <td> <div> <h3><strong>Data File Descriptions and Methods</strong></h3> <ol> <li><strong>Data file 1 [betacov_matching_IPR042578.fasta]</strong>: Representative set of 2,465 betacoronavirus S protein overlapping homologous superfamily sequences retrieved in fasta format on 4 December 2022 from the InterPro repository at https://www.ebi.ac.uk/interpro/entry/InterPro/IPR042578/.<br><br></li> <li><strong>Data File 2 [betacov_matching_IPR042578_motif.fasta]</strong>: With Data File 1 as input, extracted 98,122 furin cleavage site (FCS) output motifs of 20 amino acids length, including overlapping and redundant sequences, produced with the FindFur algorithm with preset parameters as described by (Gu, 2020). FindFur as used was deposited on 15 December 2020 at the GitHub software repository at https://github.com/chwisteeng/FindFur.<br><br></li> <li><strong>Data File 3 [table_s1s2_hits_betacov_polyf.pdf]</strong>: Compiled summary table of sequence hits (PDF) of spike S1/S2 domains across genus&nbsp;<em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></li> <li> <p><strong>Data File 4 [table_s1s2_hits_betacov_polyf.xlsx]</strong>: Compiled summary table of sequence hits (MS Excel) of spike S1/S2 domains across genus&nbsp;<em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></p> </li> <li> <p><strong>Data File 5 [betacov_s1s2_nls_pat7_furin_psort.txt]:&nbsp;</strong>Nuclear localization signal (NLS) detection output for 5 representative betacoronavirus spike sequence domains, including the positive hits for pat7 in SARS-CoV-2 and for MERS-MA30 CoV. NLS predictions used the PSORT algorithm available as a webservice at https://wolfpsort.hgc.jp/ which is based on the work of Nakai and Horton (Nakai and Horton, 1999). Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 6 [betacov_s1s2_oglyc_netogly.txt]:&nbsp;</strong>Detection output for 5 representative betacoronavirus spike sequence domains tested for Thr/Ser O-glycosite residue pairs with the standard prediction software NetOGlyc4.0 (Steentoft et al., 2013) as available at https://services.healthtech.dtu.dk/services/NetOGlyc-4.0/. Positive hits have scores above 0.5. Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 7 [betacov_s1s2_nls_pat7_furin_blastp.txt]</strong>: Comprehensive sequence database searches using were performed using the NCBI protein BLAST (blastp) algorithm with webservice available at https://blast.ncbi.nlm.nih.gov/Blast.cgi?PAGE=Proteins. The following blastp search parameters and settings were used: Word size=2; Expect value=200000; Hitlist size=500; Gapcosts=9,1; Matrix=PAM30; Filter string=F; Genetic Code=1;Window Size=40; Threshold=11; Composition-based stats=0; Database Posted date=Jan 19, 2023 2:59 AM; Number of letters=17,117,563; Number of sequences=10,766; Entrez query: Includes: Betacoronavirus (taxid:694002); Excludes: SARS-CoV-2 (taxid:2697049). The six polyfunctional input query consensus motif sequences were TXXPR(K/H/R)XRSX and TXXPRX(K/H/R)RSX.</p> </li> </ol> <h3><strong>References</strong></h3> <p>Gu, C., 2020. FindFur: A Tool for Predicting Furin Cleavage Sites of Viral Envelope Substrates. Master&rsquo;s Thesis, San Jose State University, CA, USA. doi: <a href="https://doi.org/10.31979/etd.4ahv-9jya">10.31979/etd.4ahv-9jya</a>&nbsp;</p> <p>Gangavarapu K, Latif AA, Mullen JL, Alkuzweny M, Hufbauer E, Tsueng G, Haag E, Zeller M, Aceves CM, Zaiets K, Cano M, Zhou X, Qian Z, Sattler R, Matteson NL, Levy JI, Lee RTC, Freitas L, Maurer-Stroh S; GISAID Core and Curation Team; Suchard MA, Wu C, Su AI, Andersen KG, Hughes LD. Outbreak.info genomic reports: scalable and dynamic surveillance of SARS-CoV-2 variants and mutations. Nat Methods. 2023. 20(4):512-522. doi: <a href="https://doi.org/10.1038/s41592-023-01769-3">10.1038/s41592-023-01769-3</a>.</p> <p>Nakai, K., Horton, P., 1999. PSORT: a program for detecting sorting signals in proteins and predicting their subcellular localization. Trends Biochem Sci 24, 34&ndash;36. doi: <a href="https://doi.org/10.1016/s0968-0004(98)01336-x">10.1016/s0968-0004(98)01336-x</a></p> <p>Steentoft, C., Vakhrushev, S.Y., Joshi, H.J., Kong, Y., Vester-Christensen, M.B., Schjoldager, K.T.-B.G., Lavrsen, K., Dabelsteen, S., Pedersen, N.B., Marcos-Silva, L., Gupta, R., Bennett, E.P., Mandel, U., Brunak, S., Wandall, H.H., Levery, S.B., Clausen, H., 2013. Precision mapping of the human O-GalNAc glycoproteome through SimpleCell technology. EMBO J 32, 1478&ndash;1488.&nbsp;doi: <a href="https://doi.org/10.1038/emboj.2013.79">10.1038/emboj.2013.79</a></p> </div> </td> </tr> </tbody> </table>

opencc-by-4.0Jul 2024View details →
zenodo44/100

On the Use of Artificially Degraded Manuscripts for Quality Assessment of Readability Enhancement Methods - Dataset & Code

<p>This object contains the dataset and python code used for the paper:</p> <p>S. Brenner and R. Sablatnig. On the Use of Artificially Degraded Manuscripts for Quality Assessment of Readability Enhancement Methods. Accepted for OAGM Workshop&nbsp; 2019<strong>, </strong>Steyr, Austria.</p> <p>The dataset is a modified subset of the UCL Multispectral Processed Images of Parchment Damage Dataset (<a href="http://dx.doi.org/10.14324/000.ds.1469099">10.14324/000.ds.1469099</a>). The accompanying code documents how the modified version was created and how the evaluations described in the paper were performed.</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Canada's national artificial intelligence governance system: Dataset from interviews with 20 government leaders & subject matter experts

<p><strong>Summary</strong></p> <p>Anonymized aggregate data from interviews with 20 government leaders and subject matter experts. The data was collected as part of a study of Canada's national system of artificial intelligence governance. The data was collected from February 2023 to July 2023. The dataset contains 610 topics that emerged from thematic analysis of interview transcripts from July 2023 to October 2023. The contexts, actors, resources, networks, evaluations, logics, functional bounds, rules, ecosystem-level dynamics, opportunities for improvement, and other topics contained in the dataset collectively represent the most significant components of Canada's national AI governance system that emerged over the course of the interviews with the 20 participants.</p> <p>&nbsp;</p> <p><strong>Notes for interpreting this dataset</strong></p> <p>Topics in analytical dimensions 1, 3, and 6-11 contain counts of the frequency with which aggregate topics emerged across each of the interviews with the 20 participants. Topics in analytical dimensions 2, 4, and 5 contain categories instead of frequency counts: the topics in these dimensions represent every unique actor, resource, and network that emerged over the course of the interviews instead of aggregate topics.&nbsp;</p> <p>Column titles contain the following abbreviations:<br>LEAD: Interviews with leaders of public sector AI governance initiatives.<br>SME-PS: Interviews with subject matter experts employed in the private sector.<br>SME-CS: Interviews with subject matter experts employed in the academic or civil sectors.</p> <p>&nbsp;</p> <p><strong>Full report</strong></p> <p>A report containing more information about this dataset and about the findings of our study can be found on SSRN: <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4783525">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4783525</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Archival Datasets for SuperNova Artificial Inference by Lstm neural networks (SNAIL)

<p>The spectral-observation dataset (enclosed in the file&nbsp;archival_spec_observations.tar.gz)&nbsp;is comprised of 3091 observed spectra from 361 SNe Ia,&nbsp;largely contributed from CfA (Blondin et al. 2012), BSNIP (Silverman et al. 2012), CSP (Folatelli et al. 2013) and Supernova Polarimetry Program (Wang &amp; Wheeler 2008; Cikota et al. 2019a; Yang et al. 2020).</p> <p>The spectral-template dataset (enclosed in the file&nbsp;archival_spec_templates.tar.gz)&nbsp;includes&nbsp;361 spectral templates, each of them (covering -15 to +33d with wavelength from 3800 to 7200 A)&nbsp;was generated from the available spectroscopic observations of an individual SN via a LSTM neural network model.</p> <p>The&nbsp;auxiliary photometry&nbsp;dataset&nbsp;(enclosed in the file&nbsp;archival_phot_observations.tar.gz) provides&nbsp;the B &amp; V light curves of these SNe (in total, 196 available&nbsp;SNe Ia), that&nbsp;were&nbsp;used to calibrate the synthetic B-V color of the observed spectra.</p> <p>In additional, the two master catalogs give the detailed information about the 361 SNe and their spectroscopic observations, respectively.&nbsp;</p> <p>These datasets are&nbsp;associated to the paper &quot;Spectroscopic Studies of Type Ia Supernovae Using LSTM Neural Networks&quot;&nbsp;(Hu et al. 2022, ApJ, accepted).</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

DeepBacs – Artificial labeling of E. coli membranes dataset and fnet/CARE models

<p>Training and test images of <em>E. coli </em>cells for artificial labeling of membranes in brightfield images using fnet or CARE, as well as trained models for prediction of super-resolution membranes.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>Example image shows an <em>E. coli</em> bright field image and PAINT membrane image predicted by the neural network (scale bar is 1 &micro;m).</p> <p>&nbsp;</p> <p><strong>Training and testing dataset</strong></p> <p><strong>Data type</strong>: Paired bright field and super-resolution images</p> <p><strong>Microscopy data type</strong>: Bright field and fluorescence microscopy (widefield and point accumulation for imaging in nanoscale topography (PAINT) images)</p> <p><strong>Microscope</strong>: Nikon Eclipse Ti-E equipped with an Apo TIRF 1.49NA 100x oil immersion objective</p> <p><strong>Cell type</strong>: <em>E. coli </em>K12 strain derivatives</p> <p><strong>File format</strong>: .tif (8-bit)</p> <p><strong>Image size</strong>: 512x512 px<sup>2</sup> with different pixel sizes:</p> <p>1x tube lens: 158 nm (raw) and 19.75 nm (8x upscaled for PAINT images)</p> <p>1.5x tube lens: 106 nm (raw) (widefield fluorescence only)</p> <p>&nbsp;</p> <p><strong>fnet model (PAINT membrane images)</strong></p> <p>The fnet 2D model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained for 200,000 steps on 33 paired images (image dimensions: (512 x 512 px&sup2;), patch size: (128 x 128 px&sup2;)) with a batch size of 4, a learning rate of 0.0004, 10% validation split and 4x data augmentation (flipping and rotation).</p> <p>Model weights can be used with the ZeroCostDL4Mic fnet 2D notebook.</p> <p>&nbsp;</p> <p><strong>CARE model (PAINT membrane images):</strong></p> <p>The CARE 2D model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained for 300 epochs (100 steps/epoch) on 33 paired images (image dimensions: 512 x 512 px&sup2;, patch size: 256 x 256 px&sup2;) with a batch size of 4, a learning rate of 0.0004, 90/10% train/validation split and 4x data augmentation (flipping and rotation).</p> <p>Model weights can be used with the ZeroCostDL4Mic CARE 2D notebook or the CSBDeep Fiji plugin.</p> <p><br> <strong>Author(s)</strong>: Christoph Spahn<sup>1,2</sup>, Mike Heilemann<sup>1,3</sup></p> <p><strong>Contact email</strong>: christoph.spahn@mpi-marburg.mpg.de</p> <p>&nbsp;</p> <p><strong>Affiliation(s)</strong>:&nbsp;</p> <p>1) Institute of Physical and Theoretical Chemistry, Max-von-Laue Str. 7, Goethe-University Frankfurt, 60439 Frankfurt, Germany</p> <p>2) ORCID: 0000-0001-9886-2263&nbsp;</p> <p>3) ORCID: 0000-0002-9821-3578</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

AAM: Artificial Audio Multitracks Dataset

<p>This dataset contains 3,000 artificial music audio tracks with rich annotations. It is based on real instrument samples and generated by algorithmic composition with respect to music theory.</p> <p>It provides full mixes of the songs as well as single instrument tracks. The midis used for generation are also available. The annotation files include: Onsets, Pitches, Instruments, Keys, Tempos, Segments, Melody instrument,&nbsp;Beats, and Chords.</p> <p>A <strong>presentation</strong> <strong>paper</strong> was published open-access in <a href="https://doi.org/10.1186/s13636-023-00278-7">EURASIP Journal on Audio, Speech, and Music Processing</a>.</p> <p>Current development<strong> </strong>and source code of the <strong>generator tool</strong> can be found on <a href="https://github.com/fabianostermann/ArtificialSongGenerator">GitHub</a>.</p> <p>For a <strong>tiny version</strong> for demonstration and testing purposes see: <a href="https://doi.org/10.5281/zenodo.6771120">zenodo.6771120</a></p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Artificial Intelligence: Professional reference dataset of Artificial Intelligence professional competences analysis based on the job market

<p>Artificial Intelligence&nbsp;vacancies collection to support FAIRsFAIR Artificial Intelligence&nbsp;Professional Competences<br> <br> This dataset is provided as validation and support for the analysis of Artificial Intelligence&nbsp;competences.<br> The dataset includes a collection of vacancies from the job application<br> website&nbsp;<a href="http://indeed.com/">indeed.com</a>&nbsp;that responded to the search term &quot;Artificial Intelligence&quot;.</p> <p>The used search term could be easily adjusted in the provided code at&nbsp;<a href="https://github.com/atomcracker/Competence_analysis.git">Github Repository</a>. The heavy extensive research analysis is reflected in graphs, described and reflected in&nbsp;<a href="https://scripties.uba.uva.nl/search?id=727184">Thesis</a>.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

CZI (Carl Zeiss Image) dataset with artificial test camera images with various dimension for testing libraries reading

<p>Set of CZI test images created by using a simulated microscope with a test grayscale camera (no LSM or AiryScan or RGB). The filename indicates the used dimension(s)&nbsp;for the acquisition experiment. The files can be used to test the basic functionality of libraries reading CZI files.</p> <p>Examples:</p> <ul> <li>S=2_T=3_CH=1.czi = 2 Scenes, 3 TimePoints and 1 Channel <ul> <li>Z-Stack <strong>was not</strong> activated inside acquisition experiment</li> </ul> </li> <li>S=2_T=3_Z=5_CH=2.czi = 2 Scenes, 3 TimePoints, 5-Z-Planes and 1 Channels <ul> <li>Z-Stack <strong>was </strong>activated inside acquisition experiment</li> </ul> </li> </ul> <p>The test files (so far) contain not any data with more &quot;advanced&quot; dimensions&nbsp;like AiryScan rawdata, illumination angles etc.&nbsp;Also no CZI files with&nbsp;pixel type RGB are included yet.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Global sea surface dimethyl sulfide dataset simulated by artificial neural network

<p>This dataset contains (1) the matched and binned data used for constructing an artificial neural network (ANN) model to simulate the sea surface concentration of dimethyl sulfide (DMS); (2) the simulated global daily sea surface concentrations of DMS ranging from 2005 to 2014 by ANN model and the calculated total transfer velocities (Kt) and sea-to-air fluxes; (3) the simulated global monthly sea surface concentrations of DMS ranging from 2005 to 2100 by ANN model and CMIP6 ensemble and the calculated Kt and sea-to-air fluxes; (4) the yearly mean DMS concentration of each grid in different sensitivity experiments exploring the roles different variables play in driving DMS future changes. The input variables of this ANN model include chlorophyll <em>a</em>, sea surface temperature (SST), mixed layer depth (MLD), nitrate, phosphate, silicate, dissolved oxygen (DO), downward short-wave radiation (DSWF), and sea surface salinity (SSS). The future projections (2015-2100) are subjected into two Shared Socioeconomic Pathway scenarios SSP2-4.5 and SSP5-8.5. The spatial resolution of the simulated dataset is 1&deg;&times;1&deg;. The units of DMS concentration, Kt, and flux are nmol L<sup>&ndash;1</sup>, m s<sup>&ndash;1</sup>, and &mu;mol S m<sup>&ndash;2</sup> d <sup>&ndash;1</sup>, respectively.</p> <p>Compared with the previous version (v1.0), this version is based on an updated ANN model after adjusting the data match-up between satellite and in-situ chlorophyll <em>a</em> for ANN training. In addition, the historical simulation based on CMIP6 only covers the time period from 2005 to 2014, which was from 1850 to 2014 for v1.0.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

CASM: A long-term Consistent Artificial-intelligence based Soil Moisture dataset based on machine learning and remote sensing

<p>Paper to cite:&nbsp;Skulovich, O., Gentine, P. A Long-term Consistent Artificial Intelligence and Remote Sensing-based Soil Moisture Dataset.&nbsp;<em>Sci Data</em>&nbsp;10, 154 (2023). https://doi.org/10.1038/s41597-023-02053-x</p> <p>&nbsp;</p> <p>The Consistent Artificial Intelligence (AI)-based Soil Moisture (CASM) dataset is a global, consistent, and long-term, remote sensing soil moisture (SM) dataset created using machine learning. It is based on the NASA Soil Moisture Active Passive (SMAP) satellite mission SM data as a target and is aimed at extrapolating SMAP-like quality SM data back in time with previous satellite microwave platforms. Machine learning approach, such as neural network (NN) has the advantage of being both nonlinear, and state-dependent, and naturally imposing a global distribution matching between the source and the target data. Utilizing this, the new CASM dataset was created using high-quality SMAP SM as a target and Soil Moisture and Ocean Salinity (SMOS) or Advanced Microwave Scanning Radiometer - Earth Observing System (AMSR-E/2) brightness temperature as a source, which allowed extrapolating SM data 13 years back from before SMAP mission launch. CASM represents SM in the top soil layer, defined on a global 25 km EASE-2 grid and covers 2002-2020 with a 3-day temporal resolution. The resulting dataset exhibits excellent spatial and temporal homogeneity, without compromising the interannual variability, and is in excellent agreement with the SMAP data (with a mean correlation of 0.97 between the SMAP and CASM SM for the period when the two overlap). Moreover, the input and target datasets were divided into seasonal cycle and residuals, with the NN trained on the residuals. This approach ensures that the high performance does not mask a simple seasonal cycle matching but rather exemplifies the skill targeted at&nbsp;predicting extremes; with the NN achieving a correlation of 0.75 on the test data for the residuals. Comparison to 367 global in-situ SM monitoring sites shows a SMAP-like median correlation of 0.66 between station SM and CASM SM from the corresponding grid cell. Additionally, the SM product uncertainty was assessed, and both aleatoric and epistemic uncertainties were estimated and included in the dataset. Mean epistemic uncertainty, related to the NN model structure, ranges from 0.007 m<sup>3</sup>/m<sup>3</sup>&nbsp;to 0.014 m<sup>3</sup>/m<sup>3</sup>&nbsp;and on average is close to a desired SM product stability threshold of 0.01 m<sup>3</sup>/m<sup>3</sup>&nbsp;per year. Aleatoric uncertainty, defined as input noise propagated through the system, depends on the introduced level of noise. With 10% noise applied to the residuals, the resulting mean standard deviation of the model outputs rises from 0.005 to 0.007 m<sup>3</sup>/m<sup>3</sup>. &nbsp;&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Estimation of axial loads in tie-rods: Dataset generated from Finite Element simulations for training Artificial Neural Network

<p>Dataset employed for training the Artificial Neural Networks (ANNs) presented in the cited journal article. The trained ANNs were used to estimate the tensile force in tie-rods installed in a historical structure (the church of the monastery of Sant Cugat close to Barcelona) from dynamic parameters obtained from vibration testing.</p> <p>The dataset consists of input-otput data generated using finite element (FE) simulations. A blank column has been used to separate input data from output data.</p> <p>More details on the nature of the data and how it was employed can be found in the following journal article, which is supplemented by this upload:<br> <em><strong>Makoond N, Pel&agrave; L, Molins C. Robust estimation of axial loads sustained by tie-rods in historical structures using Artificial Neural Networks.&nbsp;Structural Health Monitoring. 2022;0(0). doi:</strong></em><strong><a href="https://doi.org/10.1177/14759217221123326">10.1177/14759217221123326</a></strong></p> <p><a href="https://www.researchgate.net/publication/364098652_Robust_estimation_of_axial_loads_sustained_by_tie-rods_in_historical_structures_using_Artificial_Neural_Networks">Link to author&#39;s version of accepted manuscript</a></p> <p>This work was supported by the Servei del Patrimoni Arquitect&ograve;nic of the Generalitat de Catalunya through a project (managed by the City Council of Sant Cugat) aimed at monitoring the church of the Monastery of Sant Cugat (grant number C-10764). Financial support is also acknowledged from&nbsp;the Ministry of Science, Innovation and Universities of the Spanish Government and the ERDF (European Regional Development Fund) through the SEVERUS project (Multilevel evaluation of seismic vulnerability and risk mitigation of masonry buildings in resilient historical urban centres) (grant number RTI2018-099589-B-I00).</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Automated Segmentation of Large Image Datasets using Artificial Intelligence for Microstructure Characterisation and Damage Analysis

<p>Many properties of commonly used materials are driven by their microstructure, which can be influenced<br>by the composition and manufacturing processes. To optimise future materials, understanding the<br>microstructure is critically important. Here, we present two novel approaches based on artificial intelligence<br>that allow the segmentation of the phases of a microstructure for which simple numerical approaches, such<br>as thresholding, are not applicable: One is based on the nnU-Net neural network, and the other on generative<br>adversarial networks (GAN).<br>Using scanning electron microscopy images collected from large areas (~1 mm&sup2;) of dual-phase steels as a<br>case study, we demonstrate how both methods effectively segment intricate microstructural details,<br>including martensite, ferrite, and damage sites, for subsequent analysis.<br>Either method shows substantial generalizability across a range of image sizes and conditions, including<br>heat-treated microstructures with different phase configurations. The nnU-Net excels in mapping large<br>image areas. Conversely, the GAN-based method performs reliably on smaller images, providing greater<br>step-by-step control and flexibility over the segmentation process.<br>This study highlights the benefits of segmented microstructural data for various purposes, such as<br>calculating phase fractions, modelling material behaviour through finite element simulation, and<br>conducting geometrical analyses of damage sites and the local properties of their surrounding<br>microstructure.</p> <p>https://doi.org/10.1016/j.matdes.2024.113031</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Identifying Episodes of Hypovigilance in Intensive Care Units Using Routine Physiological Parameters and Artificial Intelligence: a Derivation Study. Open Code and Dataset

<p>The purpose of this project is to detect hypogilance using the EVEILS database.</p> <p>Database is ICU data from H&ocirc;tel-Dieu De L&eacute;vis , Qu&eacute;bec, Canada. Please cite us if you use either the data or code.&nbsp;</p> <p>This code was written during Rapha&euml;lle Gigu&egrave;re Msc in Computer Science. The goal of her project is to detect hypovigilance using machine learning in the ICU. In this repository, you have the data set before preprocessing:</p> <ul> <li>df_hypovigilance : Contains the hours, date and value of the vigilance level, using either the RASS or Ramsay and already converted using the thresholds shown in the paper.</li> <li>raw_df : Contains the raw values from the gateway for each participant. All of the identifying values have been removed.</li> </ul> <p>At the end of the preprocessing_anonymous script, you should generate a new dataset called "df_final". This dataset is used for the training_model script.</p> <p>The cross validation employs groups of random size meaning the results might differ from time to time but should stay consistent.</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Dataset: Global X Robotics & Artificial Intelligence ETF (BOTZ) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Global X Artificial Intelligence & Technology ETF (AIQ) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Dataset: Global X Artificial Intelligence & Technology ETF (AIQ) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record