Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Supplementary files for Machine learning for histological annotation and quantification of cortical layers
<div> <h2>Creators</h2> <ul> <li><a href="https://orcid.org/0009-0000-9093-9385">Meystre Julie</a></li> <li><a href="https://orcid.org/0000-0002-7100-3749">Olivier Burri</a></li> </ul> <h2>Contributors</h2> <ul> <li><a href="https://orcid.org/0009-0002-0029-7951">Jean Jacquemier</a></li> </ul> </div> <h2>Description</h2> <p>This dataset contains 7 <a href="https://qupath.github.io/">QuPath</a> projects. The raw data images linked to these projects and located in other Zenodo datasets need to be downloaded as well.</p> <p>The raw data contains images of 14 hemispheres from height animals.</p> <ul> <li>Nissl_1 : <ul> <li>animal 1413827 Right Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_2 : <ul> <li>animal 1413829 Right Hemisphere</li> <li>animal 1413828 Right Hemisphere</li> <li>animal 1413827 Left Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_3 : <ul> <li>animal 1413828 Left Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_4 : <ul> <li>animal 1443459 Right Hemisphere</li> <li>animal 1443460 Right Hemisphere</li> <li> </li> </ul> </li> <li>Nissl_5 : <ul> <li>animal 1443459 Left Hemisphere</li> <li>animal 1443460 Left Hemisphere</li> </ul> </li> </ul> <ul> <li>Nissl_6 : <ul> <li>animal 1449920 Left Hemisphere</li> <li>animal 1449921 Left Hemisphere</li> <li>animal 1449921 Right Hemisphere</li> <li>animal 1449922 Left Hemisphere</li> <li>animal 1449922 Right Hemisphere</li> <li> </li> </ul> </li> <li>QuPath_LayerBoundaries_GroundTruth_20220927: <ul> <li>This is the QuPath project that contains S1HL layers annotations done by the experts and which have been used to trained the Random forest Machine Learning method for the S1HL brain classification. It contains some images from all the eight animals.</li> </ul> </li> </ul> <p> </p> <div> <h3>Animals</h3> <p>All animal procedures were approved by the Veterinary Authorities and the Cantonal Commission for Animal Experimentation of the Canton of Vaud, according to the Swiss animal protection laws, under license number VD3516.</p> <p>Outbred Wistar Han rats (Janvier Laboratories, France) were ordered with their litter aged eight postnatal days (P8). Dams were housed individually and allowed to raise their own litters until experimentation on male offspring aged fourteen days (P14; N=8 animals; N=3 litters). Animals were housed in standard plastic laboratory cages, with bedding, nesting material and paper tube and ad libitum access to food (SAFE 150 SP-25) and water, cleaned once per week, and kept on a twelve-hour light-dark schedule with lights turned on at 06:30 AM, in rooms under controlled humidity and temperature. The sample size here is greater than those reported in other open source atlases <a href="https://www.zotero.org/google-docs/?1dkN18">(“Allen Reference Atlas - Mouse,” n.d.; “The Rat Brain in Stereotaxic Coordinates - 7th Edition,” n.d.)</a>.</p> <h3>Sample preparation</h3> <p>On postnatal day fourteen, rats were transferred to the experimental room in the morning to acclimate. The described procedure was conducted within a consistent 3-hour window of the day (09:00-12:00). Initially, the rats were deeply anesthetized using pentobarbital (intraperitoneal dose of 150 mg/kg; concentration of 150 mg/ml). This was succeeded by transcardial perfusion with ice cold 0.1 M phosphate buffer (PB; pH 7.4), followed by cold 4% paraformaldehyde (PFA) in 0.1 M PB. Subsequently, the brain was carefully removed from the skull, postfixed at 4°C in 4% PFA overnight, and then rinsed in 0.1 M PB. The brains underwent a sequential storage process: first in a 15% sucrose solution (in 0.1 M PB) at 4°C for approximately 24 hours, followed by a 30% sucrose solution at 4°C for an additional 24 hours. The hemispheres were carefully divided along the midline, after which both right and left hemispheres were precisely sliced sagittally using a cryostat (Leica, VT-1200S) at 50 µm employing an approximate angle rotation of 4 ± 1 degrees along the anterior-posterior axis to optimize alignment with apical dendrites. These brain slices were stored in a cryoprotectant solution (30% v/v ethylene glycol; 30% m/v sucrose in 0.1 M PB) at -20°C, preserving them until immunohistochemistry assays were executed (within a maximum of two weeks from extraction to immunohistochemistry).</p> <p>In order to determine the cell densities in P14 rat, brain slices were immunostained using cresyl violet, a stain specifically targeting cell bodies, including the endoplasmic reticulum, also known as Nissl substance or Nissl bodies. Free-floating sections of 50 µm thickness were transferred from cryoprotectant into 0.1 M PB to thaw and eliminate any cryoprotectant remnants. Subsequently, they were transferred into 0.01 M PB to minimize salt residues before being meticulously mounted onto SuperFrost© glass slides (Thermo Fisher Scientific Inc., Gerhard Menzel B.V. & Co. KG, GE). This mounting was carried out while considering the brain’s orientation relative to the midline, from its external to internal regions. Slide-mounted sections were processed using an automated slide stainer Tissue-Tek® Prisma Plus (Sakura Finetek-Europe, NL). These sections were incubated for 6 minutes at room temperature (RT = 20°C) in a 0.5% cresyl violet solution in water (with pH adjusted to 2.85 using acetic acid), followed by a brief wash in tap water. The sections underwent dehydration through a series of ethanol concentrations (70%, 70%, 96%, 100%, 100%) with each step lasting one minute at RT. Subsequently cleared with two steps of xylene for one minute each at RT, and the sections were mounted using Pertex (Sakura Finetek-Europe, NL) before being cover-slipped using the automated glass coverslipper Tissue-Tek® Glas™ g2 (Sakura Finetek-Europe, NL). A meticulous assessment of the coloration was conducted and if the staining appeared faint, a repeat staining procedure was carried out.</p> <p><strong>Immunostained slides were scanned using an automated slide scanner (Olympus, VS120-L100, GER) equipped with a UPLSAPO 20x/0.75 air objective (Olympus, GER) and a Pike F505 Color camera leading to a pixel size of 0.346 μm/pixel. Each brain slice was entirely scanned. Subsequently, the digital images obtained were meticulously organized and subjected to analysis using the open-source software QuPath v0.3.2 <a href="https://www.zotero.org/google-docs/?jnVnIg">(Bankhead et al., 2017)</a>. </strong></p> <p> </p> <h2>Intructions</h2> <p>The projects contained in this dataset have been created with QuPath v0.3.2 but could be opened with new QuPath version.</p> <ol> <li>Download the dataset</li> <li>untar the tar balls included in this dataset</li> <li>install <a href="https://qupath.github.io">QuPath</a></li> <li>Open QuPath</li> <li>Open a project within QuPath (Files->Project...->Open Project...)</li> </ol> </div>
R-code for publication: Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods
<p>This is the R-code as well as the underlying data needed to reproduce the results of the springer book chapter: "Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods"</p> <p>For more information contact: dlieske@mta.ca</p> <p> </p>
Global derived datasets for use in k-NN machine learning prediction of global seafloor total organic carbon
<p>This dataset includes 663 predictor grids used for k-NN global prediction of seafloor total organic carbon.</p> <p>663 predictor grids available in netCDF4 HDF5 file format. Grids are cell-centered sized 4320 x 2160. File names adhere to the naming conventions discussed below. The naming structure is partioned by underscores and periods in the following order: interface to which the gridded values refer to, quantity of values contained within the grid, units and reference values/units (e.g. meters below sea level), data source, statistic calculated (if applicable), grid pitch, and file extension.</p> <p>Possible interfaces from the top – down:</p> <p>SS – Sea surface – atmosphere interface (may also be average of the entire water column)</p> <p>SF – Seafloor – water interface (may also be denoted by GL)</p> <p>GL – Ground level (e.g. bottom of pure liquid, top of dirt)</p> <p>SC – Sediment – crust interface (e.g. sediment above, igneous/metamorphic below)</p> <p>CM – Crust – mantle interface (e.g. Mohorovicic discontinuity)</p> <p>Appropriate reference naming marker (bold), original data source, and date of last access:</p> <p><strong>Becker</strong></p> <p>Becker, J. J., Wood, W. T., & Martin, K. M. (2014). <em>Global crustal heat flow using random decision forest prediction</em>, Abstract NG31A-3788 presented at 2014 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 06/23/2015.</p> <p><strong>CRUST1</strong> </p> <p>Pasyanos, M.E., Masters, G., Laske, G. & Ma, Z. (2012). <em>LITHO1.0 - An Updated Crust and Lithospheric Model of the Earth Developed Using Multiple Data Constraints</em>, Abstract T11D-09 presented at 2012 Fall Meeting, AGU, San Francisco, California, U.S.A. Last access: 07/01/2014.</p> <p><strong>CRUST1_NOAA</strong></p> <p> As the NOAA sediment thickness database is globally not complete, data gaps in the NOAA grid with this have been supplemented by the CRUST1 sediment thickness (see above citation).</p> <p>Whittaker, J., Goncharov, A., Williams, S., Müller, R. D., & Leitchenkov, G. (2013) Global sediment thickness dataset updated for the Australian-Antarctic Southern Ocean, <em>Geochemistry, Geophysics, Geosystems. </em>https://doi.org/10.1002/ggge.2018.<em> </em>Last access: 09/02/2018.</p> <p><strong>GVP</strong></p> <p>Global Volcanism Program (2013) Volcanoes of the World. In E. Venzke (ed.). (Vol. 4.7.3). Smithsonian Institution. https://doi.org/10.5479/si.GVP.VOTW4-2013. Last access: 09/22/2014.</p> <p><strong>ETOPO2v2</strong></p> <p>National Geophysical Data Center (2006). 2-minute Gridded Global Relief Data (ETOPO2) v2. National Geophysical Data Center, NOAA. DOI: 10.7289/V5J1012Q. Last access: 02/06/2013.</p> <p><strong>PLATES</strong></p> <p>Coffin, M.F., Gahagan, L.M., & Lawver, L.A. (1998). Present-day Plate Boundary Digital Data Compilation. University of Texas Institute for Geophysics Technical Report (No. 174, pp. 5). Last access: 09/15/2014.</p> <p><strong>ONRL</strong></p> <p>Ludwig,W., Amiotte-Suchet, P., & Probst, J. L. (2011). ISLSCP II Global River Fluxes of Carbon and Sediments to the Oceans. In F. G. Hall, G. Collatz, B. Meeson, S. Los, E. Brown de Colstoun, and D. Landis (Eds.), <em>ISLSCP Initiative II Collection</em>. Oak Ridge National Laboratory Distributed Active Archive Center, Oak Ridge, Tennessee, U.S.A. http://dx.doi.org/10.3334/ORNLDAAC/1028. Last Access: 02/15/2015.</p> <p><strong>Muller</strong></p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust, <em>Geochemistry, Geophysics, Geosystems</em>, 9(4), Q04006. https://doi.org/10.1029/2007GC001743. Last accessed: 07/19/2011.</p> <p><strong>Woa13x</strong></p> <p>Boyer, T.P., Antonov, J. I., Baranova, O. K., Coleman, C., Garcia, H. E., Grodsky, A., et al. (2013) World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), <em>NOAA Atlas NESDIS 72, Technical Ed</em>. Silver Spring, MD. http://doi.org/10.7289/V5NZ85MT. Last Access: 09/18/2014.</p> <p><strong>KIM</strong></p> <p>Kim, S.S. & Wessel, P. (2011). New global seamount census from the altimetry-derived gravity data, <em>Geophysical Journal International</em>, 186, 615-631. https://doi.org/10.1111/j.1365-246X.2011.05076.x. Last access: 09/22/2014.</p> <p><strong>HYCOM</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data.Last access: 03/19/2014.</p> <p><strong>NCEDC</strong></p> <p>NCEDC (2016). Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC. Last access: 09/21/2014.</p> <p><strong>Wei2010</strong></p> <p>Wei, C.-L., Rowe, G. T., Escobar-Briones, E., Boetius, A., Soltwedel, T., Caley, M. J., et al.(2010). Global patterns and predictions of seafloor biomass using random forests. <em>PLoS ONE</em>,5(12), e15323. https://doi.org/10.1371/journal.pone.0015323 Last access: 06/20/2016.</p> <p><strong>NGA_egm2008</strong></p> <p>Pavlis, N.K., Holmes, S. A., Kenyon, S. C., & Factor, J. K. (2008). <em>The</em> <em>EGM2008 Global Gravitational Model</em>, Abstract 2008AGUFM.G22A..01P presented at the 2008 General Assembly of the European Geosciences Union, Vienna, Austria. Last access: 07/10/2014.</p> <p><strong>WAVEWATCH3</strong></p> <p>The 1/12 deg global HYCOM+NCODA Ocean Reanalysis was funded by the U.S. Navy and the Modeling and Simulation Coordination Office. Computer time was made available by the DoD High Performance Computing Modernization Program. The output is publicly available at https://hycom.org/publications/acknowledgements/ocean-reanalysis-data. Last access: 03/19/2014.</p> <p>Updated global seafloor porosity grid using our k-nearest neighbors algorithm using 5 nearest neighbors. Observed data used for prediction from Martin et al. (2015). </p> <p>Martin, K. M., Wood, W. T., & Becker, J. J. (2015). A global prediction of seafloor sediment porosity using machine learning. <em>Geophysical Research Letters</em>, 42(24), 10640. https://doi.org/10.1002/2015GL065279</p> <p>Other grids which have been generated by empirical means are latitude (and derivatives), longitude (and derivatives), Coriolis, coast_is_1.0, and the random noise grids. </p> <p>Units referenced are as follows:</p> <p>KGM3 - kilogram per cubic meter<br> MS - meters per second<br> KM - kilometer<br> M_ASL - meters above sea level (i.e. meters referenced to sea level)<br> MWM2 - milliwatt per square meter<br> TGCYR - terragram of carbon per year<br> TGYR - terragram per year<br> MA - megaannum<br> M - meters<br> MGCM2 - milligram of carbon per square meter<br> DEG - degree<br> S - seconds</p> <p>Statistics grids are calculated within a given radius (e.g. 10km, 50km, 125km, 250km, 500km, 1000km) of the respective cell-centered value. The statistics grids include mean (.men), average absolute deviation from the mean (.aad), and the common logarithm (.log) of the absolute value of the mean (.mlg). Additionally, some grids are a weighted count for given radii (e.g. seamounts) where weight is a cosine taper from the center of the grid cell. </p> <p>The grid pitch for this dataset is uniformly at 5-arc minute denoted by “.5m”. Additionally, the extension used (netCDF4) is denoted by “.nc”.</p>
Use of EO data to test machine learning algorithms
<p>This data set is composed of one Sentinel-2 image of Darwin City, Australia. The objective of this work is to use EO data to compare the performance of different machine learning algorithms.</p>
Dataset and code for "Classification of Solar Wind With Machine Learning"
<p>Matlab software and data from http://www.mlspaceweather.org/ for the paper</p> <p>https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017JA024383</p>
Modeling Autonomic Pupillary Responses from External Stimuli Using Machine Learning - Dataset
<p>This page contains the data collected for the paper: <em><strong>Modeling Autonomic Pupillary Responses from External Stimuli Using Machine Learning </strong></em>(<a href="https://doi.org/10.26717/BJSTR.2019.20.003446">DOI:10.26717/BJSTR.2019.20.003446</a>). The dataset consists of spectral and pupillometric data collected during three outdoor/indoor walks. The folders “raw”, “merged”, and “cleaned” contain data collected by the Konica Minolta CL-500A Illuminance Spectrophotometer and Tobii Pro Glasses 2 at three different stages in the data preparation process. The “raw” folder contains uncleaned and unsynchronized .csv/.json files. The “merged” folder contains uncleaned, but synchronized light and ocular data in .csv format. The “cleaned” folder contains a single .csv of cleaned and synchronized data with the derived variables: average pupil diameter and pupil diameter difference. </p> <p>The best choice of data files will depend on desired analysis. More guidance on how to handle this data can be found in the readMe files located in each subsequent folder. More information on the sensing devices used here can be found at the Minolta and Tobii information links below. </p> <p><strong>Minolta Information</strong>: <a href="https://sensing.konicaminolta.us/uploads/cl-500a_instruction217a_eng-250cl60686.pdf">https://sensing.konicaminolta.us/uploads/cl-500a_instruction217a_eng-250cl60686.pdf</a></p> <p><strong>Tobii Information</strong>: <a href="https://www.tobiipro.com/siteassets/tobii-pro/user-manuals/tobii-pro-glasses-2-user-manual.pdf/?v=1.1.3">https://www.tobiipro.com/siteassets/tobii-pro/user-manuals/tobii-pro-glasses-2-user-manual.pdf/?v=1.1.3</a></p> <p>The codes used to prepare, analyze, and visualize this data is available in the LightOcular GitHub Repository linked below. </p> <p><strong>LightOcular GitHub Repo</strong>: <a href="https://github.com/mi3nts/LightOcular">https://github.com/mi3nts/LightOcular</a></p>
Nonlocal machine-learned exchange functional for molecules and solids
<p>This dataset supplements the journal article "Nonlocal machine-learned exchange functional for molecules and solids," published in Physical Review B: <a href="https://doi.org/10.1103/PhysRevB.110.075130">DOI:10.1103/PhysRevB.110.075130</a>. It contains the machine-learned functionals developed in the study (which can be used with the <a href="https://github.com/mir-group/CiderPressLite">CiderPressLite code</a>), along with additional data and results from the study.</p> <p>Please see the README.md file for more details on the dataset, and refer to the original paper for details on methodology and funding acknowledgments.</p>
CREMP-CycPeptMPDB: Conformer-rotamer ensembles of macrocyclic peptides for machine learning with permeability annotations
<p>CREMP-CycPeptMPDB: A resource generated for the rapid development and evaluation of machine learning models for permeable macrocyclic peptides. CREMP-CycPeptMPDB contains 3,258 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 8.7 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations and with experimental membrane permeability measurements obtained from the <a href="http://cycpeptmpdb.com/" target="_blank" rel="noopener">CycPeptMPDB</a> database. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>This dataset complements the <a title="CREMP" href="../doi/10.5281/zenodo.7931444" target="_blank" rel="noopener">CREMP dataset</a>, which contains a larger selection of conformer ensembles for homodetic macrocyclic peptides.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the “pickle” folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single “pickle.tar.gz” archive. In the “sdf_and_json” folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, “sdf_and_json.tar.bz2”. A single summary CSV file is also provided containing ”sequence”, “smiles”, “num_monomers”, “num_atoms”, “num_heavy_atoms”, along with the CREST metadata “totalconfs”, “uniqueconfs”, “lowestenergy”, “poplowestpct”, “temperature”, “ensembleenergy”, “ensembleentropy”, and “ensemblefreeenergy”. The number of unique conformers with different 3D structures is given by “uniqueconfs”, while “totalconfs” includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 13 GB for "pickle.tar.gz" and 84 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>
Training Machine-Learned Density Functionals on Band Gaps
<p>This dataset contains atomic structures and molecular orbitals (computed with DFT using the PBE functional) for the systems studied in the paper "<a href="https://doi.org/10.1021/acs.jctc.4c00999">Training Machine-Learned Density Functionals on Band Gaps</a>." The <a href="https://github.com/mir-group/CiderPress">CiderPress code</a> can be used to analyze the data.</p> <p>Please see the README.md file for more details on the dataset, and refer to the original paper for details on methodology and funding acknowledgments.</p>
Data For Scalco et al. Clinicopathological correlates of quantitative Amyloid-B Pathology in the Temporal Cortex: Machine learning analysis of 131 cases from an ADRC
<p>Dataset containing 131 de-identified whole slide images (WSIs) with a respective data dictionary. </p> <p><strong>Paper</strong>: Scalco, R., Oliveira, L.C., Lai, Z. et al. Machine learning quantification of Amyloid-β deposits in the temporal lobe of 131 brain bank cases. acta neuropathol commun 12, 134 (2024). https://doi.org/10.1186/s40478-024-01827-7</p> <p><strong>Details</strong>: A total of 131 .svs. WSIs, de-identified using svs-deidentifier v 0.9.1-beta (https://github.com/pearcetm/svs-deidentifier/releases). Dataset is uploaded in batches due to Zenodo data upload limitations.</p> <p><strong>Slide curation/preparation</strong>: All samples were retrieved from archives of the University of California, Davis Alzheimer’s Disease Center Brain Bank (<a href="https://www.ucdmc.ucdavis.edu/alzheimers/">https://www.ucdmc.ucdavis.edu/alzheimers/</a>). Archival samples analyzed in this study were 5 μm formalin fixed, paraffin embedded sections of the superior and middle temporal gyrus from human brain. The tissue had been previously stained with an amyloid-β antibody (4G8, recognizing residues 17-24, BioLegend, formerly Covance) that were first pretreated with formic acid to rid samples of endogenous protein. All slides were digitized using an Aperio AT2 between 20x and 40x magnification.</p> <p><strong>Code:</strong> Please refer to <a href="https://github.com/ucdrubinet/BrainSec">https://github.com/ucdrubinet/BrainSec</a> and <a href="https://github.com/keiserlab/plaquebox-paper">https://github.com/keiserlab/plaquebox-paper</a></p>
Dataset of "Comprehensive Machine Learning Approaches for Modelling the State of Charge of Lithium-ion Batteries"
<p>This paper evaluates three ML approaches for SOC modeling in LIBs: the multilayer perceptron (MLP), long short-term memory (LSTM), and the nonlinear autoregressive with exogenous input (NARX) neural network architectures. These models were tested using an experimental dataset with multiple input variables, including electrochemical impedance spectroscopy (EIS) data, voltage, and capacity readings for commercial LIB cells. Results indicate that MLP and LSTM are more adaptable with a smaller training dataset (14 samples), while the NARX model required more than 34 out of 67 samples to achieve reasonable accuracy. Additionally, the NARX model is more sensitive to changes in the learning rate (α) and exhibits larger output error deviations. The MLP and LSTM models consistently performed well across various hidden layer sizes, showing no upper bound constraints, whereas the NARX model’s performance deteriorated with certain hidden layer configurations.</p>
Mapping a novel metric for Flash Flood Recovery using Interpretable Machine Learning
<p>This data is supplementary to the paper titled "Mapping a novel metric for Flash Flood Recovery using Interpretable Machine Learning". The file contains the main results.<br><br>For any queries, please visit <a href="https://hydrosense.iitd.ac.in" target="_blank" rel="noopener">Hydrosense Lab (IIT Delhi)</a>.</p>
GPC/m: Global Precipitation Climatology by Machine Learning; Quasi-global, Daily, and One Degree Spatial Resolution
<p>A precipitation dataset, Global Precipitation Climatology by Machine Learning (ML), GPC/m, is released.</p> <p>This new precipitation dataset has been produced by machine learning, which is daily from 1979 to 2020 (will be to present), 1° × 1° spatial resolution. Three ML methods are used. Data is produced from outgoing longwave radiation (OLR) and atmospheric circulation from reanalysis. You can download this with DOI.</p> <p>This daily precipitation dataset has been produced by machine learning (ML) methods using satellite observations and atmospheric circulations from reanalysis. The quasi-global daily precipitation dataset has been around for 42 years from 1979 to 2020, which will be updated to the present. The spatial resolution is 1° × 1° zonally global and from 40°S to 50°N. The ML methods are supervised learning, and the reference data are estimated precipitation datasets from 2001 to the present. The input data are somewhat modified based on knowledge of the climatological background. Using the trained statistical models, we predict back to 1979, when daily precipitation data was almost unavailable globally. For now, this GPC/m precipitation dataset version is GPC/m-v1-2024. This data will be updated in the future with added value. The purpose of this dataset is a challenge to produce a climatological dataset by reducing artificial gaps as much as possible for discussion of climatology, climate variability, and climate change. This dataset is very useful for statistical analysis, such as composite analysis and correlation analysis. Disadvantages should also be understood in the description paper (Takahashi, 2024c). Also, I hope that this dataset can contribute to improving the current precipitation datasets, which are based on physical or researcher-explaining algorithms.<br><br>To facilitate analysis of the dataset, it is distributed in Network Common Data Form (netCDF) format and the Grid Analysis and Display System (GrADS) format (with control file). If you would like recently updated data, please contact the creator. If it has already been created, it can be distributed.<br><br><em>Added on September 18, 2024.</em><br>More details are in the preprint paper at this link (<a href="https://doi.org/10.48550/arXiv.2409.09639">Takahashi, 2024, https://doi.org/10.48550/arXiv.2409.09639</a>).</p> <p><em>Added on March 4, 2025.</em><br><strong>Alternative Download Options</strong><br>If you experience slow download speeds from Zenodo, alternative mirrors are available for the dataset files.<br><em><span>However, we kindly request you to download the .ctl file from Zenodo for tracking purposes.</span></em><br>Download NetCDF (.nc) or Binary (.bin) from:<br><a href="https://camo.fpark.tmu.ac.jp/gpcm.html">https://camo.fpark.tmu.ac.jp/gpcm.html</a></p>
WaivOps EDM-HSE: Open Audio Resources for Machine Learning in Music
<p><strong>EDM-HSE Dataset</strong></p> <p>EDM-HSE is an open audio dataset containing a collection of code-generated drum recordings in the style of modern electronic house music. It includes 8,000 audio loops recorded in uncompressed stereo WAV format, created using custom audio samples and a MIDI drum dataset. The dataset also comes with paired JSON files containing MIDI note numbers (pitch) and tempo data, intended for supervised training of generative AI audio models.</p> <p><strong>Overview</strong></p> <p>The EDM-HSE Dataset was developed using an algorithmic framework to generate probable drum notations commonly played by EDM music producers. For supervised training with labeled data, a variational mixing technique was applied to the rendered audio files. This method systematically includes or excludes drum notes, assisting the model in recognizing patterns and relationships between drum instruments, thereby enhancing its generalization capabilities.</p> <p>The primary purpose of this dataset is to provide accessible content for machine learning applications in music and audio. Potential use cases include generative music, feature extraction, tempo detection, audio classification, rhythm analysis, drum synthesis, music information retrieval (MIR), sound design and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>8,000 audio loops (approximately 17 hours)</li> <li>16-bit WAV format</li> <li>Tempo range: 120–130 BPM</li> <li>Paired label data (WAV + JSON)</li> <li>Variational drum patterns</li> <li>Subgenre styles (Big room, electro, minimal, classic)</li> </ul> <p>A JSON file is provided for referencing and converting MIDI note numbers to text labels. You can update the text labels to suit your preferences.</p> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The EDM-HSE dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>Please note that this dataset has not been fully reviewed and may contain minor notational errors or audio defects.</p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-EDM-HSE">GitHub repository</a>.</p>
Replication package for paper: Insights on the Use of Software Design Principles in Machine Learning Pipelines
<p>This is the replication package of the paper "Insights on the Use of Software Design Principles in Machine Learning Pipelines".</p> <p>This replication package contains two files:</p> <ul> <li><a href="../api/records/13828806/draft/files/Data%20extraction.xlsx/content" target="_blank" rel="noopener noreferrer">Data extraction.xlsx</a>: file containing the details of the extracted data for each single ML project. </li> <li><a href="../api/records/13828806/draft/files/Source%20Code%20and%20Metadata.zip/content" target="_blank" rel="noopener noreferrer">Source Code and Metadata.zip</a>: zip file including the source code local copy analyzed and the repository metadata (.json) provided by GitHub API for each ML project .repository </li> </ul> <p>Reference: [1] Lidia López, Cristina Gómez, and Claudia Ayala. Insights on the Use of Software Design Principles in Machine Learning Pipelines. <em>Accepted </em>in the 2024 edition of the International Conference on Product-Focused Software Process Improvement (PROFES 2024).</p> <p><strong>Note</strong>: The licence is applicable to the excel file. "Source Code and Metadata.zip" file contains source code repositories downloaded from GitHub, the license for each repository is defined in the corresponding GitHub repository by their authors.</p>
Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data
<h2><strong>Sub-dataset: WRB soil types probabilities (part 1)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>
Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data
<h2><strong>Sub-dataset: WRB soil types probabilities (part 2)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>
Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data
<h2><strong>Sub-dataset: WRB soil types classification and relative entropy</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>
Research data for "Intermediates of Forming Transition Metal Dichalcogenides Heterostructures Revealed by Machine Learning Simulations"
<p>This dataset supports the paper "Intermediates of Forming Transition Metal Dichalcogenides Heterostructures Revealed by Machine Learning Simulations". </p> <p><strong>Included Files:</strong></p> <ul> <li><strong>ocp_active.zip</strong>: Modified version of ocp (https://github.com/Open-Catalyst-Project/ocp) tailored for active learning applications.</li> <li><strong>deployed.pth</strong>: Pre-trained model used in the experiments.</li> <li><strong>chemiscopy_run.py</strong>: Script integrating the chemiscopy and nequip modules, designed for dataset visualization.</li> <li><strong>new_energy.py</strong>: Modified version of the nequip module, featuring a repulsive potential function.</li> <li><strong>test_datasets.extxyz</strong> & <strong>train_datasets.extxyz</strong>: The test and training datasets in extxyz format.</li> </ul> <p>How to use the modified version of the nequip module:</p> <p>To train this version of the potential function, it is recommended to use nequip<=0.5.6 (on Linux). The NequIP training files need to be updated as follows:</p> <pre><code>model_builders: - new_energy.EnergyModel - StressForceOutput min_bond_len: 1.8</code></pre> <p>Then run:</p> <p><code>export PYTHONPATH=${PYTHONPATH}:$PWD</code><br><code>nequip-train config.yml # Train the potential function</code><br><code>nequip-deploy build --train-dir nequipresultsdir build.pth # Deploy the trained model</code></p>
WetCH4: A Machine Learning-based Upscaling of Methane Fluxes of Northern Wetlands during 2016-2022
<p>This dataset (WetCH<sub>4</sub>) contains methane (CH<sub>4</sub>) emissions using three different wetland maps, their uncertainties, and underlying flux intensities from northern wetlands (>45° N). The dataset is a data-driven upscaling product using observations from northern eddy covariance CH<sub>4</sub> flux sites and random forest machine learning. WetCH<sub>4</sub> provides daily CH<sub>4</sub> fluxes of northern wetlands at 10-km resolution from 2016 to 2022 and can be used to study regional CH<sub>4</sub> budgets and wetland responses to climate change. The data products are provided in netCDF format files (.nc) with more details in the attributes of the files.</p> <p>File list:</p> <p>- fch4_nmol_m2_s_10km_intensity.nc.gz and fch4_nmol_m2_s_10km_uncertainty.nc.gz:</p> <p> The underlying flux intensities and associated uncertainties.</p> <p> </p> <p>- fch4_10km_emi_wad2m.nc.gz and fch4_10km_emi_uncertainty_wad2m.nc.gz:</p> <p> Upscaled CH4 emissions and uncertainties using WAD2M monthly dynamic wetland map.</p> <p> </p> <p>- fch4_10km_emi_giems2.nc.gz and fch4_10km_emi_uncertainty_giems2.nc.gz:</p> <p> Upscaled CH4 emissions and uncertainties using GIEMS2 monthly dynamic wetland map.</p> <p> </p> <p>- fch4_10km_emi_glwd.nc.gz and fch4_10km_emi_uncertainty_glwd.nc.gz:</p> <p> Upscaled CH4 emissions and uncertainties using static GLWD v1 wetland map.</p> <p> </p> <p>Time range: 2016-01-01 - 2022-12-31</p> <p>Time steps: daily, 2557</p> <p>Geographic extent: longitude 180W - 180E, latitude 45 - 90 N</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.