Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,307

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,307 results for “libraries”

Learn how ShareScore rates datasets ↗
zenodo56/100

ClostriTof microflex Biotyper library plugin and associated raw Maldi spectra version 2.0

<p>This dataset contains the ClostriTof microflex Biotyper library&nbsp;plugin, an installation guide as well as the raw spectral data for all library and validation strains used to construct the ClostriTof library plugin.</p> <p>If you use this library for your research, please cite Asare et al., Frontiers in Microbiology, 2023; <a href="https://doi.org/10.3389/fmicb.2023.1104707">https://doi.org/10.3389/fmicb.2023.1104707</a></p> <p>We would like to thank Thomas Maier for his help with assembling version 2.0 of the ClostriTOF Database.</p>

opencc-by-4.0Mar 2023View details →
zenodo52/100

13/1 Sferamundi di Grecia. Prima parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/1 Sferamundi di Grecia. Prima parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

ELKI Multi-View Clustering Data Sets Based on the Amsterdam Library of Object Images (ALOI)

<p>These data sets were originally created for the following publications:</p> <p><em>M. E. Houle, H.-P. Kriegel, P. Kr&ouml;ger, E. Schubert, A. Zimek</em><br> <strong>Can Shared-Neighbor Distances Defeat the Curse of Dimensionality?</strong><br> In Proceedings of the 22nd International Conference on Scientific and Statistical Database Management (SSDBM), Heidelberg, Germany, 2010.</p> <p><em>H.-P. Kriegel, E. Schubert, A. Zimek</em><br> <strong>Evaluation of Multiple Clustering Solutions</strong><br> In 2nd MultiClust Workshop: Discovering, Summarizing and Using Multiple Clusterings Held in Conjunction with ECML PKDD 2011, Athens, Greece, 2011.</p> <p>The outlier data set versions were introduced in:</p> <p><em>E. Schubert, R. Wojdanowski, A. Zimek, H.-P. Kriegel</em><br> <strong>On Evaluation of Outlier Rankings and Outlier Scores</strong><br> In Proceedings of the 12th SIAM International Conference on Data Mining (SDM), Anaheim, CA, 2012.</p> <p>&nbsp;</p> <p>They are derived from the original image data available at <a href="https://aloi.science.uva.nl/">https://aloi.science.uva.nl/</a></p> <p>The image acquisition process is documented in the original ALOI work: <em>J. M. Geusebroek, G. J. Burghouts, and A. W. M. Smeulders</em>, <strong>The Amsterdam library of object images</strong>, Int. J. Comput. Vision, 61(1), 103-112, January, 2005</p> <p>Additional information is available at: <a href="https://elki-project.github.io/datasets/multi_view">https://elki-project.github.io/datasets/multi_view</a></p> <p>The following views are currently available:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>Object number</td> <td>Sparse 1000 dimensional vectors that give the <em>true</em> object assignment</td> <td><a href="6355684/files/objs.arff.gz">objs.arff.gz</a></td> </tr> <tr> <td>RGB color histograms</td> <td>Standard RGB color histograms (uniform binning)</td> <td><a href="6355684/files/aloi-8d.csv.gz">aloi-8d.csv.gz</a> <a href="6355684/files/aloi-27d.csv.gz">aloi-27d.csv.gz</a> <a href="6355684/files/aloi-64d.csv.gz">aloi-64d.csv.gz</a> <a href="6355684/files/aloi-125d.csv.gz">aloi-125d.csv.gz</a> <a href="6355684/files/aloi-216d.csv.gz">aloi-216d.csv.gz</a> <a href="6355684/files/aloi-343d.csv.gz">aloi-343d.csv.gz</a> <a href="6355684/files/aloi-512d.csv.gz">aloi-512d.csv.gz</a> <a href="6355684/files/aloi-729d.csv.gz">aloi-729d.csv.gz</a> <a href="6355684/files/aloi-1000d.csv.gz">aloi-1000d.csv.gz</a></td> </tr> <tr> <td>HSV color histograms</td> <td>Standard HSV/HSB color histograms in various binnings</td> <td><a href="6355684/files/aloi-hsb-2x2x2.csv.gz">aloi-hsb-2x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-3x3x3.csv.gz">aloi-hsb-3x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-4x4x4.csv.gz">aloi-hsb-4x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-5x5x5.csv.gz">aloi-hsb-5x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-6x6x6.csv.gz">aloi-hsb-6x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-7x7x7.csv.gz">aloi-hsb-7x7x7.csv.gz</a> <a href="6355684/files/aloi-hsb-7x2x2.csv.gz">aloi-hsb-7x2x2.csv.gz</a> <a href="6355684/files/aloi-hsb-7x3x3.csv.gz">aloi-hsb-7x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-14x3x3.csv.gz">aloi-hsb-14x3x3.csv.gz</a> <a href="6355684/files/aloi-hsb-8x4x4.csv.gz">aloi-hsb-8x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-9x5x5.csv.gz">aloi-hsb-9x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-13x4x4.csv.gz">aloi-hsb-13x4x4.csv.gz</a> <a href="6355684/files/aloi-hsb-14x5x5.csv.gz">aloi-hsb-14x5x5.csv.gz</a> <a href="6355684/files/aloi-hsb-10x6x6.csv.gz">aloi-hsb-10x6x6.csv.gz</a> <a href="6355684/files/aloi-hsb-14x6x6.csv.gz">aloi-hsb-14x6x6.csv.gz</a></td> </tr> <tr> <td>Color similiarity</td> <td>Average similarity to 77 reference colors (not histograms) 18 colors x 2 sat x 2 bri + 5 grey values (incl. white, black)</td> <td><a href="6355684/files/aloi-colorsim77.arff.gz">aloi-colorsim77.arff.gz</a> (feature subsets are meaningful here, as these features are computed independently of each other)</td> </tr> <tr> <td>Haralick features</td> <td>First 13 Haralick features (radius 1 pixel)</td> <td><a href="6355684/files/aloi-haralick-1.csv.gz">aloi-haralick-1.csv.gz</a></td> </tr> <tr> <td>Front to back</td> <td>Vectors representing front face vs. back faces of individual objects</td> <td><a href="6355684/files/front.arff.gz">front.arff.gz</a></td> </tr> <tr> <td>Basic light</td> <td>Vectors indicating basic light situations</td> <td><a href="6355684/files/light.arff.gz">light.arff.gz</a></td> </tr> <tr> <td>Manual annotations</td> <td>Manually annotated object groups of semantically related objects such as cups</td> <td><a href="6355684/files/manual1.arff.gz">manual1.arff.gz</a></td> </tr> </tbody></table> <p><strong>Outlier Detection Versions</strong></p> <p>Additionally, we generated a number of subsets for outlier detection:</p> <table> <tbody><tr> <th>Feature type</th> <th>Description</th> <th>Files</th> </tr> <tr> <td>RGB Histograms</td> <td>Downsampled to 100000 objects (553 outliers)</td> <td><a href="6355684/files/aloi-27d-100000-max10-tot553.csv.gz">aloi-27d-100000-max10-tot553.csv.gz</a> <a href="6355684/files/aloi-64d-100000-max10-tot553.csv.gz">aloi-64d-100000-max10-tot553.csv.gz</a></td> </tr> <tr> <td>&nbsp;</td> <td>Downsampled to 75000 objects (717 outliers)</td> <td><a href="6355684/files/aloi-27d-75000-max4-tot717.csv.gz">aloi-27d-75000-max4-tot717.csv.gz</a> <a href="6355684/files/aloi-64d-75000-max4-tot717.csv.gz">aloi-64d-75000-max4-tot717.csv.gz</a></td> </tr> <tr> <td>&nbsp;</td> <td>Downsampled to 50000 objects (1508 outliers)</td> <td><a href="6355684/files/aloi-27d-50000-max5-tot1508.csv.gz">aloi-27d-50000-max5-tot1508.csv.gz</a> <a href="6355684/files/aloi-64d-50000-max5-tot1508.csv.gz">aloi-64d-50000-max5-tot1508.csv.gz</a></td> </tr> </tbody></table>

opencc-by-4.0Jun 2010View details →
zenodo52/100

13/2 Sferamundi di Grecia. Seconda parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/2 Sferamundi di Grecia. Seconda parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/6 Sferamundi di Grecia. Sesta parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/6 Sferamundi di Grecia. Sesta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/5 Sferamundi di Grecia. Quinta parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/5 Sferamundi di Grecia. Quinta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/4 Sferamundi di Grecia. Quarta parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/4 Sferamundi di Grecia. Quarta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/3 Sferamundi di Grecia. Terza parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry&nbsp;<em>13/3 Sferamundi di Grecia. Terza parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

German image spectral library of urban surface materials

<p>The German image spectral library consists of 5102 labelled image spectra of urban surface materials covering the spectral wavelength range between 455 nm and 2449 nm. The spectra have been extracted from high resolution imaging spectroscopy data (HyMap) acquired over the German cities of Dresden (18/05/1999, 01/08/2000, 20/07/2003), Potsdam (18/05/1999) and Munich (17/06/2007, 25/06/2007). This image data package ensures the collection of the most typical urban surface materials including their variations due to different illumination, alteration, observation conditions, regional specifications and data processing characteristics.</p> <p>The collection was done in two main steps: (1) manual collection of spectrally pure urban surface material pixels from the Dresden and Potsdam data sets including additional information, such as the results of field investigations, a field spectral library and color infrared aerial imagery (Heiden et al., 2007 ) and subsequent reduction for redundant pixel spectra; (2) spectral dissimilarity analysis to include and label meaningful unknow spectra from the Munich data set (Jilge et al. 2017 ).&nbsp;</p> <p>The image spectra are labelled based on three sets of spectra labels: one for EAGLE land cover (EAGLE_LCC, consult the &ldquo;Explanatory Documentation of the EAGLE Concept&rdquo; from the Copernicus Land website) , one for generalized material groupings (GENLIB_LCH_BuC_MG) and one for more detailed artificial material type (GENLIB_LCH_BuC_AMT).</p> <p>While every effort was made to ensure accurate information, this data set is presented "as is" without warranties of any kind. The authors accept no liability or responsibility to any person as a consequence of any reliance upon the data presented here. The user assumes all responsibility and risk for the use of this data.</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Liquid Chromatography - Tandem Mass Spectrometry (LC-MS/MS) and Gas Chromatography - Mass Spectrometry (GC-MS) Reference Libraries from Global Natural Products Social Molecular Networking (GNPS) and National Institute of Standards and Technology (NIST) WebBook Processed for Spectral Library Matching

<div>In order to obtain a high-quality LC-MS/MS reference database for spectral library matching, we selected 22 high-quality GNPS tandem mass spectrometry databases generated under the positive ion mode. Further preprocessing similar to Huber et al involving mass-to-charge (m/z) and intensity filtering yields the database found in the file LCMS_GNPS_reference_library.csv which contains 14,705 electrospray ionization (ESI) mass spectra, each of which corresponds to a unique compound. The NIST WebBook database was used to construct GC-MS database contained in the file GCMS_NIST_WebBook.csv. This database contains 23,721 electron ionization (EI) mass spectra, each of which corresponds to a unique non-hyphenated Chemical Abstract Service (CAS) Registry Number.</div> <div>&nbsp;</div> <div>Both LC-MS/MS and GC-MS databases are organized into three columns: one for the identifier, one for the m/z values, and one for the intensity values. For example, if spectrum A has 20 ion fragments, then there will be 20 rows corresponding to spectrum A in the corresponding database with the identifier A repeated 20 times with the corresponding m/z and intensity values.</div>

opencc-by-4.0Jul 2024View details →
zenodo52/100

BiVib - Audio-Tactile Piano Sample Library

<p><strong>BiVib</strong> is an extensive piano sample library consisting of <strong>bi</strong>naural sounds and keyboard <strong>vib</strong>ration signals.<br>Samples were acquired with high-quality audio and vibration measurement equipment on two <a href="https://en.wikipedia.org/wiki/Disklavier">Yamaha Disklavier pianos</a> (one grand and one upright model) by means of computer-controlled playback of each key at ten different MIDI velocity values.<br>Project files (<em>instruments</em> and <em>multis</em>) are provided for use with the software sampler <a href="https://www.native-instruments.com/en/products/komplete/samplers/kontakt-6/">Native Instruments Kontakt</a> (version 5 and above, available for Windows and Mac OS).<br>The nominal specifications of the equipment used in the acquisition chain are reported in a companion document, allowing researchers to calculate physical quantities (e.g. acoustic pressure, vibration acceleration) from the recordings.<br>The library is especially suited for acoustic and vibration research on the piano, as well as for research on multimodal interaction with musical instruments.</p>

opencc-by-nc-sa-4.0Jan 2019View details →
zenodo52/100

Khotanese Manuscripts from Chinese Turkestan in the British Library (XML records)

<p>The file contains XML records matching the print edition of Skjaervo&#39;s catalogue,&nbsp;in TEI schema P4.</p> <p>The records in this file are a <strong>draft version</strong>. They&nbsp;have not yet been proofed and checked against physical holdings, which will be done with the next version release.</p> <p>The XML records&nbsp;have been produced as part of the work for the&nbsp;project <em>Beyond Boundaries: Religion, Region, Language and the State</em> (An ERC Synergy project from the European Research Council under the EU&#39;s 7th Framework Programme (FP7/2007-2013)/ERC grant agreement no.609823)</p>

opencc-by-4.0Aug 2019View details →
zenodo52/100

NMR screen reveals the diverse structural landscape of a G- quadruplex library

<p>This is the NMR dataset for the manuscript '<span>NMR screen reveals the diverse structural landscape of a G-</span><br><span>quadruplex library</span>'</p> <p>Abstract</p> <p><span>G-quadruplexes are noncanonical nucleic acid structures</span><br><span>formed by stacked guanosine tetrads. Despite their functional and</span><br><span>structural diversity, a single consensus model is typically used to</span><br><span>describe</span><span> </span><span>sequences</span><span> </span><span>with</span><span> </span><span>the</span><span> </span><span>potential</span><span> </span><span>to</span><span> </span><span>form</span><span> </span><span>G-quadruplex</span><br><span>structures. We are interested in developing more specific sequence</span><br><span>models</span><span> </span><span>for</span><span> </span><span>G-quadruplexes.</span><span> </span><span>In</span><span> </span><span>previous</span><span> </span><span>work,</span><span> </span><span>we</span><span> </span><span>functionally</span><br><span>characterized each sequence in a 496-member library of variants of a</span><br><span>monomeric</span><span> </span><span>reference</span><span> </span><span>G-quadruplex</span><span> </span><span>for</span><span> </span><span>the</span><span> </span><span>ability</span><span> </span><span>to</span><span> </span><span>bind</span><span> </span><span>GTP,</span><br><span>promote a model peroxidase reaction, generate intrinsic fluorescence,</span><br><span>and to form multimers. Here we used NMR to obtain a broad overview</span><br><span>of the structural features of this library. After determining the</span><span> </span><span>1</span><span>H NMR</span><br><span>spectrum of each of these 496 sequences, spectra were sorted into</span><br><span>multiple classes, most</span><span> </span><span>of</span><span> </span><span>which could be rationalized based on</span><br><span>mutational patterns in the primary sequence. A more detailed screen</span><br><span>using representative sequences provided additional information about</span><br><span>spectral classes, and confirmed that the classes determined based on</span><br><span>analysis of</span><span> </span><span>1</span><span>H NMR spectra are correlated with functional categories</span><br><span>identified in previous studies. These results provide new insights into</span><br><span>the surprising structural diversity of this library. They also show how</span><br><span>NMR can be used to identify classes of sequences with distinct</span><br><span>mutational signatures and functions.</span></p> <p><span>Link to journal article: <a href="https://doi.org/10.1002/chem.202401437"><span>https://doi.org/10.1002/chem.202401437</span></a></span></p>

opencc-by-4.0Sep 2024View details →
zenodo52/100

Datasets for evaluating scalable supervised learning for synthesize-on-demand chemical libraries

<p>This repository contains datasets for the manuscript &quot;Evaluating scalable supervised learning for synthesize-on-demand chemical libraries&quot;:</p> <ul> <li><strong>ams_all_preds.csv.gz</strong>: The AMS dataset predictions when using an RF or baseline model trained on the training dataset. Includes the predicted score and rank from each model for each compound. We started with 8,434,707 AMS compounds and detected that 247,025 were in the LC or MLPCN training data. These were removed from the AMS list, leaving 8,187,682 compounds to score. The compound matching was done on the SMILES that we canonicalized in rdkit.</li> <li><strong>ams_order_results.csv.gz</strong>: Information about the 1,024 compounds purchased from the AMS library. Excludes the 4 AMS compounds that were incompletely dissolved. Includes the chemical feature representation, information from the vendor, RF and baseline model predictions, screening results, and clustering results.</li> <li><strong>baseline_weight.npy</strong>: The saved Similarity Baseline model, which consists of the active compounds in the training data. This model was used to score the AMS library. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a>&nbsp;for code to load the model and make predictions on new compounds.</li> <li><strong>cdd_training_data.tar.gz</strong>: The LC1234 and MLPCN PriA-SSB screening data exported from CDD.</li> <li><strong>enamine_costs_clustered_v3_with_nneighbor.csv.gz</strong>: Contains 5,620 Enamine compounds that were selected based on the RF prediction score and availability. This file also contains the Taylor-Butina cluster ID when clustering the training compounds, 1,024 tested AMS compounds, and top-ranked Enamine compounds at a 0.4 threshold. The nearest neighbor compounds in the training and AMS sets are also included along with compound information from Enamine, RF model scores, and chemical feature representations.</li> <li><strong>enamine_dose_response_curve_plots.xlsx</strong>: Images of the dose response curves from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, multiple curves are shown in the same plot. The compound structure images and SMILES are exported from CDD, not generated with RDKit.</li> <li><strong>enamine_dose_response_curves.tsv</strong>: The dose response curve summaries from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, only the highest-quality dose response curve was used.</li> <li><strong>enamine_final_list.csv.gz</strong>: The final 100 filtered compounds from&nbsp;<code>enamine_top_10000.csv.gz</code>. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>enamine_PriA-SSB_dose_response_data.tar.gz</strong>: The dose response screening data from all three runs on the 68 Enamine compounds. The 2021-06-16 run was originally screened on 2020-08-24. 2021-06-16 is the date the compound identities were corrected. This run contains two 1,536 well plates.</li> <li><strong>enamine_top_10000.csv.gz</strong>: Top 10,000 predictions from the Enamine REAL dataset using the selected RF model. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>master_df.csv.gz</strong>: The output of preprocessing the files in&nbsp;<code>cdd_training_data.tar.gz</code>. Contains 441,900 rows.</li> <li><strong>random_forest_classification_139.pkl</strong>: The saved RF classification model with&nbsp;hyperparameter ID 139. This model was used to score the AMS and Enamine REAL libraries. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a> directory for code to load the model and make predictions on new compounds.</li> <li><strong>train_ams_real_cluster.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering at a 0.4 threshold applied to the training compounds, 1,024 tested AMS compounds, and top-ranked compounds from Enamine. Includes the chemical features, dataset to which the compound belongs, leader compound for each cluster, and whether the compound is a known hit.</li> <li><strong>training_df_single_fold.csv.gz</strong>: This is all ten folds in&nbsp;<code>training_folds.tar.gz</code>&nbsp;merged for convenience. Contains 427,300 compounds.</li> <li><strong>training_df_single_fold_with_ams_clustering.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering applied to the 427,300 training compounds and the 1,024 tested AMS compounds. Different clustering results are shown at the 0.2, 0.3, and 0.4 thresholds. Includes the leader compound for each cluster. Although the training and AMS compounds were clustered jointly, only the training compounds&#39; clusters are shown. The AMS compounds&#39; clusters are in&nbsp;<code>ams_order_results.csv.gz</code>.</li> <li><strong>training_folds.tar.gz</strong>: The LC1234 and MLPCN training data split into ten folds. This dataset with 427,300 compounds was used for cross validation and model selection. This dataset is derived from&nbsp;<code>master_df.csv.gz.</code></li> </ul> <p>If you use&nbsp;these&nbsp;datasets in a publication, please cite:</p> <p>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</p> <p>See&nbsp;PubChem AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>, AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1918986">1918986</a>,&nbsp;and the associated publications for details about the PriA-SSB screening data. The screening datasets were compiled from three separate sources that should all be cited if the training dataset is used in a publication:</p> <ul> <li>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</li> <li>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.8b00363">Practical model selection for prospective virtual screening</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2018.</li> <li>Andrew F. Voter<sup>+</sup>, Michael P. Killoran<sup>+</sup>, Gene E. Ananiev, Scott A. Wildman, F. Michael Hoffmann, James L. Keck.&nbsp;<a href="https://doi.org/10.1177/2472555217712001">A high-throughput screening strategy to identify inhibitors of SSB protein&ndash;protein interactions in an academic screening facility</a>.&nbsp;<em>SLAS Discovery</em>&nbsp;2018.</li> </ul> <ul> </ul>

opencc-by-4.0Oct 2021View details →
zenodo48/100

Thermal infrared emissivity spectral library of silicates measured under the Mercury simulated environment

<p>This is the thermal emissivity spectral library of silicates measured as a function of temperature under Mercury simulated environment. Data is measured at the Planetary Spectroscopy Laboratory (PSL), Institute of Planetary Research, German Aerospace Center (DLR), Berlin. The spectral library will be used for mineral identification of Mercury surface using MERTIS datasets. The manuscript related to this work is submitted to Icarus on the title &quot;<strong>Thermal Infrared Spectroscopy (7-14 &micro;m) of Silicates under Simulated Mercury Daytime Surface Conditions and their Detection: Supporting MERTIS onboard the BepiColombo Mission&quot;.</strong></p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes

<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here:&nbsp;<a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). &nbsp;&nbsp;</p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2&nbsp;Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name:&nbsp;</strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe:&nbsp;</strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' &nbsp;values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' &nbsp;values over 50%: FLAG_RP63.</li> <li><strong>flag_sum:&nbsp;</strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted:&nbsp;</strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value).&nbsp;</p> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis.&nbsp;</p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p>&nbsp;</p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0&nbsp; functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes.&nbsp;<br><br></p>

opencc-by-4.0Jun 2023View details →
zenodo48/100

Vis-NIR Soil Spectral Library of the Hungarian Soil Degradation Observation System

<p>Since soil spectroscopy is considered to be a fast, simple, accurate and non-destructive analytical method, its application can be integrated with wet analysis as an alternative. Therefore, development of national-level soil spectral libraries containing information about all soil types represented in a country is continuously increasing to serve as a basis for calibrated predictive models capable of assessing physical and chemical parameters of soils at multiple spatial scales. In this article, we present a database containing laboratory and visible-near infrared spectral data of legacy soil samples from the Hungarian Soil Degradation Observation System (HSDOS). The published data set includes the following parameters measured in 5,490 soil samples: pH<sub>KCl</sub>, soil organic matter (SOM), calcium carbonate (CaCO<sub>3</sub>), total salt content (TSC), total nitrogen (N<sub>total</sub>), soluble phosphorus (P<sub>2</sub>O<sub>5</sub>-AL), soluble potassium (K<sub>2</sub>O-AL), plasticity index according to Hungarian standard (PLI), soil profile depth and reflectance data between 350 and 2,500 nm wavelength. The presented database can be a complement for further soil related research on continental, national or regional scales to support sustainable soil management.</p> <p>Uploaded CSV file contains variables for general information, soil parameters and reflectance data of spectral bands between 350-2,500 nm. Details about variables can be found in Table 1.</p> <table> <tbody> <tr> <td> <p>Column name</p> </td> <td> <p>Description</p> </td> <td> <p>Method</p> </td> <td> <p>Unit</p> </td> </tr> <tr> <td> <p>SAMPLE_TDR_ID</p> </td> <td> <p>Original TDR IDs</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>SAMPLING_DATE</p> </td> <td> <p>Date of sampling</p> </td> <td> <p>-</p> </td> <td> <p>YYYY-MM-DD</p> </td> </tr> <tr> <td> <p>NORTHING_EOV</p> </td> <td> <p>Northing coordinate of sampling area centroids</p> </td> <td> <p>-</p> </td> <td> <p>m</p> </td> </tr> <tr> <td> <p>EASTING_EOV</p> </td> <td> <p>Easting coordinate of sampling area centroids</p> </td> <td> <p>-</p> </td> <td> <p>m</p> </td> </tr> <tr> <td> <p>LON_WGS84</p> </td> <td> <p>Longitude of sampling area centroids</p> </td> <td> <p>-</p> </td> <td> <p>&deg;</p> </td> </tr> <tr> <td> <p>LAT_WGS84</p> </td> <td> <p>Latitude of sampling area centroids</p> </td> <td> <p>-</p> </td> <td> <p>&deg;</p> </td> </tr> <tr> <td> <p>pH_KCl</p> </td> <td> <p>pH</p> </td> <td> <p>Potentiometer (MSZ&ndash;08 0206-2: 1978)<sup>40</sup></p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>SOM</p> </td> <td> <p>Soil organic matter</p> </td> <td> <p>E4/E6 ratio<sup>41</sup><sup>,</sup><sup>42</sup> (MSZ&ndash;08-0452:1980)<sup>43</sup></p> </td> <td> <p>%</p> </td> </tr> <tr> <td> <p>CaCO3</p> </td> <td> <p>Calcium carbonate</p> </td> <td> <p>Scheibler type calcimeter (MSZ&ndash;08 0206-2:1978)<sup>40</sup></p> </td> <td> <p>%</p> </td> </tr> <tr> <td> <p>TSC</p> </td> <td> <p>Total salt content</p> </td> <td> <p>EC-TDS electrode (MSZ&ndash;08-0206-2:1978)<sup>40</sup></p> </td> <td> <p>w/w %</p> </td> </tr> <tr> <td> <p>TN</p> </td> <td> <p>Total nitrogen</p> </td> <td> <p>Kjeldahl method<sup>44</sup> (ISO 11261:1995)<sup>45</sup></p> </td> <td> <p>mg kg<sup>-1</sup></p> </td> </tr> <tr> <td> <p>P2O5_AL</p> </td> <td> <p>Soluble phosphorus</p> </td> <td> <p>AL extract atomic adsorption spectrophotometry (MSZ 20135:1999)<sup>46</sup></p> </td> <td> <p>mg kg<sup>-1</sup></p> </td> </tr> <tr> <td> <p>K2O_AL</p> </td> <td> <p>Soluble potassium</p> </td> <td> <p>AL extract, flame photometer (MSZ 20135:1999)<sup>46</sup></p> </td> <td> <p>mg kg<sup>-1</sup></p> </td> </tr> <tr> <td> <p>PLI</p> </td> <td> <p>Plasticity index according to Hungarian standard</p> </td> <td> <p>Yarn test of Arany (MSZ&ndash;08 0205-2:1978)<sup>47</sup></p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>PROFILE_LEVEL</p> </td> <td> <p>Soil profile depth level</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>SPC350:2500</p> </td> <td> <p>spectral reflectance in the range of 350 and 2500 nm</p> </td> <td> <p>ASD FieldSpec 4 spectroradiometer</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> <p>Table 1. Summary of included attributes and data set structure with laboratory test methods applied on the soil samples.</p> <p>&nbsp;</p> <p><strong>For more details / to cite this dataset please use:</strong></p> <p>M&eacute;sz&aacute;ros, J., Kov&aacute;cs, Zs., L&aacute;szl&oacute;, P., Vass-Meyndt, Sz., Ko&oacute;s, S., Pirk&oacute;, B., Szűcs-V&aacute;s&aacute;rhelyi, N., Bakacsi, Zs., Laborczi, A., Balog, K., &amp; P&aacute;sztor, L. (2024). Vis-NIR soil spectral library of the Hungarian Soil Degradation Observation System. <em>Sci Data</em>&nbsp;<strong>12</strong>, 363 (2025). https://doi.org/10.1038/s41597-025-04667-9</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

UB2030 | The Future of Research Libraries as Knowledge Hubs | Interviews

<p>UB2030 is a podcast about the innovation of the university library through technological changes and shifts in research and education. In this podcast, David Oldenhof and Maurice Vanderfeesten discuss different subjects. They will be accompanied by guests who bring in an outside perspective.</p> <p>This release contains 11 episodes.</p> <p>Follow for more at <a href="https://ubvu.github.io/ub2030/">https://ubvu.github.io/ub2030/</a></p> <p>📖 <a href="https://doi.org/10.5281/zenodo.14615659"><strong>CLICK TO READ THE REPORT</strong></a></p> <p>🎧<strong>Listen on</strong> <a href="https://soundcloud.com/vu-library-live/sets/ub2030-the-future-of-research-libraries"><strong>SoundCloud</strong></a>, <a href="https://open.spotify.com/show/7dgTKn69lE3cnvs7CKv59v"><strong>Spotify</strong></a>, or your favorite <a href="https://antennapod.org/">(open)</a> podcast app.</p> <ul> <li>Authors: Maurice Vanderfeesten, David Oldenhof</li> <li>Client: Joeri Both</li> <li>Organization: <a href="https://vu.nl/nl/over-de-vu/diensten/universiteitsbibliotheek">University Library, Vrije Universiteit Amsterdam</a></li> <li>Date: 2023-05-19</li> </ul> <p><a href="https://doi.org/10.5281/zenodo.14615659" target="_blank" rel="noopener">DOI:10.5281/zenodo.14615659</a>&nbsp;(Rapport)</p> <p><a href="https://doi.org/10.5281/zenodo.10666049" target="_blank" rel="noopener">DOI:10.5281/zenodo.10666049</a> (Data)</p> <p><a href="https://ubvu.github.io/ub2030/">Project Page</a> | <a href="https://soundcloud.com/vu-library-live/sets/ub2030-the-future-of-research-libraries">Listen on SoundCloud</a> | <a href="https://feeds.soundcloud.com/users/soundcloud:users:527805591/sounds.rss">Podcast RSS</a> | <a href="https://forms.office.com/e/KX08BEenpu">Listener Feedback</a></p> <h2>Reason (Why Now)</h2> <p>The world is rapidly changing technologically, around AI, blockchain (NFTs), and linked data. As a library, you want to remain relevant for state-of-the-art research and education. We need to not only implement existing projects but also explore the horizon of opportunities and threats that await us. The client for this project is Joeri Both.</p> <h2>Project Goal (Why and How)</h2> <ul> <li>This project provides input for the next multi-year plan, creating a roadmap with a horizon up to 2030.</li> <li>With this project, we aim to increase the knowledge level of the UB by identifying key innovative/disruptive developments, to become a full-fledged partner for providers of (new) technological applications.</li> <li>From there, we translate innovative developments/trends into UB practice in broad terms, offering suggestions for workable/realistic pilot projects that contribute to the UB ambitions for researchers.</li> <li>The approach is to deliver an innovation sub-report each month, including a podcast episode, with the aim of making innovation an actively discussed topic within the UB.</li> <li>This gives the management team insight into the wide range of possibilities and developments, allowing them to make strategic choices for starting pilots and better embedding the innovation process in the organization.</li> </ul> <h2>Project Scope (What is and isn't included)</h2> <p>Through interviews, we gather information and ideas from each department and external experts. This information is linked to the ambition themes in the multi-year plan, focusing mainly on technological developments and their impact on our work processes and product/service offerings. The collected information is available in the form of an audio recording/podcast and interview report. For the Research Support department, these ideas are further developed into pilot proposals on three implementation levels: short, medium, and long term. The interviews are scheduled per department, divided into the themes of the ambitions in the multi-year plan.</p> <h2>Episodes</h2> <ul> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-01-introductie/"><strong>Episode 01 -- Introduction</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-02-rdm/"><strong>Episode 02 -- Future of Research Data Management and Research Software Management</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-03-research_intelligence/"><strong>Episode 03 -- Future of Research Intelligence</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-04-open_science/"><strong>Episode 04 -- Future of Open Science</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-05-research_support/"><strong>Episode 05 -- Future of Research Support</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-06-digital_services_and_infrastructures/"><strong>Episode 06 -- Future of Digital Services and Infrastructures</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-07-education_support/"><strong>Episode 07 -- Future of Education Support</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-08-special_collections/"><strong>Episode 08 -- Future of Special Collections</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-09-information_services/"><strong>Episode 09 -- Future of Information Services (not public)</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-10-library_desk_services/"><strong>Episode 10 -- Future of Library Desk Services</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-11-aquisition_and_metadata/"><strong>Episode 11 -- Future of Acquisition and Metadata (canceled)</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-12-society/"><strong>Episode 12 -- Future of the Library in Society</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-bonus-01-desci/"><strong>Episode 13 -- BONUS DeSci: Future of Open Science Ecosystems</strong></a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/ubvu/ub2030/compare/v1.1...v1.6">https://github.com/ubvu/ub2030/compare/v1.1...v1.6</a></p>

opencc-zeroJun 2024View details →
zenodo48/100

MALDI MS data and metadata from "A biocodicological analysis of the medieval library and archive from Orval Abbey, Belgium"

<p>See <a href="https://doi.org/10.1098/rsos.210210">Ruffini-Ronzani et al</a>.</p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

A list of items in the FAIRsFAIR training library

<p>A file containing a list of items,&nbsp;with basic metadata, that were included in the&nbsp;&nbsp;<a href="https://www.fairsfair.eu/competence-centre/training-library">FAIRsFAIR training library</a>&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record