Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,307
datasets available to search
ShareScore release 0.7.1
Dataset results
1,307 results for “libraries”
Library of simulated root images, with different noise levels
<p>This depository contains the images and associated RSML files used in the paper entitled "Using a structural root system model to evaluate and improve the accuracy of root image analysis pipelines" from the same authors.</p> <p> </p> <p>It contains: </p> <p>- 30.000 simulated root images of 10.000 different root systems, with different noise levels (0=null, 1=medium, 3=high);</p> <p>- 10.000 corresponding RSML files;</p> <p>- .csv files containing the ground-truth data for each modelled root system (500-data.csv);</p> <p>- .csv files containing the image descriptors extracted using RIA-J (500-descriptors.csv);</p> <p> </p> <p>The codes used to generate and analyse this dataset is available here: http://doi.org/10.5281/zenodo.208499</p> <p>Note about the metrics computed from the RSML files (contained in 500-data.csv): the root systems were simulated in a constrained 2D space (rhizotron-like). Therefore, the RSML_reader plugin computed the lengths of the different roots in 2D only (despite the fact that a Z coordinate is present in the the RSML files). </p> <p> </p> <p><strong>CORRECTION</strong>: There is a scale issue in the 500-descriptors.csv file. "area" and "convexhull" columns values should be divided by 10 </p> <p> </p> <p> </p>
PanDDA analysis of BRD1 screened against 3D-Fragment-Consortium Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of BRD1 screened against 3D-Fragment-Consortium Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48769 .</p> <p> </p>
PanDDA analysis of JMJD2D screened against Zenobia Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of JMJD2D screened against Zenobia Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48770 .</p>
PanDDA analysis of SP100 screened against selection of Maybridge Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of SP100 screened against selection of Maybridge Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48771 .</p>
PanDDA analysis of BAZ2B screened against Zenobia Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of BAZ2B screened against Zenobia Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48768 .</p>
Matrix multiplication software and results bundle for paper "Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library" for P^3MA submission
<p>This is the archive containing the matrix multiplication software and the results of the publication "<em>Tuning and optimization for a variety of many-core architectures without changing a single line of implementation code using the Alpaka library</em>" submitted to the P^3MA workshop 2017.</p> <p><strong>The archive has the following content:</strong></p> <ul> <li>Source code for the (tiled) matrix multiplication in "src": <ul> <li>regular version in "src/matmul": <ul> <li>Remote: https://github.com/theZiz/matmul.git (copy will be removed)</li> <li>Branch: topic-compatible-alpaka-0-1-0</li> <li>Commit: a63ba4810d6bfcca62c68dd57408af15028e78a3</li> </ul> </li> <li>forked version for XL in "src/matmul": <ul> <li>Remote: https://github.com/theZiz/matmul.git (copy will be removed)</li> <li>Branch: topic-xl-workaround</li> <li>Commit: 1fee028eccb8cf7b677e8071233e08aa9f81846a</li> </ul> </li> </ul> </li> <li>The compiled binaries and the results of the tuning and scaling runs are in "runs" in sub folders for each type of run and architectures.</li> </ul>
British Library fragment Or.8210/S.9498 collated with དབའབཞེད་ (version 1.1)
<p>Comparison of BL S.9498+S.13683 with the DBA' BZHED MS (=DBA' 2000 in our referencing system). The purpose of the comparison is three fold: a) to tentatively reconstruct the disposition of lines and content across the folio from which the fragment came, b) to determine if there is enough space in the reconstructed folio to accommodate the names of the three ministers sent to investigate Śāntarakṣita, and c) to determine the likely position of ན་ visible as a tail in the top missing line.</p>
Heat wave and cold snap event library under various technical choices for NERC subregions in the conterminous U.S. (1980 - 2024)
<p>This <strong>Extreme Thermal Event Library</strong> provides comprehensive records of heat wave and cold snap events from 1980 to 2024, aggregated at the North American Electric Reliability Corporation (NERC) subregion level for the conterminous United States. A map of NERC subregions with county-level mean temperatures is included in the file <strong>'NERC_subregions_with_mean_temperature.tif'</strong>.</p> <p>The heat wave and cold snap events were identified using temperature data simulated by the Thermodynamic Global Warming (TGW) model. The raw TGW hourly temperature data (at a 12-km resolution) was aggregated to the county level using spatial averaging within county boundaries. Daily mean, maximum, and minimum temperatures were then derived from the hourly surface air temperature data for each county. Subsequently, the county-level daily temperatures were spatially aggregated to the NERC subregion level using three distinct spatial aggregation methods:</p> <ol> <li><strong>Simple Mean (SM)</strong>: A simple average of county-level temperatures.</li> <li><strong>Area-Weighted Mean (MWA)</strong>: Weighted by the area of each county.</li> <li><strong>Population-Weighted Mean (MWP)</strong>: Weighted by the population of each county.</li> </ol> <p>The database includes separate zipped files for each spatial aggregation method:</p> <ul> <li><strong>"heat_wave_library_NERC_average.zip"</strong> and <strong>"cold_snap_library_NERC_average.zip"</strong> contain events based on the SM method.</li> <li><strong>"heat_wave_library_NERC_average_area.zip"</strong> and <strong>"cold_snap_library_NERC_average_area.zip"</strong> contain events based on the MWA method.</li> <li><strong>"heat_wave_library_NERC_average_pop.zip"</strong> and <strong>"cold_snap_library_NERC_average_pop.zip"</strong> contain events based on the MWP method.</li> </ul> <p>Each zipped file includes 12 event libraries, corresponding to 12 different event definitions. Details about these definitions are provided in the file <strong>'Event definitions.docx'</strong>.</p> <p> </p> <h3><strong>Event Library Structure</strong></h3> <p>Each row in an event library represents one detected thermal event. The columns in the library are defined as follows:</p> <ul> <li><strong>start_date</strong>: Start date of the event.</li> <li><strong>end_date</strong>: End date of the event.</li> <li><strong>centroid_date</strong>: The centroid date, calculated as the midpoint between the start and end dates.</li> <li><strong>highest_temperature / lowest_temperature</strong>: The highest daily maximum temperature (for heat waves) or lowest daily minimum temperature (for cold snaps), in Kelvin.</li> <li><strong>duration</strong>: Duration of the event in days.</li> <li><strong>NERC_ID</strong>: Identifier for the NERC subregion.</li> <li><strong>spatial_coverage</strong>: The spatial coverage of the event within the NERC subregion, expressed as the percentage of counties experiencing the event relative to the total number of counties in the subregion.</li> </ul>
Comparative Analysis of Anthraquinone and Chalcone Derivatives-Based Virtual Combinatorial Library. A Cheminformatics "Proof-of-Concept" Study
<p>This computational “proof-of-concept” study illustrated the combinatorial approach used to explain how the selected natural products' structures undergo molecular diversity analysis. A virtual combinatorial library (1.6M) based on 20 anthraquinones and 24 chalcones were enumerated. The resulting compounds were optimized to the near drug-likeness properties and the physicochemical descriptors were calculated for all datasets including FDA, Non-FDA, and natural products (NPs) datasets from ZINC 15. UMAP and principal component analysis (PCA) were applied to compare and represent the chemical space coverage of each dataset. Subsequently, the Laplacian score, and Gini coefficient, were applied to delineate feature selection, and selectivity among properties respectively. Finally, we demonstrated the diversity between the datasets by employing Murcko’s, and central scaffolds systems, calculated three fingerprint descriptors, and analyzed their diversity by PCA and self-organizing maps (SOM). The optimized enumeration resulted in 1,610,268 compounds with NP-Likeness, and synthetic feasibility mean scores close to FDA, Non-FDA, and NPs datasets. The overlap between the chemical space of 1.6M was more prominent with NPs. Laplacian score has prioritized NP-likeness and hydrogen bond acceptor properties (1.0 and 0.923) respectively, while the Gini coefficient showed that all properties have selective effects on datasets (0.81 to 0.93). Scaffold and fingerprint diversity indicated that the descending order for the tested datasets was FDA, Non-FDA, NPs, 1.6M. Virtual combinatorial libraries based on NPs can be considered as a source of the combinatorial compound with NP-likeness properties. Furthermore, measuring molecular diversity is supposed to be performed by different methods to allow for comparison and better judgment. </p> <p>This link provides an illustration of the whole virtual combinatorial library using the TMAP algorithm in addition to the complete dataset. TMAP is a recent algorithm applied to visualize ultra-large high-dimensional chemical libraries for structures and physicochemical properties (Probst & Reymond, 2020). This approach creates and distributes intuitive tree representations of big data sets with arbitrary dimensionality in the order of 10<sup>7</sup>.</p> <p><strong>To visualize the whole library of compounds, download the "index(2).rar", then extract the index.html that pop-up in the WinRAR application.</strong></p>
Cloud Forest Library
<p><strong>Cloud Forest Library</strong><br>This dataset is a point cloud collection of trees, shrubs, herbaceous plants, and other landscape elements. Each specimen is available as in .laz, .e57, .pcd, .ply, .xyz, and .3dm format. See the collection online at <a href="https://xyz.cct.lsu.edu/">xyz.cct.lsu.edu</a>.</p> <p><strong>License</strong><br>This dataset is released under the <a href="https://creativecommons.org/publicdomain/zero/1.0/">Creative Commons Zero 1.0 Universal Public Domain Dedication</a> by Brendan Harmon.</p>
A Versioned Literature Corpus derived from Biodiversity Heritage Library hash://md5/b3cd9de0685deeebf57a5d225e59c10f
<p>THIS DESCRIPTION IS A WORK IN PROGRESS, THE CONTENT OF THIS VERSIONED CORPUS IS FINAL.</p> <p>Biodiversity Heritage Library (BHL, https://biodiversitylibrary.org) contains hundreds of thousands of digital works related to biodiversity. This publications contains a versioned snapshot of work metadata and the digital signatures of their associated pdfs as seen in period 2025-03-26/2025-06-25 . </p>
Artifacts of the CGO 2024 Paper: EasyTracker: A Python Library for Controlling and Inspecting Program Execution
<p>This is the archive of the artifacts for the CGO 2024 Paper <em>"EasyTracker: A Python Library for Controlling and Inspecting Program Execution"</em></p> <p>The EasyTracker library is an open source project, refer to the Home page and Gitlab repository for up-to-date versions: </p> <ul> <li>Home Page: <a href="https://corse.gitlabpages.inria.fr/easytracker">https://corse.gitlabpages.inria.fr/easytracker</a></li> <li>Repository: <a href="https://gitlab.inria.fr/CORSE/easytracker">https://gitlab.inria.fr/CORSE/easytracker</a></li> </ul> <p>The details of the artifacts generation are described in the paper appendix or in the <code>README.md</code> file included in the main artifacts archive <code>easytracker-artifacts-cgo-2024-v1.2.0.tar.gz</code>.</p> <p>Summary of artifacts construction steps (execution in a Docker container):</p> <ul> <li>ensure Docker is installed with <code>docker --version</code>,</li> <li>download the artifacts archive <code>easytracker-artifacts-cgo-2024-v1.2.0.tar.gz</code>,</li> <li>extract with <code>tar xvzf easytracker-artifacts-cgo-2024-v1.2.0.tar.gz</code>,</li> <li>change dir with <code>cd eastracker-artifacts-cgo-2024</code>,</li> <li>download the EasyTracker sources archive <code>easytracker-archive-dfe8aa888f.tar.gz</code><a href="../api/records/10428215/draft/files/easytracker-archive-dfe8aa888f.tar.gz/content" target="_blank" rel="noopener noreferrer">,</a></li> <li>extract with <code>tar xvzf easytracker-archive-dfe8aa888f.tar.gz</code>,</li> <li>download the docker image <code>docker-image-easytracker-1.2.0.tar</code>,</li> <li>load the Docker image with <code>docker load -i docker-image-easytracker-1.2.0.tar</code>, </li> <li>generate all artifacts with <code>./in-docker.sh ./run-all.sh</code>,</li> <li>all artifacts are generated in <code>figure-*/</code> directories,</li> <li>refer to the artifacts archive <code>README.md</code> file for more details, or refer to the paper artifacts appendix.</li> </ul> <p>Note that this artifacts archive is extracted from the artifacts repository at tag <code>v1.2.0</code>: <a title="Opens in new tab" href="https://gitlab.inria.fr/CORSE/easytracker-artifacts-cgo-2024/-/tree/v1.2.0" target="_blank" rel="noopener">https://gitlab.inria.fr/CORSE/easytracker-artifacts-cgo-2024/-/tree/v1.2.0 </a></p>
Supplementary files for the dingo Python library
<p>We compare the <a href="https://drops.dagstuhl.de/opus/frontdoor.php?source_opus=13820">Multiphase Monte Carlo flux Sampling (MMCS)</a> feature of the <a href="https://github.com/GeomScale/dingo">dingo</a> library against the combined method of <a href="https://gitlab.com/csb.ethz/PolyRound">PolyRound</a> (for rounding) followed by <a href="https://modsim.github.io/hopsy/">hopsy</a> (for sampling) on a set of 7 models with a ranging dimension (<em>ext_data.zip</em>).</p> <p>The <em>simpl_transf_polytopes.zip</em> contains the polytopes retrieved after the <em>simplify()</em> and <em>transform() </em>functions of the PolyRound library. <br>These polytopes were used as input for the dingo implementation of the MMCS algorithm asking for an ESS of 1000. <br>Under the <em>dingo_samples_on_simpl_transf_polytopes.zip</em> the resulting samples from <em>dingo</em> can be found using the MMCS algorithm. </p> <p>Similarly, <em>polyrounded_polytopes.zip contains </em>the polytopes retrieved after applying <em>simplify()</em>, <em>transform()</em> and <em>round()</em> functions of the PolyRound library. <br>These polytopes were used as input for the <a href="https://modsim.github.io/hopsy/">hopsy</a> library, again, asking for an ESS of 1000. <br>The <em>hopsy_samples.zip</em> folder contains the resulting samples from hopsy<em> </em>library, using a thinning of 100<em>d </em>;<em> </em>only in the case of Recon3D a thinning of 200<em>d </em>was used as suggested by the authors. Under the <em>hopsy_samples_ess_1000.zip </em>folder, we provide the <em>hopsy</em> samples with an ESS of 1000. <br>Last, the samples produced using the efficient Billiard Walk implementation of <em>dingo</em> can be found under the <em>BWRsamples.zip </em>file. <br>In this case, 20000 points were sampled for each model.</p> <p>Further, the <em>sars_samples.zip </em>file contains <em>dingo</em> samples from the solution space of the SARS-CoV-2 integrated model of <a href="https://doi.org/10.1093/bioinformatics/btaa813">Renz et <em>al</em> (2020)</a> for the following cases: </p> <ul> <li>unbiased; where the zero vector has been used as the objective function of the model</li> <li>after maximising for the human biomass </li> <li>after maximising for the virus biomass objective function (VBOF)</li> </ul> <p>The following Python scripts to perform these experiments are included:</p> <ul> <li><em>polyround_preproces.py </em>: runs the <em>PolyRound </em>functions and builds the simplified and transformed polytopes that <em>dingo </em>will use as well as the simplified, transformed and rounded polytopes <em>hopsy</em> uses</li> <li><em>hopsy_on_polyrounded_polytopes.py </em>: performs sampling with <em>hopsy </em></li> <li><em>dingo_on_simpl_transf_polytopes.py </em>: performs sampling with <em>dingo </em></li> <li><em>run_bwr_exp.py</em> computes samples using the efficient Billiard Walk of <em>dingo</em></li> <li><em>binary_search.py</em> : a function to return the index in the chain where ESS becomes 1000</li> <li><em>compute_ess.py: </em>based on a model's <em>hopsy</em> samples (under the <em>hopsy_samples.zip </em>folder) it retrieves the samples with an ESS of 1000 and the corresponding required time for <em>hopsy </em>to build them. The script requires the total time of the <em>hopsy </em>experiment recorded in the model's corresponding <em>.txt </em>file (you can find this under the <em>hopsy_samples.zip)</em></li> <li><em>compute_ess_psrf_per_phase.py </em>computes ESS and PSRF in specific indices which correspond to those when MMCS switches from a phase a to next one</li> </ul> <p>A <a href="https://github.com/hariszaf/dingo/blob/vbof/tutorials/vbof.ipynb">notebook</a> is available for how the integrated model was sampled. </p> <p> </p>
GNPS Bile acid modifications MS/MS spectral library
<p>Bile acids are important signaling molecules with impact on host health and metabolism. However, the diversity of bile acids furnished by both the host and microbes is incompletely characterized. To address this knowledge gap, we created a reusable resource of candidate tandem mass spectrometry (MS/MS) spectra by filtering approximately 1.2 billion publicly available MS/MS spectra from 2,706 untargeted metabolomics projects for bile acid-specific MS/MS ion patterns. Provided here are two MGFs/mzMLs consisting of 594,431 MS/MS spectra directly obtained from the public data on GNPS/MassIVE using MassQL queries designed for non-, mono-, di-, tri-, tetra- and penta-hydroxylated bile acids and MS/MS spectra from synthetic standards of 38 amino acids and 28 polyamines conjugated to bile acids. An annotation table (.tsv file) with delta masses and GNPS library matches is also provided. The MGFs/mzMLs can be selected as a group while running molecular networking jobs on GNPS (<a href="https://gnps.ucsd.edu/ProteoSAFe/static/gnps-splash.jsp">https://gnps.ucsd.edu/ProteoSAFe/static/gnps-splash.jsp</a>) and used in conjugation with the annotation table to gain unprecedented insights into the new biology of bile acids.</p>
The IMITATOR benchmarks library 2.1: A benchmarks library for extended parametric timed automata
<p>We present here the IMITATOR benchmarks library 2.1: A benchmarks library for extended parametric timed automata</p> <p> </p> <p>We present two archives:</p> <p>- one (benchmarks.zip) with the models and the properties</p> <p>- one (full.zip) with the benchmarks and all the results: the expected results, generated PDF and graphics, and a whole standalone Web page (more or less equivalent to <a href="https://www.imitator.fr/static/library.html">www.imitator.fr/static/library.html</a>) summarizing all benchmarks</p> <p> </p> <p>See a full description in the TAP 2021 paper ("<a href="https://link.springer.com/10.1007/978-3-030-79379-1_3">A Benchmarks Library for Extended Parametric Timed Automata</a>")</p>
Automatic Classification of Final Assignments at the Nuclear Polytechnic Library
<p>This study aimed to look for a method to automatically classify the final projects of Indonesian Nuclear Technology Polytechnic students.</p>
An 8-(Diazomethyl) Quinoline Derivatized Acyl-CoA In Silico Mass Spectral Library Reveals the Landscape of Acyl-CoA in the Aging Mouse Organs (Data Supplement)
<p>Data supplement for publication "An 8-(Diazomethyl) Quinoline Derivatized Acyl-CoA In Silico Mass Spectral Library Reveals the Landscape of Acyl-CoA in the Aging Mouse Organs (Data Supplement)" (2024)</p> <p>Jinhui Yu<sup>1†</sup>, Menghao Guo<sup>1,3†</sup>, Sha Li<sup>5</sup>, Jian Ni<sup>2</sup>, Yu-Qi Feng<sup>4,5</sup>*, Jun Ding<sup>1,2</sup>*</p> <p>1. CAS Key Laboratory of Plant Germplasm Enhancement and Specialty Agriculture, Wuhan Botanical Garden, Chinese Academy of Sciences, Wuhan, 430074, PR China.</p> <p>2. Renmin Hospital of Wuhan University, Wuhan University, 430072, Wuhan, P. R. China.</p> <p>3. College of Life Sciences, Wuhan University, Wuhan 430072, China.</p> <p>4. School of Bioengineering and Health, Wuhan Textile University, Wuhan 430200, China.</p> <p>5. Frontier Science Center for Immunology and Metabolism, Wuhan University, Wuhan, 430071, China.</p> <p> </p> <p>†The authors contribute equally to this work.</p> <p>* Corresponding author Email: <a href="mailto:dingjun@wbgcas.cn">dingjun@wbgcas.cn</a>, <a href="mailto:yqfeng@whu.edu.cn">yqfeng@whu.edu.cn</a></p> <p>----<br>Content:<br>1) Developement: XLS template sheet for development (can be used to adjust or create new library)<br>2) MSP library: spectra in NIST MSP format (use with MS-Dial or NIST MS Search)<br>3) NIST library: NIST23 compatible 8-DMQ-acyl-CoA library (use with NIST MS Search)<br>4) Reference spectra msp: 8-DMQ-acyl-CoA authentic MS/MS spectrum in NIST MSP format (generated by Thermo QE HF-X MS (HCD), for searching NIST MS-Search)</p> <p>Version 1.0<br>April 22 2024</p>
Morphing libraries, QSAR models, and compounds predicted to be active on the Glucocorticoid receptor (GR)
<p>This repository contains datasets and files related to the computational drug discovery project of the chemical space exploration of the Glucocorticoid receptor. The accompanying Python code is freely available in the GitHub repository (<a title="https://github.com/Iagea/GRML_analyses" href="https://github.com/Iagea/GRML_analyses" target="_blank" rel="noreferrer noopener">https://github.com/Iagea/GRML_analyses</a>).</p> <p><strong>Morphing Libraries:</strong></p> <ul> <li><strong>GRML_library.csv:</strong> The GRML library is the collection of 999,015 virtual compounds generated by Molpher [1-2] starting from GR ligands with unique Bemis-Murcko scaffolds collected from the ChEMBL17 and IMG libraries.</li> <li><strong>RML_library.csv:</strong> The RML library is the collection of 1,346,310 virtual compounds generated by Molpher starting from compounds with unique Bemis-Murcko scaffolds randomly selected from the ZINC database.</li> </ul> <p><strong>IMG library:</strong></p> <ul> <li><strong>IMG_non_proprietary.csv</strong>: The non-proprietary IMG library subset containing 12,956 compounds and their corresponding B-scores from the primary screen.</li> </ul> <p><strong>Molpher inputs:</strong></p> <ul> <li><strong>GR_inputs.csv</strong>: The GR inputs are the ligands used to create the GRML library, 204 compounds from ChEMBL17 (95 compounds) and the non-proprietary dataset from IMG (109 compounds).</li> <li><strong>Random_inputs.csv</strong>: The random inputs are 249 random ZINC compounds used to create the Random library.</li> </ul> <p><strong>Model's training sets:</strong></p> <ul> <li><strong>Model33_training_set.csv</strong>: Random forest classification model training set, it includes 865 compounds; known GR actives and inactives from ChEMBL33 (738 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>Model17_training_set.csv</strong>: Random forest classification model training set, it includes 601 compounds; known GR actives and inactives from ChEMBL17 (474 compounds) and non-proprietary active ligands from the IMG library (127 compounds).</li> <li><strong>RFR_training_set.csv</strong>: Random forest regression model training set, it includes 89 compounds; known GR actives and inactives from ChEMBL33 that fit into the GR pharmacophore with the four features we describe in our paper.</li> </ul> <p><strong>Models:</strong></p> <ul> <li><strong>Model33.pkl:</strong> Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL33 and IMG libraries.</li> <li><strong>Model17.pkl</strong>: Python pickle file containing the trained Random forest classification models used along with Mondrian cross-conformal prediction to classify GR actives/inactives. This model was trained with ChEMBL17 and IMG libraries.</li> <li><strong>RFR_models.pkl</strong><em>: </em>Python pickle file containing the 100 trained random forest regression models used to rank the proposed active morphs. These models were trained with the RFR_training_set.csv.</li> </ul> <p><strong>Active predicted morphs:</strong></p> <ul> <li><strong>all_morphs_actives</strong><em><strong>_</strong></em><strong>predicted.xlsx:</strong> An Excel spreadsheet containing two sheets. 1) All 22,524 GRML active predicted morphs. 2) All 4,341 RML active predicted morphs. The QED, NIBR Severity Score, and Molskill Score are given for each morph.</li> </ul> <p><strong>Proposed GR active ligands:</strong></p> <ul> <li><strong>designed_ligands.xlsx</strong>: An Excel spreadsheet containing two sheets. 1) All 54 designed GR ligands with their QED, NIBR severity score, MolSkill score, consensus ranking from the 100 RFR models, and the result of the manual annotation and remarks, if available. 2) The structure of the 54 ligands based on their manual annotation and presence or not in ChEMBL33 database.</li> </ul> <p>Researchers and professionals in the field of drug discovery and cheminformatics may find these resources useful for further analysis and investigations.</p> <p><strong>Bibliography</strong></p> <p>[1] Hoksza, D., Škoda, P., Voršilák, M. <em>et al.</em> Molpher: a software framework for systematic chemical space exploration. <em>J Cheminform</em> <strong>6</strong>, 7 (2014). https://doi.org/10.1186/1758-2946-6-7</p> <p>[2] <a href="https://github.com/lich-uct/molpher-lib">https://github.com/lich-uct/molpher-lib</a></p>
quickSparseM: a library for memory- and time-efficient computation on large, sparse matrices with application to omics data
<p>This page contains the code and datasets used in "quickSparseM: a library for memory- and time-efficient computation on large, sparse matrices with application to omics data".</p> <p>File <strong>test_datasets.zip</strong> containes three datasets:</p> <ul> <li><em>D1.RData</em>: scRNA-seq omics data derived from Salcher et. al (2022)</li> <li><em>D2.RData</em>: scRNA-seq omics data derived from Pineda et al. (2024)</li> <li><em>D3.RData</em>: in silico WGS SNP data.</li> </ul> <p>File <strong>test_scripts.zip</strong> containes the code to reproduce the results.</p>
Reference Sequence Library Resources - Maine-eDNA
<p>The following files and resources are associated with the Maine-eDNA Reference Library Research Group - aiming to create reference sequence library resources for researchers part of Maine-eDNA or otherwise interested in leveraging eDNA tools for research in the Gulf of Maine.</p> <p>These include:</p> <ul> <li> <p>RoughWorkflow.zip</p> </li> <ul> <li> <p>contains a NCBI scraping script to build reference databases based on an input species list, a configuration file for the script, and genbankr version - most useful for shorter species lists (time-intensive)</p> </li> </ul> <li> <p>12S_REFDB.fasta</p> </li> <ul> <li> <p>A DADA2-compliant reference library for 12S sequences, built with the RoughWorkflow based on the GNRMaineSpecies_May2024 species list</p> </li> </ul> <li> <p>COI_REFDB.fasta</p> </li> <ul> <li> <p>A DADA2-compliant reference library for COI sequences, built with the RoughWorkflow based on the GNRMaineSpecies_May2024 species list</p> </li> </ul> <li> <p>GitHub Repo - referee - <a href="https://github.com/BigelowLab/referee">https://github.com/BigelowLab/referee</a></p> </li> <ul> <li> <p>Scripts for building reference databases for the Maine-eDNA project through downloading GenBank - this workflow is suggested especially for large species lists</p> </li> </ul> <li> <p>GitHub Repo - refdbtools - <a href="https://github.com/BigelowLab/refdbtools">https://github.com/BigelowLab/refdbtools</a> </p> </li> <ul> <li> <p>R language package to assist in making eDNA reference databases</p> </li> </ul> <li> <p>SpeciesListMeta_Shareable.xlsx</p> </li> <ul> <li> <p>Describes the sources from which the GNRMaineSpecies_May2024.csv and GNRMaineTaxonomiedSpecies_May2024.csv species lists were compiled - sources not associated with a link were found as separate files and are hosted elsewhere. Species list sources are courtesy of public datasets, Maine-eDNA researchers, and collaborators</p> </li> </ul> <li> <p>GNRMaineSpecies_May2024.csv</p> </li> <ul> <li> <p>A Maine (and surrounding area) species list ran through taxize’s gnr_resolve to resolve species names and fill out taxonomy (full results)</p> </li> </ul> <li> <p>GNRMaineTaxonomiedSpecies_May2024.csv</p> </li> <ul> <li> <p>A Maine (and surrounding area) species list ran through taxize’s gnr_resolve to resolve species names and fill out taxonomy (only species results that could be resolved with taxonomy)</p> </li> </ul> </ul> <p> </p> <p>Contact Beth Y. Davis - bethy.davis4@gmail.com for questions</p> <p> </p> <p>###</p> <p>Changelog:</p> <p>May 16, 2023 (Version 1.0) - Initial upload</p> <p>July 11, 2023 (Version 1.1) - Did additional cleaning to the MaineSpeciesList_Clean file and uploaded the new version - July2023_SpeciesList</p> <p>May 20, 2024 (Version v3) - Additional cleaning to correct deduplication errors and ran taxize's gnr_resolve to resolve names and fill out taxonomy. The version update adds two files, GNRMaineSpecies_May2024.csv containing the full result of gnr_resolve, and GNRMaineTaxonomiedSpecies_May2024.csv contains only those species from the original list that could be resolved with taxonomy. The SpeciesListMeta_Shareable.csv has not been updated but is still an accurate tracker of the sources from which the species names were gained from.</p> <p>November 27, 2024 (Version 4.0) - Updated Zenodo description and added the RoughWorkflow R files, 12S and COI files</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.