Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

13,064

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

13,064 results for “Prediction”

Learn how ShareScore rates datasets ↗
zenodo44/100

RNA-Protein Interaction Prediction Using Network-Guided Deep Learning

<p>RNA-protein interactions are critical to various life processes, including fundamental translation and gene regulation. Identifying these interactions is vital for understanding the mechanisms underlying life processes. Then, ZHMolGraph is an advanced pipeline that integrates graph neural network sampling strategy and unsupervised large language models to enhance binding predictions for novel RNAs and proteins.</p> <div>&nbsp;</div>

openmit-licenseJul 2024View details →
zenodo44/100

Machine Learning Predicts Meter-Scale Laboratory Earthquakes

<p>Numerical data and Python scripts used to make Figures in the manuscript entitled "Machine Learning Predicts Meter-Scale Laboratory Earthquakes". We provide the data and scripts to create Figures 1 to 6 from the main text and Figures S1 to S15 from the supplementary information. Note that you should download the experimental catalog of Yamashita et al. (Nature Comm, 2021) and address the request for shear force data (experimental number LB12-011) to Futoshi Yamashita, the organizer of the target experiment. This code is developed using Python 3.11.7.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Predicted times, spatial coordinates of bow shock crossings and shock geometry at Mars from the NASA/MAVEN mission, using spacecraft ephemerides and magnetic field data, with a predictor-corrector algorithm

<p><strong>CHARACTERISTICS</strong><br>Planet: <strong>Mars</strong><br>Radius: <strong>R<sub>M</sub> = 3389.5 km</strong> (volumetric mean planetary radius)<br>Spacecraft: <strong>NASA/Mars Atmosphere and Volatile Evolution (MAVEN)</strong><br>Spacecraft coordinates system: <strong>Mars Solar Orbital (MSO)</strong> equivalent to <em>Sun-State </em>coordinate system:</p> <ul> <li>+<em>X<sub>MSO</sub></em>&nbsp;points towards the Sun from the planet&rsquo;s centre,</li> <li>+<em>Z<sub>MSO</sub></em>&nbsp;towards Mars&rsquo; North pole and perpendicular to the orbital plane defined as the&nbsp;<em>X<sub>MSO</sub></em>&ndash;<em>Y<sub>MSO</sub></em>&nbsp;plane passing through the centre of Mars,</li> <li><em>Y<sub>MSO</sub></em>&nbsp;completes the orthogonal system.</li> </ul> <p>Time span:&nbsp;<strong>01/11/2014 to 30/04/2024</strong> (Mars Years MY32 to MY36 included, part of MY37).<br>Total number N of candidate bow shock crossings in the database: <strong>N = 20107</strong></p> <p><strong>ORIGINAL DATASETS USED</strong><br>The original MAVEN/MAG data repository on which these algorithms&nbsp;were applied is available on NASA's Planetary Data System (PDS) at&nbsp;<a href="https://doi.org/10.17189/1414178">https://doi.org/10.17189/1414178</a>.&nbsp;For this study, 1-Hz magnetic field data was used.</p> <p><strong>METHOD</strong><br>To construct this database from the original datasets above, the&nbsp;predictor and predictor-corrector algorithms used are described in:<br>Simon Wedlund, C., Volwerk, M., Beth, A., Mazelle, C.,&nbsp;M&ouml;stl, C., Halekas, J., Gruesbeck, J. and Rojas-Castillo, D.,&nbsp;(2022), A Fast Bow Shock Location Predictor-Estimator From 2D&nbsp;and 3D Analytical Models: Application to Mars and the MAVEN&nbsp;mission,&nbsp;<em>Journal of Geophysical Research</em>, <strong>127</strong>, 1-33,&nbsp;e2021JA029942,&nbsp;<a href="https://doi. org/10.1029/2021JA029942">https://doi. org/10.1029/2021JA029942</a>.&nbsp;</p> <p>Also available at: <a href="https://doi.org/10.1002/essoar.10507942.1">https://doi.org/10.1002/essoar.10507942.1 </a>&nbsp;and as arXiv e-print:&nbsp;<a href="https://doi.org/10.48550/arXiv.2109.04366">https://doi.org/10.48550/arXiv.2109.04366</a></p> <p>These algorithms consist of two consecutive steps:&nbsp;</p> <ol> <li>Predictor geometric algorithm based on J. Gruesbeck's 3D model&nbsp;(<a href="https://doi.org/10.1029/2018JA025366">Gruesbeck et al. 2018</a>) for prediction of Mars bow shock&nbsp;position</li> <li>Corrector algorithm based on magnetic field measurements (magnitude and fluctuations).</li> </ol> <p><strong>REMARK ON VERSIONS</strong><br>From Version 3 onwards, we also provide the angle between the average Interplanetary Magnetic Field (IMF) vector upstream of the shock and the shock normal, noted \(\theta_{Bn}\)(ThetaBn). Assuming a smooth shock surface and&nbsp;the 3D model of Gruesbeck et al. (2018, all points), this gives a&nbsp;first indication of the geometry of the shock, so that:</p> <ul> <li>45<sup>∘</sup>&lt;<em>&theta;</em><sub><em>B</em><em>n</em></sub>&lt;135<sup>∘</sup>: quasi-perpendicular shock condition</li> <li><em>&theta;</em><sub><em>B</em><em>n</em></sub>&le;45<sup>∘</sup> and <em>&theta;</em><sub><em>B</em><em>n</em></sub>&ge;135<sup>∘</sup>: quasi-parallel shock condition</li> </ul> <p>Uncertainty on these angles is estimated to be &plusmn; 5&ordm;.&nbsp;</p> <p>From Version 4 onwards, we also added the solar longitude Ls (in degrees).</p> <p>For details, see Simon Wedlund et al. (2022) above, &sect;2.3 pp. 10-12.&nbsp;Note that due to minor adjustments in the code, some of the&nbsp;ThetaBn angles calculated here for the examples of Fig. 6 in&nbsp;Simon Wedlund et al. (2022) may slightly differ from the values&nbsp;quoted in the paper.</p> <p><strong>VARIABLES DESCRIPTION</strong><br>This database contains the following ASCII variables:</p> <ul> <li>Bow shock times in MAVEN's database (1-s resolution): <em>T</em><sub>bs</sub></li> <li>Mars Solar Orbital coordinates of the shock, in&nbsp;units of Mars radius <em>R</em><sub><em>M</em>&nbsp;</sub>(<em>R<sub>M</sub></em> = 3389.5 km):<br><em>X<sub>MSO</sub></em>,<sub>&nbsp;</sub><em>Y<sub>MSO</sub></em>,&nbsp;<em>Z<sub>MSO</sub></em>&nbsp;and Euclidean&nbsp;distance&nbsp;\(R_{MSO} = \sqrt{X_{MSO}^2 + Y_{MSO}^2 + Z_{MSO}^2}\)&nbsp;(in&nbsp;<em>R<sub>M</sub></em>)</li> <li>Solar Zenith angle in degrees:&nbsp;<em>SZA</em> = \(\tan^{-1}{Y_{MSO}^2+Z_{MSO}^2 \over X_{MSO}^2}\)&nbsp;(in&nbsp;&ordm;)&nbsp;</li> <li>Angle between average B-field direction and&nbsp;shock&nbsp;normal assuming a smooth shock surface \(\theta_{Bn}\) (ThetaBn,&nbsp;in &ordm;) <ul> <li>45 &lt; ThetaBn &lt;&nbsp; 135 deg: quasi-&perp; shock</li> <li>ThetaBn &le;45 deg &amp; ThetaBn &ge; 135 deg: quasi-|| shock</li> </ul> </li> <li>Solar longitude Ls, in degrees.</li> <li>Flag for crossing: <ul> <li>sheath&nbsp;\(\longrightarrow\)&nbsp;solar wind, flag = 0.</li> <li>solar wind \(\longrightarrow\)&nbsp;sheath, flag = 1.</li> </ul> </li> </ul> <p><strong>WARNING</strong><br>This database is based on an automatic statistical&nbsp;geometrical estimate, further refined by constraints on magnetic&nbsp;field. It is aimed at giving a first approximation of the shock area times in the MAVEN data. It is particularly suited to&nbsp;statistical studies and region identification in the MAVEN&nbsp;datasets. As such, this database should be used as a <em>first&nbsp;indicator</em> of the shock location, and <em>with</em> <em>caution</em>: it <strong>CANNOT</strong>, and <strong>WILL NOT&nbsp;</strong>substitute, especially in case studies, for a careful analysis&nbsp;of the full magnetometer and plasma suite bow shock signatures.&nbsp;Moreover, the algorithm is optimised for detecting the first disturbance observed in&nbsp;the magnetic field immediately ahead of the shock's foot (in the foreshock area), and not for the detection of&nbsp;other structures in the shock, such as the shock ramp. The&nbsp;"shock"&nbsp;location is therefore given here with typical uncertainties of about 0.075 R<sub>M</sub>&nbsp;(with R<sub>M</sub>&nbsp;= 3389.5 km, i.e., about 250 km in the radial direction). Finally, for multiple shock crossings, the algorithm chooses the first occurrence of the shock starting from the undisturbed&nbsp;solar wind.</p> <p>Current formatting optimised for MATLAB.</p> <p><strong>ACKNOWLEDGEMENTS</strong><br>C. Simon Wedlund and M. Volwerk thank the Austrian Science Fund&nbsp;(FWF) project P32035-N36. C. M&ouml;stl thanks the Austrian Science&nbsp;Fund FWF projects P31659-N27, P31521-N27. A. Beth thanks the&nbsp;Swedish National Space Agency (SNSA) and its support with the&nbsp;grant 108/18.&nbsp;This database was notably used to add to the Helio4Cast database&nbsp;which monitors solar wind parameters in the solar system&nbsp;(<a href="https://doi.org/10.6084/m9.figshare.6356420">https://doi.org/10.6084/m9.figshare.6356420</a>). Helio4Cast is&nbsp;available at <a href="http://www.helioforecast.space/icmecat">www.helioforecast.space/icmeca</a>t and&nbsp;<a href="http://www.helioforecast.space/sircat">www.helioforecast.space/sircat</a>. &nbsp; &nbsp;&nbsp;</p> <p><strong>LICENSE AND RIGHTS</strong><br>This database is shared under a Creative Commons CC-BY-4.0 license.</p> <p>Version 1 (c) Cyril Simon Wedlund @ Space Research Institute of Graz (IWF),&nbsp;<br>&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Austrian Academy of Sciences (&Ouml;AW), 2021-09-08<br>Version 2 (c) CSW @ &Ouml;AW/IWF, 2021-11-30 -- Addition of R_MSO and SZA<br>Version 3 (c) CSW @ &Ouml;AW/IWF, 2022-02-09 -- Addition of ThetaBn<br>Version 4 (c) CSW @ &Ouml;AW/IWF, 2025-03-20 -- Addition of Ls, Bx, By, Bz and Bt.</p> <p>&nbsp;</p> <p><br>Contact email: &nbsp; &nbsp; &nbsp; &nbsp;cyril.simon.wedlund@gmail.com</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network

<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Datasets for Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes

<p>All used Datasets to pair with &quot;Supervised Learning Model Predicts Protein Adsorption to Carbon Nanotubes&quot; by Nicholas Ouassil*, Rebecca L. Pinals*, Jackson Travis Del Bonis-O&#39;Donnell, Jeffrey W. Wang, and Markita P. Landry</p> <p>*Co-authors</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction - Datasets

<p>Datasets to NeurIPS 2021 accepted paper &quot;Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction&quot;.</p> <p>Datasets are pytorch files containing a dictionary with training, validation and test sets. Train, validation and test sets are custom dataset classes which inherit from the standard torch dataset class. Corresponding code an be found at https://github.com/HSG-AIML/NeurIPS_2021-Weight_Space_Learning.</p> <p>Datasets 41, 42, 43 and 44 are our dataset format wrapped around the zoos from Unterthiner et al, 2020 (https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy)<br> <br> Abstract:<br> Self-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn neural representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Distributed Predictive Drone Swarms in Cluttered Environments

<p>This folder contains data, videos, and supplementary material for the article titled &quot;Distributed Predictive Drone Swarms in Cluttered Environments&quot;.</p> <p>In the article, we present a Distributed Model Predictive Control (DMPC) algorithm for drone swarm navigation in two types of cluttered environments, i.e., a forest and a funnel-like environment.&nbsp;<br> &nbsp;</p> <p>The material in `zenodo_upload` is organized as follows.<br> 1. a data folder, with the logs of simulation and hardware experiments;<br> 2. an analysis folder, with Matlab scripts that analyze the logs in the data folder;<br> 3. a plotting folder, with Matlab functions used by the analysis scripts;<br> 4. an mp4 video file, on simulation and hardware experiments;<br> 5. a pdf, with supplementary materials.<br> <br> The `qp_swarm` folder contains MATLAB code for simulation experiments.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Supplementary data for "Single extreme storm sequence can offset decades of predicted shoreline retreat by sea-level rise"

<p>This dataset comprises topography and bathymetric data at three coastal locations in Australia (Narrabeen), UK (Perranporth) and used for the publication&nbsp;&quot;Single extreme storm sequence can offset decades of predicted shoreline retreat by sea-level rise&quot;. Please refer to readme files for metadata</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

KiSSim: Predicting off-targets from structural similarities in the kinome

<p><strong>KiSSim: Predicting off-targets from structural similarities in the kinome</strong></p> <p><strong>Project description.</strong></p> <p>KiSSim (Kinase Structural&nbsp;Similarity) is&nbsp;a novel fingerprint designed specifically for kinase pockets, allowing for similarity studies across the structurally covered kinome. The kinase fingerprint is based on the <a href="https://klifs.net/">KLIFS</a>&nbsp;pocket alignment, which defines 85 pocket residues for all kinase structures. This enables a residue-by-residue comparison without a computationally expensive alignment step.</p> <p>The pocket fingerprint encodes each pocket residue&rsquo;s spatial and physicochemical properties. The spatial properties describe the residue&rsquo;s position in relation to the kinase pocket center and important kinase subpockets, i.e. the hinge region, the DFG region, and the front pocket. The physicochemical properties encompass for each residue its size and pharmacophoric features, solvent exposure, and side chain orientation.</p> <p>Some datasets are not part of the `kissim_app` GitHub repository due to their size but can be downloaded from here to the respective kissim_app folders.</p> <p><strong>Data.</strong></p> <ul> <li>`20210902_KLIFS_HUMAN.tar.gz` --- save in `kissim_app/data/external/structures`</li> <li>`complete_SiteAlign.txt.gz` --- save in `kissim_app/data/external/sitealign`</li> </ul> <p><strong>Results.</strong></p> <ul> <li>`results.tar.bz2`--- save as `kissim_app/results`</li> </ul> <p>These are the KiSSim results:&nbsp;fingerprints,&nbsp;feature/fingerprint distances, kinase matrices, and kinase trees&nbsp;for structures in all (`all`), DFG-in (`dfg_in`), and DFG-out (`dfg_out`)&nbsp;conformation. In the case of the DFG-in conformation, we also have KiSSim runs with fingerprint subsets based on only residues that interact with certain ligands in KLIFS IFPs: Erlotinib (`dfg_in_IRE`), Imatinib (`dfg_in_STI`), Bosutinib (`dfg_in_DB8`), and Dopamapimod (`dfg_in_B96`). The folder contains README with a detailed file list.</p> <p><strong>Usage.</strong></p> <p>This dataset can be used to run the notebooks available on&nbsp;<a href="https://github.com/volkamerlab/kissim_app">https://github.com/volkamerlab/kissim_app</a>.</p> <ol> <li>Clone the kissim_app&nbsp;repository.</li> <li>Download the files provided here.</li> <li>If applicable, extract the archive content to the&nbsp;folders as indicated above and run the notebooks.</li> </ol> <pre><code class="language-bash">cd /path/to/your/download tar -xvf results.tar.bz2 -C /path/to/kissim_app/ tar -xvf 20210902_KLIFS_HUMAN.tar.bz2 -C /path/to/kissim_app/data/external/structures/ # In case you want the raw SiteAlign data mv complete_SiteAlign.txt.gz /path/to/kissim_app/data/external/sitealign</code></pre> <p><strong>Citation.</strong></p> <p>These&nbsp;datasets are&nbsp;part of the KiSSim publication: TBA</p>

openmit-licenseDec 2021View details →
zenodo44/100

19th Century United States Newspaper images predicted as Photographs with labels for "human", "animal", "human-structure" and "landscape"

<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (<a href="https://chroniclingamerica.loc.gov/">chroniclingamerica.loc.gov/</a>).&nbsp;</p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in&nbsp;<em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the&nbsp;<a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a>&nbsp;crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is &#39;photographs&#39;. This dataset contains a sample of these images with additional labels indicating if the photograph has one or more of the following labels: &quot;human&quot;, &quot;animal&quot;, &quot;human-structure&quot; and &quot;landscape&quot;</p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in `images.zip`</li> <li>`newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>`multi_label.csv` contains the labels for the images as a CSV file</li> <li>`annotations.csv` conains the labels for the images with additional metadata</li> </ul> <p>This dataset was created for use in an under-review Programming Historian tutorial (<a href="http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2">http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2</a>) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong></p> <p>The metadata CSV file contains the following columns:</p> <p>- filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url</p>

openother-openJan 2022View details →
zenodo44/100

Simulated NGS read datasets for prediction of novel fungal pathogens and multiple pathogen classes

<p>This repository contains simulated Illumina read datasets for novel fungal pathogen prediction and real-time detection of multiple pathogen classes. They were used to train the models hosted at <a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>.<br> The reads were simulated with Mason (<a href="https://www.seqan.de/apps/mason/">https://www.seqan.de/apps/mason/</a>) from genomes downloaded from NCBI, based on metadata stored in a manually curated database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>).</p> <p>We provide the following:</p> <p>1) An rds file describing assignment of fungal species from the database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>) to training, validation and test sets (TrainValTest_fungi.rds). A second rds file (TrainValTest_temporal.rds) includes species added within 12 weeks after the original datasets were compiled. Those species were used for a temporal benchmark.</p> <p>2) Fungal validation and test sets. Each contains 1.25 million, 250bp-long reads simulated from non-overlapping sets of human (&quot;pathogenic&quot;) or non-human (&quot;nonpathogenic&quot;) pathogens. The test set contains paired reads (&quot;_1&quot; and &quot;_2&quot; for the first and second mate). The number of reads per species is proportional to the respective genome length. An additional, temporal test set (*temporal*fasta.gz) includes 15 species added after 12 weeks from the consturction of the original datasets.</p> <p>3) Fungal training sets. They contain 250bp-long reads simulated from species not present in the validation or test sets. There are four variants:<br> 3a) &quot;low-coverage, linear&quot; - 20 million reads, number of reads per species proportional to genome length<br> 3b) &quot;low-coverage, logarithmic&quot; - 20 million reads, number of reads per species proportional to the logarithm of genome length (&quot;log&quot;)<br> 3c) &quot;high-coverage, linear&quot; - 240 million reads, number of reads per species proportional to genome length&nbsp; (&quot;24&quot;)<br> 3d) &quot;high-coverage, logarithmic&quot; - 240 million reads, number of reads per species proportional to the logarithm of genome length&nbsp; (&quot;24log&quot;)</p> <p>4) Training, validation and test sets for the multiclass models. They should be used together with the &quot;pathogenic&quot; read sets hosted at <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>. Here, we share sets for two of the four total classes:<br> 4a) The &#39;non-pathogen&#39; class is a mixture of &quot;nonpathogenic&quot; biacterial and viral read sets, concatenated and downsampled to the original read number (20M for training, 1.25M for validation and test). The training and validation sets contain mixed-length (25-20bp) simulated subreads (original sets hosted here: <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>). The test set contains 250bp long reads based on the test sets from here: <a href="https://zenodo.org/record/3678563">https://zenodo.org/record/3678563</a> and here: <a href="https://zenodo.org/record/4312525">https://zenodo.org/record/4312525</a>; it was also sorted by species.<br> 4b) Mixed-length versions of the &quot;pathogenic&quot; fungal training and validation sets, prepared by random shortening of the &quot;low-coverage&quot; read sets in the &quot;linear&quot; (_rn_) and &quot;logarithmic&quot; (_rn_*log_) flavours.</p> <p>See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a></p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Output data of the models used in "Comparison of six approaches to predicting droplet activation of surface active aerosol. Part 1: moderately surface active organics" by Vepsäläinen et al. (2022)

<p>Output data of the different models used in &quot;Comparison of six approaches to predicting droplet activation of surface active aerosol. Part 1: moderately surface active organics&quot; by Veps&auml;l&auml;inen et al. (2022).</p> <p>Output data is included for 50 nm particles containing malonic acid (mna), succinic acid (sca) and glutaric acid (glutarica), mixed with ammonium sulphate (AS) in different organic mass fractions.&nbsp;</p> <p>A plotter that allows the user to plot the K&ouml;hler curves, surface tensions and organic<br> partitioning factors during droplet growth from the model output data provided is included.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Snow equi-temperature metamorphism described by a phase-field model applicable on micro-tomographic images: prediction of microstructural and transport properties

<p>This dataset provides data described and used in the article submitted to Journal of Advances in Modeling Earth Systems &quot;Snow equi-temperature metamorphism described by a phase-field model applicable on micro-tomographic images: prediction of microstructural and transport properties&quot;.</p> <p>It contains .csv files with different properties computed on outputs of the model Snow3D simulating equi-temperature metamorphism. This micro-scale model was used here with experimental micro-tomographic snow images as input and returns series of 3-D images of snow showing features of equi-temperature metamorphism at different time steps as output.</p> <p>In this dataset, you will find two types of files:</p> <p>- the microstructural properties (density, specific surface area, covariance lengths, mean curvature) computed on&nbsp; the simulated images at different time steps.</p> <p>- the transport properties (effective conductivity, normalizes effective vapor diffusion coefficient, permeability) of the simulated images at different time steps.</p> <p>Finally, metadata_simulations.csv gather the information relative to the simulations.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Predicting Exon Criticality from Protein Sequence

<p>Exon ByPASS (predicting Exon-skipping Based in Protein amino acid SequenceS), predictions on test exons from Human and Mouse transcripts. The exons in the test set from the two genomes are those that are not predicted to be skippable in hg38 and mm10 annotation and are also exons that in-frame when skipped. The preprocessed data includes the ensemble transcript id and exon rank as well the amino acid sequence for the upstream, downstream, and exon of interest. Additionally, the table contains the output probability from the model in the last column. The input data is&nbsp;transformed data of the amino acid sequence that is the Exon ByPASS model can use as an input.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Predicted maps

<p>This dataset is the final product of research through the projects ANTARES (grant agreement No. 739570) and&nbsp;CYBELE (grant agreement No. 825355).&nbsp;The dataset consists of yield, protein content and selective harvesting soya maps&nbsp;at a resolution of 10 m.&nbsp;These maps were created by satellite images and soil properties&nbsp;data using machine learning algorithms. Maps are located in the Upper Austria region. Files with the name of &quot;map yield&quot; contain information about yield amount per pixel,&nbsp;while files with &quot;map protein&quot; denote&nbsp;parcels with predicted protein content.&nbsp;Also, the same protein map files contain an additional class column. Class 1&nbsp;indicates pixels where soya have good quality (protein content &gt; 41), while class 2 represents poorer quality.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "

<p>Datasets for&nbsp;preprint&nbsp;entitled &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi&quot;</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of&nbsp;<em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi&quot;.&nbsp;&nbsp;&nbsp;</p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict.&nbsp;For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22.&nbsp;In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using&nbsp;Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes.&nbsp;</p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset&nbsp;is made up of predicted protein tertiary structures representing the main member of each up-regulated&nbsp;<em>V. inaequalis</em> effector candidate family. Structures were predicted using&nbsp;Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>).&nbsp;In cases where&nbsp;the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2.&nbsp;Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold&nbsp;(<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>)&nbsp;open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input.&nbsp;</p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the &quot;The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence&nbsp;proteins from other fungi&quot; study. These structures were predicted using&nbsp;Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input.&nbsp;</p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>

opencc-by-2.0Feb 2022View details →
zenodo44/100

PaRoutes: a framework for benchmarking retrosynthesis route predictions

<p>PaRoutes is a framework for benchmarking multi-step retrosynthesis methods, i.e. route predictions.</p> <p>It provides:</p> <ul> <li>A curated reaction dataset for building one-step retrosynthesis models</li> <li>Two sets of 10,000 routes</li> <li>Two sets of stock molecules to use as stop-criterion for the search</li> </ul> <p>Homepage:&nbsp;<a href="https://github.com/MolecularAI/PaRoutes">https://github.com/MolecularAI/PaRoutes</a></p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Prediction stock of soil organic carbon in Argentina

<p>We standardized the Stocks soil organic carbon (SOC)&nbsp;at 0-30 cm depth for 5,073 soil samples. We spatially predicted SOC stock (kg/m2) using regression forest and associated prediction uncertainties using quantile regression forest at 1000 m resolution.&nbsp;Global accuracy based on cross-validation. We obtained a&nbsp;RMSE 2.624 and&nbsp;Rsquared&nbsp;0.464.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS

<p>We provide&nbsp;significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS.&nbsp;We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within &plusmn;1Mb, which we assumed they act through cis mechanisms. We include&nbsp;the model&rsquo;s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value &lt; 0.05 and R<sup>2</sup> &gt; 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Predicted proteome of Paratrimastix pyriformis

<p>The upload contains a predicted proteome from a genomic assembly of flagellate <em>Paratrimastix pyriformis</em> (Metamonada, Excavata). The publication describing the genomic study in detail is in progress. The final&nbsp;<em>P.&nbsp;pyriformis</em> genome was assembled into 650 scaffolds spanning 56,722,987 bp, with an N50&nbsp;=&nbsp;268,802 bp and a GC content of 60.92%. Manual and automatic gene prediction resulted in 13,532&nbsp;predicted protein-coding genes, which are the subject of this upload. Below we briefly describe, how the data were generated.&nbsp;&nbsp;</p> <p><strong>DNA isolation: </strong>Monoeukaryotic, xenic culture of <em>P. pyriformis</em> (strain RCP-MX, ATCC 50935) was maintained in the Sonneborn&#39;s Paramecium medium ATCC 802 at room temperature. The DNA was isolated from 15 litres of culture using two different kits. The gDNA samples for PacBio, Illumina HiSeq, and Illumina MiSeq sequencing were each isolated using the Qiagen DNeasy Blood &amp; Tissue Kit (Qiagen). The isolated gDNA was further ethanol-precipitated to increase the concentration and remove any contaminants. For nanopore sequencing, the DNA was isolated using Qiagen MagAttract HMW DNA Kit (Qiagen) according to the manufacturer&rsquo;s protocol.</p> <p><strong>Sequencing:</strong>&nbsp;We used three platforms to generate the sequence data - PacBio (RSII sequencer), Illumina (HiSeq and MiSeq) and Oxford Nanopore (two flow cells, MinION Mk1B).&nbsp;</p> <p><strong>Assembling:</strong> Sequencing quality was assessed with FastQC (Andrew 2010).&nbsp; For the Illumina data, adapter and quality trimming was performed using Trimmomatic 0.36 (Bolger et al. 2014), with a quality threshold of 15. For the nanopore data, trimming and removal of chimeric reads was performed using Porechop v0.2.3 (https://github.com/rrwick/Porechop).&nbsp;The initial assembly of the genomes was made only with the Nanopore and PacBio generated reads using Canu v1.7.1 assembler (Koren et al. 2017), with the corMinCoverage and corOutCoverage set to 0 and 100000 respectively. After assembly, the data were binned using tetraESOM (Haddad et al. 2009). The resulting eukaryotic bins were also checked using a combination of BLASTn and BLASTp and a scoring strategy based on the identity and coverage of the scaffold as described in (Treitli et al. 2019). After binning, the resulted genomic bins were polished in two phases. In the first phase, the scaffolds were polished using the raw reads generated by nanopore with Nanopolish (Loman et al. 2015). In the second phase, the resulting scaffolds generated by Nanopolish were further corrected using Illumina short reads with Pilon v1.21 (Walker et al. 2014). Finally, the genome assembly of <em>P. pyriformis</em> was further scaffolded with raw RNA-seq reads using Rascaf (Song et al. 2016).&nbsp;</p> <p><strong>Gene prediction:</strong>&nbsp;For <em>de novo</em> prediction of genes, first, we manually re-trained Augustus using a manually curated set of gene models. After re-training of Augustus, intron hints were generated from the RNAseq data and gene prediction was performed on repeat masked genomes using Augustus 3.2.3 (Stanke and Waack 2003). For polishing of the predicted genes, we mapped the transcriptome assemblies to the genome using PASA (Haas et al. 2003) and used the assembled transcripts by PASA as evidence for gene model polishing with EVM (Haas et al. 2008).&nbsp;</p> <p><strong>Protein annotation: </strong>Automatic annotation of the proteins was performed using KEGG Automatic Annotation Server (Moriya et al. 2007), as well as similarity searches using BLAST against NCBI nr protein database. Manual search and annotation were performed by searching the predicted proteome using BLAST and HMMER (Finn et al. 2011). Proteins of interest were manually investigated and if possible, the gene models were manually corrected.</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record