Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

43

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

43 results for “Pre-trained models”

Learn how ShareScore rates datasets ↗
zenodo48/100

Pre-trained models for segmentation and tracking of Coronal Bright Fronts from SDO AIA Base Difference images

<p>Here we present pretrained U-NET-based models followed by SDO AIA Base Difference(BD) validation set after intensity tresholding [-50;150] with predicted feature masks samples. &nbsp; &nbsp;&nbsp;<br>We provide a command-line Python utility for image segmentation using our CNNs designed to process images of solar eruptive phenomena. The https://gitlab.com/iahelio/helios_cnn repository includes regularly updated and newly published models.&nbsp;</p> <p>First model we present is designed to predict the likelihood of each pixel belonging to a certain class or feature in the solar image. A probabilistic output allows for a more nuanced interpretation of ambiguous region. The output can be converted into binary masks through thresholding. The range of values also gives insights into the model's confidence</p> <p>We also present sample segmentation results and the second model designed to produce binary masks.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Ecore Metamodels and EcoreBERT Pre-trained Language Model

<p>This dataset contains ecore metamodels from the MAR dataset&nbsp;transformed into tree representations.&nbsp;The original dataset can be found here:&nbsp;<a href="http://mar-search.org/experiments/models20/">http://mar-search.org/experiments/models20/</a></p> <p>The data contained in this repository were used to conduct the experiments in the paper: <strong>Recommending Metamodel Concepts during Modeling Activities with Pre-Trained Language Models.&nbsp;</strong>Link to the paper:&nbsp;<a href="https://arxiv.org/abs/2104.01642">https://arxiv.org/abs/2104.01642</a></p> <p>The data are organized as follows:</p> <ul> <li>model : our model trained on the tree representations of metamodels with RoBERTa architecture.</li> <li>tokenizers : the byte-level BPE tokenizer we used to train our model.</li> <li>train : the training data separated into a training and validation set.</li> <li>test : the test data of all experiments conducted in the paper.</li> </ul> <p>This data repository is linked with the following Github repository containing our code:&nbsp;<a href="https://github.com/mweyssow/ecore-bert">https://github.com/martiwey/metamodel-concepts-bert</a></p>

opencc-by-4.0Apr 2021View details →
zenodo40/100

BioVAE: a pre-trained latent variable language model for biomedical text mining

<p>We release BioVAE, the first large-scale pre-trained latent variable language model for the biomedical domain, which uses the OPTIMUS framework to train on large volumes of biomedical text.</p> <p>This version contains&nbsp;the pre-trained models for text mining tasks such as named entity recognition or&nbsp;relation extraction, and text generation task.</p> <p>Explanation of each file: (lt32: latent_size = 32, beta05: beta=0.5)</p> <ul> <li>pm-full-lt32-beta00</li> <li>pm-full-lt32-beta05</li> <li>pm-full-lt768-beta00</li> <li>pm-full-lt768-beta05</li> <li>pm-full-generation</li> </ul>

openapache2.0Nov 2021View details →
zenodo40/100

A novel strategy for fully automated segmentation of supratentorial meningiomas: Use of pre-trained models and inclusion of normal brain images

<p>This repository is accompanying MRI datasets under the journal, titled: <strong>A novel strategy for fully automated segmentation of supratentorial meningiomas: Use of pre-trained models and inclusion of normal brain images</strong>.&nbsp;</p> <p>Nii_data.tar.gz (zipped)&nbsp;file includes MRI images of&nbsp;all patients described in the paper that are formatted as .nii.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Training and validation data used to produce the pre-trained model for the TomoTwin paper.

<p>This datasets represents the training and validation data that was used to produce the pre-trained model for the TomoTwin paper. Please see 10.5281/zenodo.6637357 for the raw tomograms.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Semi-Supervised Pre-trained Foundation Model for 3D Structural Feature Analysis of Seismic Images

<p>Codes, trained model, and datasets for the paper "Semi-Supervised Pre-trained Foundation Model for 3D Structural Feature Analysis of Seismic Images".</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

An Empirical Comparison of Pre-Trained Models of Source Code

<p>The replication package&nbsp;of the paper &quot;An Empirical Comparison of Pre-Trained Models of Source Code&quot;. For the source code, please refer to&nbsp;<a href="https://github.com/NougatCA/FineTuner">https://github.com/NougatCA/FineTuner</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

UniFMIR: Pre-training a Foundation Model for Universal Fluorescence Microscopy Image Restoration

<p>This repository contains the preprocessed dataset for&nbsp;[UniFMIR](https://github.com/cxm12/UNiFMIR/).&nbsp;All training and test data involved in the experiments are publicly available datasets. Licenses of the original dataset are applied.&nbsp;You can refer to the Github repository for details.</p> <p>* The 3D denoising/isotropic reconstruction/projection datasets can be downloaded from [Content Aware Image Restoration dataset](https://publications.mpi-cbg.de/publications-sites/7207/). `Projection_Flywing/train_data/my_training_data.npz` are generated according to the [CSBDeep](http://csbdeep.bioimagecomputing.com/doc/).</p> <p>* The SR dataset can be downloaded from [BioSR dataset](https://doi.org/10.6084/m9.figshare.13264793). The dataset is&nbsp;augmented&nbsp;according to the instructions in [DFCAN](https://github.com/qc17-THU/DL-SR/tree/main#train-a-new-model) and `my_training_data.npz` files are&nbsp;generated&nbsp;following [CSBDeep](http://csbdeep.bioimagecomputing.com/doc/datagen.html).&nbsp;</p> <p>* The Volumetric reconstruction dataset are from [VCD-LFM dataset](https://doi.org/10.5281/zenodo.4390067).&nbsp; The dataset is prepared according to the instructions in [VCD-Net](https://github.com/feilab-hust/VCD-Net).</p> <p>* DeepBacs dataset can be downloaded from [DeepBacs dataset](https://zenodo.org/record/6460867). We split the dataset into 5 folds for cross-validation. Shareloc dataset can be downloaded from [Shareloc dataset](https://zenodo.org/record/7234161).</p> <p>&nbsp;</p> <p>The data paths should be as follows:</p> <p>```</p> <p>VCD/vcdnet/</p> <p>CSB/DataSet/</p> <p>&nbsp; &nbsp; Denoising_Planaria/</p> <p>&nbsp; &nbsp; Denoising_Tribolium/</p> <p>&nbsp; &nbsp; Isotropic/Isotropic_Liver/</p> <p>&nbsp; &nbsp; Projection_Flywing/</p> <p>&nbsp; &nbsp; BioSR_WF_to_SIM/DL-SR-main/dataset/</p> <p>&nbsp; &nbsp; Synthetic_tubulin_gfp/</p> <p>&nbsp; &nbsp; Synthetic_tubulin_granules/</p> <p>DeepBacs/</p> <p>Shareloc/</p> <p>```</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Phenanthrene: TD-DFTB datasets, pre-trained SchNet models and initial coniditions for TSH

<p><em>Data associated with the paper entitled </em></p> <p><strong>On application of Deep Learning to simplified quantum-classical dynamics in electronically excited states</strong></p> <ol> <li>Three TD-DFTB datasets&nbsp;(<strong>sX_10_force.db</strong>) have been produced using the <a href="https://wiki.fysik.dtu.dk/ase/">Atomic Simulation Environment</a> (ASE) coupled to <a href="http://demon-nano.ups-tlse.fr/">deMon-Nano</a> code for the linear response Time-Dependent Density Functional based Tight-Binding (TD-DFTB) calculations. Each dataset contains 10000 TD-DFTB electronic structure calculations for a given excited singlet state (S<sub>2</sub>/S<sub>3</sub>/S<sub>4</sub>) of a neutral phenanthrene molecule. Each database entry contains Cartesian atomic coordinates as well as potential energy and atomic forces for a given excited state at a given geometry. Since ASE has been used, all physical quantities are stored in the corresponding units (e.g. eV for energy or eV/&Aring; for forces). The file format is SQLite as provided by the ASE;</li> <li>Three pre-trained Deep Learning models (<strong>best_model_sX</strong>) for a given excited singlet state have been produced using <a href="https://schnetpack.readthedocs.io/en/stable/">SchNetPack</a> package, which implements the SchNet architecture for atomistic simulations. Each model has been trained using the corresponding TD-DFTB dataset from #1. The file format is binary as provided by the SchNetPack;</li> <li><a href="https://zenodo.org/api/files/f1925cb5-66a8-4c6f-809b-3414f0cbc1d5/500_init_conditions.tar.gz"><strong>500_init_conditions.tar.gz</strong>&nbsp;</a> contains 500 initial conditions (Cartesian coordinates and velocities), which can be used for Trajectory Surface Hopping (TSH) simulations with or without the pre-trained models from #2.</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Pre-trained word2vec models for ``Easy over Hard: A Case Study on Deep Learning''

<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds.</p> <p> </p> <p>More details about how to use it, please see paper </p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

MOST-GAN Pre-trained Model

<p><strong>Introduction</strong></p> <p>Recent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningful and disentangled way. Prior approaches aim to achieve such semantic control and disentanglement within the latent space of a previously trained GAN. In contrast, we propose a framework that a priori models physical attributes of the face such as 3D shape, albedo, pose, and lighting explicitly, thus providing disentanglement by design. Our method, MOST-GAN, integrates the expressive power and photorealism of style-based GANs with the physical disentanglement and flexibility of nonlinear 3D morphable models, which we couple with a state-of-the-art 2D hair manipulation network. MOST-GAN achieves photorealistic manipulation of portrait images with fully disentangled 3D control over their physical attributes, enabling extreme manipulation of lighting, facial expression, and pose variations up to full profile view.&nbsp;</p> <p>To foster further research into this topic, we are publicly releasing our pre-trained model for MOST-GAN. Please see our AAAI paper titled [MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image Manipulation](https://arxiv.org/abs/2111.01048) for details.</p> <p><strong>At a Glance</strong></p> <p>-The size of the unzipped model is ~300MB.</p> <p>-The unzipped folder contains: (i) a README.md file and (ii) ./checkpoints/checkpoint01.pt pre-trained model. The pre-trained model could be loaded in our publicly released MOST-GAN implementation.</p> <p><strong>Citation</strong></p> <p>If you use the MOST-GAN data in your research, please cite our paper:</p> <pre><code>@inproceedings{medin2022most, title={MOST-GAN: 3D morphable StyleGAN for disentangled face image manipulation}, author={Medin, Safa C and Egger, Bernhard and Cherian, Anoop and Wang, Ye and Tenenbaum, Joshua B and Liu, Xiaoming and Marks, Tim K}, booktitle={Proceedings of the AAAI conference on artificial intelligence}, volume={36}, number={2}, pages={1962--1971}, year={2022} } </code></pre> <p><strong>License</strong></p> <p>The MOST-GAN data is released under&nbsp;<a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA-4.0 license</a>.</p> <p>All data:</p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2022,2023 SPDX-License-Identifier: CC-BY-SA-4.0 </code></pre> <p>&nbsp;</p>

opencc-by-sa-4.0Aug 2023View details →
zenodo36/100

LUVLi Pre-trained Model

<p><strong>Introduction</strong></p> <p>Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting landmark locations, associated uncertainties of these predicted locations, and landmark visibilities. We model these as mixed random variables and estimate them using a deep network trained with our proposed Location, Uncertainty, and Visibility Likelihood (LUVLi) loss. In addition, we release an entirely new labeling of a large face alignment dataset with over 19,000 face images in a full range of head poses. Each face is manually labeled with the ground-truth locations of 68 landmarks, with the additional information of whether each landmark is unoccluded, self-occluded (due to extreme head poses), or externally occluded. Not only does our joint estimation yield accurate estimates of the uncertainty of predicted landmark locations, but it also yields state-of-the-art estimates for the landmark locations themselves on multiple standard face alignment datasets. Our method&rsquo;s estimates of the uncertainty of predicted landmark locations could be used to automatically identify input images on which face alignment fails, which can be critical for downstream tasks.</p> <p>To foster further research into this topic, we are publicly releasing our pre-trained LUVLi models. Please see our CVPR 2020 paper titled <a href="https://arxiv.org/abs/2004.02980">LUVLi Face Alignment: Estimating Landmarks&rsquo; Location, Uncertainty, and Visibility Likelihood</a> for details</p> <p><strong>At a Glance</strong></p> <p>-The size of the unzipped model is ~700MB.</p> <p>-The unzipped folder contains: (i) a README.md file and (ii) pre-trained models and logs. The pre-trained models could be loaded in our publicly released LUVLi implementation.</p> <p><strong>Other Resources</strong></p> <p><strong>Citation</strong></p> <p>If you use the LUVLi data in your research, please cite our paper:</p> <pre><code>@inproceedings{kumar2020luvli, title={{LUVLi} Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood}, author={Kumar, Abhinav and Marks, Tim K. and Mou, Wenxuan and Wang, Ye and Jones, Michael and Cherian, Anoop and Koike-Akino, Toshiaki and Liu, Xiaoming and Feng, Chen}, booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year={2020} } </code></pre> <p><strong>License</strong></p> <p>The LUVLi data is released under&nbsp;<a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA-4.0 license</a>.</p> <p>All data:</p> <pre><code>Created by Mitsubishi Electric Research Laboratories (MERL), 2022,2023 SPDX-License-Identifier: CC-BY-SA-4.0 </code></pre>

opencc-by-sa-4.0Mar 2024View details →
zenodo36/100

Pre-trained Models for SMP Classification and Segmentation

<p>This dataset provides access to pre-trained models that were used for SnowMicroPen profile classification and segmentation. The models were trained on a part of the MOSAiC SMP dataset, available on <a href="https://doi.pangaea.de/10.1594/PANGAEA.935554">https://doi.pangaea.de/10.1594/PANGAEA.935554</a>. The labeled training data consists mostly of profiles from leg three of the expedition (January - May 2020), some profiles from leg one and two, and no profiles from leg four. Please refer to the snowdragon GitHub repository (<a href="https://github.com/liellnima/snowdragon">https://github.com/liellnima/snowdragon</a>) to access the models&#39; training code and be directed to current publications.</p> <p>The following trained models are available here (alphabetically ordered):</p> <ul> <li>Artificial neural networks <ul> <li>Bi-directional long short-term memory <em>(blstm.hdf5)</em></li> <li>Encoder-decoder <em>(enc_dec.hdf5)</em></li> <li>Long short-term memory <em>(lstm.hdf5)</em></li> </ul> </li> <li>Baseline <ul> <li>Majority vote classifier <em>(baseline.model)</em></li> </ul> </li> <li>Semi-supervised models <ul> <li>Cluster-then-predict models: <ul> <li>Bayesian Gaussian mixture model <em>(gmm.model)</em></li> <li>Bayesian mixture model <em>(bmm.model)</em></li> <li>K-means clustering <em>(kmeans.model)</em></li> </ul> </li> <li>Label propagation <em>(label_spreading.model)</em></li> <li>Self-trained classifier <em>(self_trainer.model)</em></li> </ul> </li> <li>Supervised models <ul> <li>Balanced random forest <em>(rf_bal.model)</em></li> <li>Easy ensemble <em>(easy_ensemble.model)</em></li> <li>K-nearest neighbors <em>(knn.model)</em></li> <li>Random forest <em>(rf.model)</em></li> <li>Support vector machines <em>(svm.model)</em></li> </ul> </li> </ul> <p><br> <em>Loading Instructions:</em><br> The models with the file-ending &quot;.model&quot; are pickeled Python objects and can be loaded with ``pickle.load(your_model.model)``. The random forest must be loaded with ``joblib.load(rf.model)``. All artificial neural networks are h5py.File objects (tf.keras models) and can be loaded with ``tf.keras.models.load_model(your_ann.model)``.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Simulations dataset and pre-trained models of "Deep learning in real-time on the astrophysical data obtained from the Čerenkov CTA Observatory" Ph.D. project

<p>Ph.D. project datasets and models release, <br><em>Deep learning in real-time on the astrophysical data obtained from the Čerenkov CTA Observatory.</em></p>

opencc-by-4.0May 2024View details →
zenodo36/100

MS2Query pre-trained embeddings and models

<p>The models, embeddings, sqlite file with metadata and classifiers identifiers needed for running MS2Query (https://github.com/iomega/ms2query) MS2Query positive mode model and library. The library consists of the GNPS library, MassBank, MoNA and Brungs et al. 's library.&nbsp; A bug with the compound classes was fixed. This version is the downloaded model with MS2Query version &gt;= 1.5.3</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Towards safe human-to-robot handovers of unknown containers: pre-trained models and 3D hand keypoints annotations

<p>This repository contains additional data to be used with the implementation of the real-to-simulation framework of the paper <em>Towards safe human-to-robot handovers of unknown containers</em>. The data include pre-trained models and annotations of the 3D hand poses for selected recordings from the public training and testing sets of <a href="http://corsmal.eecs.qmul.ac.uk/containers_manip.html">CORSMAL Container Manipulation (CCM) dataset</a>. The pre-trained models are used for classifying the filling type and filling level of a container. 3D hand poses are annotated as 21 keypoints based on the <a href="https://github.com/CMU-Perceptual-Computing-Lab/openpose">OpenPose</a>&nbsp;format.</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

Pre-trained models for Vietnamese Natural Resources and Environment Domain

<p>Two Pre-trained models for Vietnamese Natural Resources and Environment Domain:</p> <p>- fastText</p> <p>- Word2vec</p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

"An efficient ptychography reconstruction strategy through fine-tuning of large pre-trained deep learning model" train and test data

<ul><li>Model for the article "An efficient ptychography reconstruction strategy through fine-tuning of large pre-trained deep learning model".</li><li>The &nbsp;.pth file is the pre-trained PtyNet-S model and the fine-tuned PtyNet-B model.</li><li>Please contact panxy@ihep.ac.cn if you have any questions.</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)

<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)

<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record