Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
99
datasets available to search
ShareScore release 0.9.0
Dataset results
99 results for “Transfer Learning”
A multi-station volcano-tectonic earthquakes monitoring based on Transfer Learning techniques.
<p>A multi-station volcano-tectonic earthquakes monitoring based on Transfer Learning techniques.</p> <p>Manuel Titos (1), Ligdamis Gutiérrez (2,3), Carmen Benítez (1), Pablo Rey Devesa (2,3), Ivan Koulakov (4) and Jesús. M. Ibáñez (2,3)</p> <p><br> <strong>Institutions associated:</strong></p> <p>(1) CITIC, Department of Signal Processing, Telematic and Communications, University of Granada, 18071. Granada. Spain.<br> (2) Department of Theoretical Physics and Cosmos. Science Faculty. Avd. Fuentenueva s/n. University of Granada. 18071. Granada. Spain.<br> (3) Andalusian Institute of Geophysiscs. Campus de Cartuja. University of Granada. C/Profesor Clavera 12. 18071. Granada. Spain.<br> (4) Laboratory for Seismic Forward and Inverse Problems, Institute of Petroleum Geology and Geophysics, Siberian Branch of the Russian Academy of Sciences, Novosibirsk, Russia</p> <p><br> <strong>Acknowledgment:</strong></p> <p>This is a short text to acknowledge the contributions of specific colleagues, institutions, or agencies that aided the efforts of the authors.</p> <p>a) This work is part of the research by the Spanish FEMALE project (PID2019-106260GB-I00). <strong>FEMALE </strong>(<em>Forecasting Volcanic Eruptions Using Signal Processing and Machine Learning Techniques on Seismic Signals</em>) https://femalevolcanoes.es/</p> <p><br> b) JMI and LG were partially funded by the Spanish project PROOF-FOREVER (EUR2022.134044).</p> <p><br> <strong>Keywords:</strong></p> <p>Automatic volcanic monitoring, real-time monitoring, Artificial Intelligence, Transfer Learning, Recurrent Neural Networks, Temporal Convolutional Networks.</p> <p> </p> <p><strong>Data availability statement:</strong></p> <p>Seismic data from Bezymianny volcano (2017), Kamchatka, Russia.</p> <p> </p> <p><strong>Contents:</strong></p> <p>Seismic Data from Bezymianny volcano recorded at stations BZ01, BZ02, BZ06 and BZ10.<br> The data represent the vertical component of the seismic signal, associated to the period analyzed in the study:</p>
Data from: RockNet: Rockfall and earthquake detection and association via multitask learning and transfer learning
Open the record for dataset details and reuse information.
Within and cross species predictions of plant specialized metabolism genes using transfer learning
<p>Datasets for <em>Within and cross species predictions of plant specialized metabolism genes using transfer learning.</em> </p> <ol> <li>Dataset 1: All gene features used in the full feature machine learning models including expression, co-expression, evolutionary, duplication, and protein domain features.</li> <li>Dataset 2: All gene features used in shared feature machine learning models including expression, evolutionary, duplication, and protein domain features.</li> <li>Table S1: All gene annotations from TomatoCyc or manual annotation.</li> <li>Table S2: All model scores.</li> <li>Table S3: All gene scores and predictions from each model.</li> <li>Table S4: Feature importance for 5 models.</li> <li>Table S5: Statistical analysis between classes for binary and continuous feature data.</li> <li>Table S6: RNAseq datasets used in analysis.</li> </ol>
Fussy Transfer Learning dataset
<p>Companion dataset repository to be used with <a href="https://github.com/MatZar01/Fussy-Transfer-Learning">Fussy Transfer Learning - Github</a> project methods.</p>
Biological and environmental data for a study on transferability of statistical and machine learning models using North Sea Macrozoobenthos
<p>General</p> <p>Data documented here are not the product of our research but was scraped from various sources and processed - so no genuine reupload. This collection is a contribution to reproduceable reseach. All datasets are given in "RData" binary format</p> <p> </p> <p>Data description</p> <p>majornorthseabenthos </p> <p>This is macrozoobenthos data as data frame scraped from the GBIF repository (gbif.org). Species are Corbula gibba, Tellina fabula, Turritella communis, Euspira pulchella, Corystes cassive- launus, Upogebia deltaura, Lanice conchilega, Nephtys hombergii, Echinocardium cordatum, and Amphiura filiformis. data was postprocessed to have only single occurrence fon the approxinatel 1x1 km grid used for this study. Also, occurrences closer than 5 km close to shore were removed - including occurrences on land.</p> <p> </p> <p>Predictors</p> <p>A SpatialPixelsDataFrame in EPSG 4326 with five layers: Median grain size in micrometers, mud content in percent (both MUDAB database), water depth in meters above MSL (Weatherall et al, 2015), modelled average bottom shear stress from waves in N/sqrm (The Wamdi Group, 1988) and climatologival average winter bottom water temperature in deg. C (Stips et al, 2004).</p> <p> </p> <p> </p> <p>References</p> <p>Stips A, Bolding K, Pohlmann T, Burchard H (2004) Simulating the temporal and spatial dy- namics of the North Sea using the new model GETM (general estuarine transport model). Ocean Dynamics 54(2):266–283</p> <p>The Wamdi Group (1988) The WAM model-a third generation ocean wave prediction model. Journal of Physical Oceanography 18(12):1775–1810</p> <p>Weatherall P, Marks K, Jakobsson M, Schmitt T, Tani S, Arndt JE, Rovere M, Chayes D, Ferrini V, Wigley R (2015) A new digital bathymetric model of the world’s oceans. Earth and Space Science 2(8):331–345</p> <p> </p>
Data and script pipeline for: Common to rare transfer learning (CORAL) enables inference and prediction for a quarter million rare Malagasy arthropods
<p>The scripts and the data provided in this depository demonstrate how to apply the approach described in the paper "<strong>Common to rare transfer learning (CORAL) enables inference and prediction for a quarter million rare Malagasy arthropods</strong>" by Ovaskainen et al. Here we summarize how to use the software with a small, simulated dataset, with running time less than a minute in a typical laptop (Demo 1); (2) how to apply the analyses presented in the paper for a small subset of the data, with running time of ca. one hour in a powerful laptop (Demo 2); how to reproduce the full analyses presented in the paper, with running time up to several days, depending on the computational resources (Demo 3). The Demos 1 and 2 are aimed to be user-friendly starting points for understanding and testing how to implement CORAL. The Demo 3 is included mainly for reproducibility.</p> <p> </p> <p><strong><u>System requirements</u></strong></p> <p> </p> <p>· The software can be used in any operating system where R can be installed.</p> <p>· We have developed and tested the software in a windows environment with R version 4.3.1.</p> <p>· Demo 1 requires the R-packages phytools (2.1-1), MASS (7.3-60), Hmsc (3.3-3), pROC (1.18.5) and MCMCpack (1.7-0).</p> <p>· Demo 2 requires the R-packages phytools (2.1-1), MASS (7.3-60), Hmsc (3.3-3), pROC (1.18.5) and MCMCpack (1.7-0).</p> <p>· Demo 3 requires the R-packages phytools (2.1-1), MASS (7.3-60), Hmsc (3.3-3), pROC (1.18.5) and MCMCpack (1.7-0), jsonify (1.2.2), buildmer (2.11), colorspace (2.1-0), matlib (0.9.6), vioplot (0.4.0), MLmetrics (1.1.3) and ggplot2 (3.5.0).</p> <p>· The use of the software does not require any non-standard hardware.</p> <p> </p> <p><strong><u>Installation guide</u></strong></p> <p> </p> <p>· The CORAL functions are implemented in Hmsc (3.3-3). The software that applies the is presented as a R-pipeline and thus it does not require any installation other than installation of R.</p> <p> </p> <p><strong><u>Demo 1: Software demo with simulated data</u></strong></p> <p><strong> </strong></p> <p>The software demonstration consists of two R-markdown files:</p> <p> </p> <p>· <strong>D01_software_demo_simulate_data.</strong> This script creates a simulated dataset of 100 species on 200 sampling units. The species occurrences are simulated with a probit model that assumes phylogenetically structured responses to two environmental predictors. The pipeline saves all the data needed to data analysis in the file allDataDemo.RData: XData (the first predictor; the second one is not provided in the dataset as it is assumed to remain unknown for the user), Y (species occurrence data), phy (phylogenetic tree), studyDesign (list of sampling units). Additionally, true values used for data generation are save in the file trueValuesDemo.RData: LF (the second environmental predictor that will be estimated through a latent factor approach), and beta (species responses to environmental predictors).</p> <p>· <strong>D02_software_demo_apply_CORAL</strong>. This script loads the data generated by the script D01 and applies the CORAL approach to it. The script demonstrates the informativeness of the CORAL priors, the higher predictive power of CORAL models than baseline models, and the ability of CORAL to estimate the true values used for data generation.</p> <p> </p> <p>Both markdown files provide more detailed information and illustrations. The provided html file shows the expected output. The running time of the demonstration is very short, from few seconds to at most one minute.</p> <p> </p> <p><strong><u>Demo 2: Software demo with a small subset of the data used in the paper</u></strong></p> <p><strong> </strong></p> <p>The software demonstration consists of one R-markdown file:</p> <p> </p> <p><strong>MA_small_demo.</strong> This script uses the CORAL functions in HMSC to analyze a small subset of the Malagasy arthropod data. In this demo, we define rare species as those with prevalence at least 40 and less than 50, and common species as those with prevalence at least 200. This leaves 51 species to the backbone model and 460 rare species modelled through the CORAL approach. The script assess model fit for CORAL priors, CORAL posteriors, and null models. It further visualizes the responses of both the common and the rare species to the included predictors.</p> <p> </p> <p><strong><u>Scripts and data for reproducing the results presented in the paper (Demo 3)</u></strong></p> <p> </p> <p>The input data for the script pipeline is the file “allData.RData”. This file includes the metadata (meta), the response matrix (Y), and the taxonomical information (taxonomy). Each file in the pipeline below depends on the outputs of previous files: they must be run in order. The first six files are used for fitting the backbone HMSC model and calculating parameters for the CORAL prior:</p> <p> </p> <p>· <strong>S01_define_Hmsc_model</strong> - defines the initial HMSC model with fixed effects and sample- and site-level random effects.</p> <p>· <strong>S02_export_Hmsc_model</strong> - prepares the initial model for HPC sampling for fitting with Hmsc-HPC. Fitting of the model can be then done in an HPC environment with the bash file generated by the script. Computationally intensive.</p> <p>· <strong>S03_import_posterior</strong> – imports the posterior distributions sampled by the initial model.</p> <p>· <strong>S04_define_second_stage_Hmsc_model</strong> - extracts latent factors from the initial model and defines the backbone model. This is then sampled using the same S02 export + S03 import scripts. Computationally intensive.</p> <p>· <strong>S05_visualize_backbone_model</strong> – check backbone model quality with visual/numerical summaries. Generates Fig. 2 of the paper.</p> <p>· <strong>S06_construct_coral_priors</strong> – calculate CORAL prior parameters.</p> <p> </p> <p>The remaining scripts evaluate the model:</p> <p>· <strong>S07_evaluate_prior_predictionss </strong>– use the CORAL prior to predict rare species presence/absences and evaluate the predictions in terms of AUC. Generates Fig. 3 of the paper.</p> <p>· <strong>S08_make_training_test_split</strong><em> </em>– generate train/test splits for cross-validation ensuring at least 40% of positive samples are in each partition.</p> <p>· <strong>S09_cross-validate</strong><em> </em>– fit CORAL and the baseline model to the train/test splits and calculate performance summaries. Note: we ran this once with the initial train/test split and then again with on the inverse split (i.e., <em>training = ! training</em> in the code, see comment). The paper presents the average results across these two splits. Computationally intensive.</p> <p>· <strong>S10_show_cross-validation_results</strong><em> </em>– Make plots visualizing AUC/Tjur’s R<sup>2</sup> produced by cross-validation. Generates Fig. 4 of the paper.</p> <p>· <strong>S11a_fit_coral_models</strong><em> </em>– Fit the CORAL model to all 250k rare species. Computationally intensive.</p> <p>· <strong>S11b_fit_baseline_models</strong><em> </em>– Fit the baseline model to all 250k rare species. Computationally intensive.</p> <p>· <strong>S12_compare_posterior_inference</strong><em> ­­</em>– compare posterior climate predictions using CORAL and baseline models on selected species, as well as variance reduction for all species. Generates Fig. 5 of the paper.</p> <p> </p> <p>Pre-processing scripts:</p> <p>· <strong>P01_preprocess_sequence_data.R </strong>– Reads in the outputs of the bioinformatics pipeline and converts them into R-objects.</p> <p>· <strong>P02_download_climatic_data.R </strong>– Downloads the climatic data from "sis-biodiversity-era5-global” and adds that to metadata.</p> <p>· <strong>P03_construct_Y_matrix.R </strong>– Converts the response matrix from a sparse data format to regular matrix. Saves “allData.RData”, which includes the metadata (meta), the response matrix (Y), and the taxonomical information (taxonomy).</p> <p> </p> <p>Computationally intensive files had runtimes of 5-24 hours on high-performance machines. Preliminary testing suggests runtimes of over 100 hours on a standard laptop.</p> <p><em> </em></p> <p><strong><u>ENA Accession numbers</u></strong></p> <p>All raw sequence data are archived on mBRAVE and are publicly available in the European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena; project accession number PRJEB86111; run accession numbers ERR15018787-ERR15009869; sample IDs for each accession and download URLs are provided in the file ENA_read_accessions.tsv).</p>
Data set: Al-Biruni Earth Radius Optimization with Deep Transfer Learning based Scene Image Classification on Remote Sensing Imagery
Open the record for dataset details and reuse information.
Transfer learning reveals sequence determinants of the quantitative response to transcription factor dosage
<p>Processed data and code for "Transfer learning reveals sequence determinants of the quantitative response to transcription factor dosage," Naqvi et al 2025.</p> <p>Directory is organized into the following subfolders, each tar'ed and gzipped:</p> <p><strong>data_analysis.tar.gz - Processed data for modulation of TWIST1 levels and calculation of RE responsiveness to TWIST1 dosage</strong></p> <ul> <li>atac_design.txt - design matrix for ATAC-seq TWIST1 titration samples</li> <li>all.sub.150bpclust.greater2.500bp.merge.TWIST1.titr.ATAC.counts.txt - ATAC-seq counts from all samples over all reproducible ATAC-seq peak regions, as defined in Naqvi et al 2023</li> <li>atac_deseq_fitmodels_moded50.R - R code for calculating new version of ED50 and response to full depletion from TWIST1 titration data (note, uses drm.R function from <a href="https://doi.org/10.5281/zenodo.7689948">10.5281/zenodo.7689948</a>, install drc() with this version to avoid errors)</li> </ul> <p><strong>baseline_models.tar.gz - Code and data for training baseline models to predict RE responsiveness to SOX9/TWIST1 dosage</strong></p> <ul> <li>{sox9|twist1}.{0v100|ed50}.{train|valid|test}.txt - Training/testing/validation data (ED50 or full TF depletion effect for SOX9 or TWIST1), split into train/test/validation folds</li> <li>HOCOMOCOv11_core_HUMAN_mono_jaspar_format.all.sub.150bpclust.greater2.500bp.merge.minus300bp.p01.maxscore.mat.cpg.gc.basemean.txt.gz - matrix of predictors for all REs. Quantitative encoding of PWM match for all HOCOMOCO motifs + CpG + GC content, plus unperturbed ATAC-seq signal</li> <li>train_baseline.R - R code to train baseline (LASSO regression or random forest) models using predictor matrix and the provided training data. <ul> <li>Note: training the random forest to predict full TF depletion is computationally intensive because it is across all REs, if doing this run on CPU for ~6 hrs. </li> </ul> </li> </ul> <p><strong>chrombpnet_models.tar.gz - Remainder of code, data, and models for fine-tuning and interpreting ChromBPNet mdoels to </strong><strong>predict RE responsiveness to SOX9/TWIST1 dosage</strong></p> <ul> <li>Fine-tuning code, data, models <ul> <li>{all|sox9.direct|twist1.bound.down}.{train|valid|test}.{ed50|0v100.log2fc}.txt - Training/testing/validation data (ED50 or full TF depletion effect for SOX9 or TWIST1), split into train/test/validation folds</li> <li>pretrained.unperturbed.chrombpnet.h5 - Pretrained model of unperturbed ATAC-seq signal in CNCCs, obtained by running ChromBPNet (https://github.com/kundajelab/chrombpnet) on DMSO-treated SOX9/TWIST1-tagged ATAC-seq data</li> <li>finetune_chrombpnet.py - code for fine-tuning the pretrained model for any of the relevant prediction tasks (ED50/ effect of full TF depletion for SOX9/TWIST1)</li> <li>best.model.chrombpnet.{0v100|ed50}.{sox9|twist1}.h5 - output of finetune_chrombpnet.py, best model after 10 training epochs for the indicated task</li> <li>chrombpnet.{0v100|ed50}.{sox9|twist1}.contrib.{h5|bw} - contribution scores for the indicated predictive model, obtained by running chrombpnet contribs_bw on the corresponding model h5 file.</li> <li>chrombpnet.{0v100|ed50}.{sox9|twist1}.contrib.modisco.{h5|bw} - TF-MoDIsCo output from the corresponding contribution score file</li> </ul> </li> <li>Interpretation code, data, models <ul> <li>contrib_h5_to_projshap_npy.py - code to convert contrib .h5 files into .npy files containing projected SHAP scores (required because the CWM matching code takes this format of contribution scores)</li> <li>sox9.direct.10col.bed, twist1.bound.down.10col.uniq.bed - regions over which CWMs will be matched (likely direct targets of each TF)</li> <li>match_cwms.py - Python code to match individual CWM instances. Takes as input: modisco .h5 file, SHAP .npy file, bed file of regions to be matched. Output is a bed file of all CWM matches (not pruned, contains many redundant matches).</li> <li>chrombpnet.ed50.{sox9|twist1}.contrib.perc05.matchperc10.allmatch.bed - output of match_cwms.py </li> <li>take_max_overlap.py - code to merge output of match_cwms.py into clusters, and then take the maximum (length-normalized) match score in each cluster as the representative CWM match of that cluster. Requires upstream bedtools commands to be piped in, see example usage in file. </li> <li>chrombpnet.ed50.{sox9|twist1}.contrib.perc05.matchperc10.allmatch.maxoverlap.bed - output of take_max_overlap.py. These CWM instances are the ones used throughout the paper.</li> </ul> </li> </ul> <p><strong>modisco_reports.zip -</strong><strong> TF-MoDIsCo reports from running on the fine-tuned ChromBPNet models</strong></p> <ul> <li>modisco_report_{sox9|twist1}_{0v100|ed50}: folders containing images of discovered CWMs and HTMLs/PDFs of summarized reports from running TF-MoDisCo on the indicated fine-tuned ChromBPNet model</li> </ul> <p><strong>chrombpnet_models_supp.tar.gz - Alternative ChromBPNet mdoels to </strong><strong>predict SOX9/TWIST1 ED50 using varying definitons of direct targets</strong></p> <ul> <li> <p>best.model.chrombpnet.ed50.twist1.3hdn.h5 - TWIST1 direct targets defined using response to full 3h depletion (as was done for SOX9 throughout the rest of the paper)</p> </li> <li> <p>best.model.chrombpnet.ed50.sox9.v5chip.h5 - SOX9 direct targets defined using V5 ChIP-seq from SOX9-tagged lines (as was done for TWIST1 throughout the paper)</p> </li> </ul> <p><strong>mirny_model.tar.gz - Code and data for analyzing and fitting Mirny model of TF-nucleosome competition to observed RE dosage response curves</strong></p> <ul> <li>twist1.strong.multi.only.ed50.cutoff.true.hill.txt - ED50 and signed hill coefficients for all TWIST1-dependent REs with only buffering Coordinators (mostly one or two) and no other TFs' binding sites. "ed50_new" is the ED50 calculation used in this paper. </li> <li>twist1.strong.weak{1|2|3}.ed50.cutoff.true.hill.txt - ED50 and signed hill coefficients for all TWIST1-dependent REs with only buffering Coordinators (mostly one or two) and the indicated number of sensitizing (weak) Coordinators and no other TFs' binding sites. "ed50_new" is the ED50 calculation used in this paper. </li> <li>MirnyModelAnalysis.py - Python code for analysis of Mirny model of TF-nucleosome competition. Contains implementations of analytic solutions, as well as code to fit model to observed ED50 and hill coefficients in the provided data files.</li> </ul> <p><strong>nucleoatac.tar.gz - Output files from running NucleoATAC on merged ATAC-seq from each of 5 TWIST1 dosages</strong></p> <ul> <li>TWIST1_{dosage}_merge.nucmap_combined.bed.gz - see NucleoATAC docs for output format</li> </ul>
Glacial Lake Image Dataset for "Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset"
<p>Glacial lake dataset for the paper <strong>"Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset" (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.jhydrol.2025.133072" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.jhydrol.2025.133072</a>)</strong></p> <p>The GLID dataset contains a total of 18,367 samples, and the size of each sample is 512*512. Each sample consists of an image and a corresponding label. Four glacial lake types including supraglacial lake, proglacial lake, ice-marginal lake, and unconnected glacial lake are involved. The pixel value of the glacier lake in annotation map is labeled 255 and background is labeled 0.</p> <p>GLID.rar contains the training dataset (16,000 samples), and val.zip contains the validation dataset (2,367 samples).</p> <p>GLID_annotation.zip contains the annotated shapefile of GLID with a CRS of WGS 84.</p> <p>Optical_images_source.xlsx contains the optical images ID/names and acquisition time of each platform (e.g., WV2, LC08, S2B, and GF02) used in GLID.</p> <p>Transferability validation.zip contains the images, labels, and predictions for transferability validation, which is independent of GLID (not used for model training or validation). The file structure is shown below:</p> <p>Transferability validation.zip</p> <ul> <li>images <ul> <li>AS.tif</li> <li>GL.tif</li> <li>NA.tif</li> <li>SA.tif</li> </ul> </li> <li>labels <ul> <li>AS_gt.tif</li> <li>GL_gt.tif</li> <li>NA_gt.tif</li> <li>SA_gt.tif</li> </ul> </li> <li>predictions <ul> <li>AS_pred.tif</li> <li>GL_pred.tif</li> <li>NA_pred.tif</li> <li>SA_pred.tif</li> </ul> </li> </ul> <p>AS, GL, NA, and SA represent Asia, Greenland, North America, and South America, respectively. Four high-quality Landsat-8/9 images (each cloud cover less than 6%) were used for testing, and we manually annotated the glacial lakes in each image as labels. The pixel value of the glacier lake in annotation map is labeled 255 and background is labeled 0. Files in transferability validation.zip have a same CRS of WGS 84.</p> <p>If you find this dataset is helpful in your research, please consider <strong><em>cite </em></strong>this paper:</p> <blockquote> <p><em>Ma D, Li J, Jiang L. 2025. Efficient glacial lake mapping by leveraging deep transfer learning and a new annotated glacial lake dataset. Journal of Hydrology 657: 133072.</em></p> </blockquote>
Replication Package for the Paper: Transfer Learning with Time Series Data: A Systematic Mapping Study
<p>This is a replication package for the paper "Transfer Learning with Time Series Data: A Systematic Mapping Study".</p> <p>It provides</p> <ul> <li>a documentation of the conducted electronic literature search,</li> <li>exports of the search results from each literature database,</li> <li>and an excel file on the included literature and extracted data.</li> </ul>
Data accompanying MetaChrom and "Annotating functional effects of non-coding variants in neuropsychiatric cell types by Deep Transfer Learning"
<p>This is the data accompanying the paper " Annotating functional effects of non-coding variants in neuropsychiatric cell types by Deep Transfer Learning" and the GitHub repository https://github.com/bl-2633/MetaChrom. </p> <p><strong>/data/bed_files/ </strong>contains the unprocessed bed file used in analysis</p> <p><strong>/data/seq_data/</strong> contains processed data from the bed files with corresponding partition and labels for each sequence segment.</p> <p><strong>/trained_model/MetaChrom_model/</strong> contains the pre-trained MetaChrom model on neural developmental context</p> <p><strong>/trained_models/MetaFeat_model/ </strong>contains the MetaFeat model used in training</p> <p><strong>/tool/</strong> contains files and software necessary for the processing pipeline.</p>
Classification of Eye Images by Personal Details With Transfer Learning Algorithms
<p>During the data collection phase of the research, first of all, a brief information was given to the participants about the study, and how the data would be used and what to do. Photographs of the eye area were collected from participants consisting of a total of 96 different people aged between 3-64. It has been clearly stated that there will be no situations that will define them during the photo shoot. Then, at least ten images of the right eye area of each person were taken. In addition, at least ten photographs of the left eye area were taken. Along with these photographs, no data other than the age and gender of the persons was recorded. Below are images of two people of different genders.</p> <p>A total of 1980 images were obtained from the participants, as in the figure above. More than ten images were obtained from some people. For this reason, there is a difference in the number of photos of people. Care has been taken to use different angles and lights so that each photograph does not form the same frame. Thus, photographs that were not, all the same, were collected. In order for each photograph not to be confused with another photograph, a naming rule has been developed to express the person, age, gender and the number of the photograph taken. An underscore ("_") character is inserted between each expression. Each expression used in the naming convention is given below in order.</p> <ul> <li>Person ID: It is a unique code value for each person photographed. This value ranges from 1 to 100.</li> <li>Age: The age is written directly as a number to express how old the person is. This value varies between 3-64.</li> <li>Gender ID: The value of 1 is expressed if the person photographed is male, and the value of 0 if it is a woman.</li> <li>Photo ID: Due to the fact that more than one photo was taken for each person, each photo was numbered sequentially from 1-10.</li> </ul> <p>If this dataset is used, reference should be made to the article below.</p> <ul> <li>Aktürk, C., Aydemir, E., Hama Rashid, Y. M. 2022. Classification of eye images according to person details with the transfer learning algorithms. Acta Informatica Pragensia, DOI: 10.18267/j.aip.190</li> </ul>
iTRADE: image-based TRAnsfer learning for Drug Effects
<p><code>iTRADE: <strong>i</strong>mage-based <strong>TRA</strong>nsfer learning for <strong>D</strong>rug <strong>E</strong>ffects</code></p> <p>This dataset supports <em>Berker et al. (2022) IEEE Trans Med Imaging</em>, <a href="https://doi.org/10.1109/TMI.2022.3205554">https://doi.org/10.1109/TMI.2022.3205554</a>.</p> <p>It comprises microscopy images, layout information and metabolic readouts for a total of 18 acquisitions: 1 control experiment (CE) and 17 drug screens (DS).</p> <pre><code>. ├── Images ├── Layouts ├── MeanOfStack ├── Metabolic └── README.md 4 directories, 1 file </code></pre> <p><strong><code>Images</code>: microscopy images</strong></p> <p>In total, the dataset contains 1 × 420 + 17 × 1848 = 31836 TIFF images (2123194264 bytes) in 18 folders.</p> <pre><code>Images ├── BT-40_V2_DS1 ├── BT-40_V3_DS1 ├── BT-40_V3_DS2 ├── HD-MB03_V1_DS1 ├── HD-MB03_V1_DS2 ├── HD-MB03_V2_DS1 ├── HD-MB03_V2_DS2 ├── INF_R_1021_relapse1_V1_DS2 ├── INF_R_1025_primary_V2_DS1 ├── INF_R_1123_primary_V1_DS1 ├── INF_R_153_CE ├── INF_R_153_V2_DS1 ├── INF_R_153_V3_DS1 ├── NCI-H3122_V2_DS1 ├── SJ-GBM2_V2_DS1 ├── SMS-KCNR_V1_DS1 ├── SMS-KCNR_V2_DS1 └── SMS-KCNR_V2_DS2 18 directories, 0 files </code></pre> <p><em>Control experiment</em></p> <p>The <code>INF_R_153_CE</code> folder contains images for a control experiment using the INF<em>R</em>153 cell line subjected to DMSO in 7 × 14 = 98 wells and staurosporine (STS) in 8 × 14 = 112 wells on a single 384-well plate (ignoring border wells). Each well is represented by 2 images (maximum intensity projection, <code>proj</code>, and mid-z image, <code>midz</code>), respectively. In total, this folder contains (7 + 8) × 14 × 2 = 420 images (26 megabytes).</p> <pre><code>Images/INF_R_153_CE ├── [ 60K] INF_R_153_CE_P1_B05_midz_224x224.tif ├── [ 58K] INF_R_153_CE_P1_B05_proj_224x224.tif ┆ ... ├── [ 62K] INF_R_153_CE_P1_O23_midz_224x224.tif └── [ 59K] INF_R_153_CE_P1_O23_proj_224x224.tif 0 directories, 420 files </code></pre> <p><em>Drug screens</em></p> <p>Other folders, such as <code>BT-40_V2_DS1</code>, are named for screen identifiers, which consist of a sample identifier (name of cell line, <code>INF_R_153|BT-40|HD-MB03|NCI-H3122|SJ-GBM2|SMS-KCNR</code>, or <em>INFORM</em> pseudonym of primary patient-derived sample, <code>INF_R_[0-9]{3,}_(primary|relapse[0-9])</code>) followed by <code>_V[0-9]_DS[0-9]</code> indicating the biological (<code>V</code>) and the technical (<code>DS</code>) replicate. For the three patient-derived samples included (<code>INF_R_[0-9]{4}</code>), <code>V1</code> signifies fresh viable tissue shipped immediately after biopsy while <code>V2</code> indicates a cell culture shipped after establishment.</p> <p>Each drug-screen folder contains images for a single drug screen consisting of 3 plates per screen, 14 × 22 = 308 wells per plate (namely, a 384-well plate ignoring all border wells), and 2 images per well. In total, each folder contains 3 × 14 × 22 × 2 = 1848 images (112 to 146 megabytes).</p> <pre><code>Images/BT-40_V2_DS1 ├── [ 62K] BT-40_V2_DS1_P1_B02_midz_224x224.tif ├── [ 58K] BT-40_V2_DS1_P1_B02_proj_224x224.tif ┆ ... ├── [ 60K] BT-40_V2_DS1_P3_O23_midz_224x224.tif └── [ 54K] BT-40_V2_DS1_P3_O23_proj_224x224.tif 0 directories, 1848 files Images/BT-40_V3_DS1 ├── [ 64K] BT-40_V3_DS1_P1_B02_midz_224x224.tif ├── [ 61K] BT-40_V3_DS1_P1_B02_proj_224x224.tif ┆ ... ├── [ 63K] BT-40_V3_DS1_P3_O23_midz_224x224.tif └── [ 62K] BT-40_V3_DS1_P3_O23_proj_224x224.tif 0 directories, 1848 files ... </code></pre> <p><em>Image files</em></p> <p>Image file names of the form <code>$ScreenID_P[123]_[B-O][0-9]{2}_(midz|proj)_[0-9]+x[0-9]+.tif</code> include the plate number (<code>1</code>, <code>2</code>, <code>3</code>), the well coordinates (<code>B02</code> to <code>O23</code>), the image type (<code>midz|proj</code>) and the size of the square images (<code>224x224</code>).</p> <p>Images are stored in Tagged Image File Format (TIFF), using a 16-bit integer in little-endian encoding for each pixel value. Files have been read from original TIFF image files and downscaled using the <code>Keras-Preprocessing</code> (v1.1.2) <code>load_img</code> function, and resaved (from plain image arrays without any metadata) using the <code>opencv-python</code> (v4.5.5.62) <code>imwrite</code> function using Adobe Deflate as a compression algorithm.</p> <p><strong><code>Layouts</code>: layout information</strong></p> <pre><code>Layouts ├── Drugs.csv ├── Layout_CE.csv └── Layout_DS.csv 0 directories, 3 files </code></pre> <p>Two text files, <code>Layout_CE.csv</code> and <code>Layout_DS.csv</code>, describe the layout of control experiments (<code>CE</code>) and drug screens (<code>DS</code>), respectively. Note that <code>Layout_DS.csv</code> has been generated from the imaging layout file published with the iTReX web app (available at <a href="https://itrex.kitz-heidelberg.de/">https://itrex.kitz-heidelberg.de/</a>), which can be downloaded from GitHub or iTReX. See <code>itrade.util.layouts.convert_itrex_ds_layout()</code> for details.</p> <p>A third text file, <code>Drugs.csv</code>, maps drug names as used in the iTReX-based <code>Layout_DS.csv</code> to drug names, abbreviations and drug (sub-)classes used throughout the manuscript. This file is used only by <code>plots.R</code>.</p> <p>All layout information is stored in long-table format. Text files are stored as Comma-Separated Values (CSV) with UTF-8 character encoding and Unix-style (<code>LF</code>) line endings.</p> <p><strong><code>MeanOfStack</code>: mean-of-stack measurements</strong></p> <p>For each drug screen represented by a folder named after the screen identifier, mean-of-stack computations produced from the full-resolution (<code>2048x2048</code> pixels) are stored in matrix format using one text file per plate, named <code>SID_P[123]_$Barcode.txt</code>, e.g., <code>BT-40_V2_DS1_P1_H104-03N1A98.txt</code>. Text files are stored as Tab-Separated Values (TSV) with UTF-8 character encoding and Unix-style (<code>LF</code>) line endings. This folder comprises a total of 17 × 3 = 51 text files.</p> <pre><code>MeanOfStack ├── BT-40_V2_DS1 │ ├── BT-40_V2_DS1_P1_H104-03N1A98.txt │ ├── BT-40_V2_DS1_P2_H104-03N2A98.txt │ └── BT-40_V2_DS1_P3_H104-03N3A98.txt ├── BT-40_V3_DS1 │ ├── BT-40_V3_DS1_P1_H104-03N1D07.txt │ ├── BT-40_V3_DS1_P2_H104-03N2D07.txt │ └── BT-40_V3_DS1_P3_H104-03N3D07.txt ┆ ... └── SMS-KCNR_V2_DS2 ├── SMS-KCNR_V2_DS2_P1_H104-03N1D02.txt ├── SMS-KCNR_V2_DS2_P2_H104-03N2D02.txt └── SMS-KCNR_V2_DS2_P3_H104-03N3D02.txt 17 directories, 51 files </code></pre> <p><strong><code>Metabolic</code>: metabolic readouts</strong></p> <p>Similar to mean-of-stack computations, metabolic readouts are included for each drug screen. Folder and file names and file formats are identical to the <code>MeanOfStack</code> folder.</p> <pre><code>Metabolic ├── BT-40_V2_DS1 │ ├── BT-40_V2_DS1_P1_H104-03N1A98.txt │ ├── BT-40_V2_DS1_P2_H104-03N2A98.txt │ └── BT-40_V2_DS1_P3_H104-03N3A98.txt ├── BT-40_V3_DS1 │ ├── BT-40_V3_DS1_P1_H104-03N1D07.txt │ ├── BT-40_V3_DS1_P2_H104-03N2D07.txt │ └── BT-40_V3_DS1_P3_H104-03N3D07.txt ┆ ... └── SMS-KCNR_V2_DS2 ├── SMS-KCNR_V2_DS2_P1_H104-03N1D02.txt ├── SMS-KCNR_V2_DS2_P2_H104-03N2D02.txt └── SMS-KCNR_V2_DS2_P3_H104-03N3D02.txt 17 directories, 51 files </code></pre>
Data for paper: Transfer learning for cross-context prediction of protein expression from 5'UTR sequence
<p>This depsit contains data for the paper entitled: "<strong>Transfer learning for cross-context prediction of protein expression from 5'UTR sequence</strong>".</p> <p>The <strong>rebeca.zip</strong> file contains a snapshot of the rebeca package which can be used to train, fine tune and test the CONV-LSTM model used in this study.</p> <p>The <strong>datasets.zip</strong> file contains the compiled sequence to expression datasets from across all Flow-seq expressions considered in this study. </p> <p>The <strong>analysis.zip</strong> file contains all data files and jupyter notebooks necessary to reproduce our analysis. Each Flow-seq study has a dedicated folder (e.g., `fepB') with two sub-folders: 1. The `data\_split' folder, which contains the steps necessary to split the Flow-seq data for our ML experiments (a `readme.txt' file describes the input and output files and a jupyter notebook is available to reproduce the data split); 2. The `data\_analysis' folder, which contains a jupyter notebook and the necessary input files to reproduce the analysis of our experiments.</p>
Cell type matching across species using protein embeddings and transfer learning
<p>TACTiCS is a method to transfer and align cell types in cross-species data. This repository contains the protein sequences, protein embeddings, count matrices and trained models for human, mouse and marmoset.</p>
DFT torsiondrive data for: MACE-OFF23: Transferable Machine Learning Force Fields for Organic Molecules
<p>MACE-OFF23: Transferable Machine Learning Force Fields for Organic Molecules</p> <div><a href="https://arxiv.org/search/physics?searchtype=author&query=Kov%C3%A1cs,+D+P">Dávid Péter Kovács</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Moore,+J+H">J. Harry Moore</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Browning,+N+J">Nicholas J. Browning</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Batatia,+I">Ilyes Batatia</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Horton,+J+T">Joshua T. Horton</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Kapil,+V">Venkat Kapil</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Witt,+W+C">William C. Witt</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Magd%C4%83u,+I">Ioan-Bogdan Magdău</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Cole,+D+J">Daniel J. Cole</a>, <a href="https://arxiv.org/search/physics?searchtype=author&query=Cs%C3%A1nyi,+G">Gábor Csányi </a><a href="https://doi.org/10.48550/arXiv.2312.15211">https://doi.org/10.48550/arXiv.2312.15211</a></div> <p> </p> <p>Supporting data including raw outputs from SPICE consistent torsion drives on the TorsionNet500 and OpenFF Biaryl datasets and HDF5 versions formated to be consistent with the rest of the SPICE dataset. See the <a href="../records/10975225">SPICE release</a> for more details. </p>
Data Set: Evaluation of Domain Randomization Techniques for Transfer Learning - Part 0
<p><strong>Transfer Learning Data Set:</strong><br> The total data set contains <strong>1.44m synthetic images</strong> and <strong>10k real-world images </strong>of size 299x224.</p> <p>Due to file size limitation, the data set is split into two sets. <strong>This data set</strong> (part 0: DOI 10.5281/zenodo.2581311) contains the real-world images and the synthetic images of perspective 00. The second part (part 1: DOI 10.5281/zenodo.2581469) contains perspective 01 and 10 of the synthetic images.</p> <p><strong>1. 10k real world images</strong><br> the filename labels the image:<br> -first 6 digits represent the id.<br> -8. digit labels the grasp: 0 -> no grasp, 1 -> grasp<br> -10. digit is empty (reserved for 2nd perspective)<br> -12. digit labels the graspbox: 0 -> green box, 1 -> yellow box<br> -14. digit labels the distractors : 0 -> no distractors, 1 -> distractors<br> -c in the end stands for acolor image, d is reserved for depth (not in use).<br> <br> <strong>2. 1.44m synthetic images</strong><br> the folder labels the enabled technique:<br> -first 2 digits label the perspective: 00 -> standard, 01 -> shake, 10 -> random<br> -3. digit labels the graspbox: 0 -> defaul green box, 1 -> random box<br> -4. digit labels the distractors: 0 -> no distractors, 1 -> distractors<br> -5. digit labels the lighting: 0 -> default lighting, 1 -> random lighting<br> -6. digit labels the mesh randomization: 0 -> default mesh color, 1 -> random mesh color<br> <br> the filename consists of 3 parts:<br> -first 6 digits represent the id.<br> -8. digit labels the grasp: 0 -> no grasp, 1 -> grasp<br> -10.-15. digit equals the folder name and represents the enabled technique<br> -c in the end stands for a color image, d is reserved for depth (not in use).</p>
Data Set: Evaluation of Domain Randomization Techniques for Transfer Learning - Part 1
<p><strong>Transfer Learning Data Set:</strong><br> The total data set contains <strong>1.44m synthetic images</strong> and <strong>10k real-world images</strong> of size 299x224.</p> <p>Due to file size limitation, the data set is split into two sets. The first part (part 0: DOI 10.5281/zenodo.2581311) contains the real-world images and the synthetic images of perspective 00. <strong>This data set </strong>(part 1: DOI 10.5281/zenodo.2581469) contains perspective 01 and 10 of the synthetic images.</p> <p><strong>1. 10k real world images</strong><br> the filename labels the image:<br> -first 6 digits represent the id.<br> -8. digit labels the grasp: 0 -> no grasp, 1 -> grasp<br> -10. digit is empty (reserved for 2nd perspective)<br> -12. digit labels the graspbox: 0 -> green box, 1 -> yellow box<br> -14. digit labels the distractors : 0 -> no distractors, 1 -> distractors<br> -c in the end stands for acolor image, d is reserved for depth (not in use).<br> <br> <strong>2. 1.44m synthetic images</strong><br> the folder labels the enabled technique:<br> -first 2 digits label the perspective: 00 -> standard, 01 -> shake, 10 -> random<br> -3. digit labels the graspbox: 0 -> defaul green box, 1 -> random box<br> -4. digit labels the distractors: 0 -> no distractors, 1 -> distractors<br> -5. digit labels the lighting: 0 -> default lighting, 1 -> random lighting<br> -6. digit labels the mesh randomization: 0 -> default mesh color, 1 -> random mesh color<br> <br> the filename consists of 3 parts:<br> -first 6 digits represent the id.<br> -8. digit labels the grasp: 0 -> no grasp, 1 -> grasp<br> -10.-15. digit equals the folder name and represents the enabled technique<br> -c in the end stands for a color image, d is reserved for depth (not in use).</p>
Training, Validation and Test Sets for paper 'A Little Data goes a Long Way: Automating Seismic Phase Arrival Picking at Nabro Volcano with Transfer Learning'
<p>Training, Validation and Test Data for model presented in paper 'A Little Data Goes A Long Way: Automating Seismic Phase Arrival Picking at Nabro Volcano with Transfer Learning', submitted to Journal of Geophysical Research: Solid Earth.</p> <p>Files:</p> <p>- train_events_2498.h5 = training set of seismic waveforms (events with P-/S-wave labelled arrivals only, i.e., no noise waveforms)</p> <p>- train_events_2498.pkl = event training set metadata (UTC P-/S-wave phase arrival times)</p> <p>- train_noise_2498.h5 = training set of seismic waveforms (noise sections only, i.e., no event waveforms)</p> <p>- train_noise_2498.pkl = noise training set metadata (UTC time for training noise waveforms)</p> <p>- val_events.h5 = validation set of seismic waveforms (events with P-/S-wave labelled arrivals only, i.e., no noise waveforms)</p> <p>- val_events.pkl = event validation set metadata (UTC P-/S-wave phase arrival times)</p> <p>- val_noise.h5 = validation set of seismic waveforms (noise sections only, i.e., no event waveforms)</p> <p>- val_noise.pkl = noise validation set metadata (UTC time for validation noise waveforms)</p> <p>- test.h5 = test set of seismic waveforms (events and noise)</p> <p>- test_events.pkl = event test set metadata (UTC P-/S-wave phase arrival times for test event waveforms)</p> <p>- test_noise.pkl = noise test set metadata (UTC time for test noise waveforms)</p> <p>- nabro_2011-247.mseed = 24 hours seismic data from Nabro Urgency Array (2011-09-04), saved in mseed format (e.g., can be read with obspy)</p> <p>- nabro_2011-269.mseed = 24 hours seismic data from Nabro Urgency Array (2011-09-26), saved in mseed format (e.g., can be read with obspy)</p> <p> </p> <p>Further details and code for reading and using these files can be found at the GitHub repo for this paper: <a href="https://github.com/sachalapins/U-GPD">https://github.com/sachalapins/U-GPD</a></p> <p> </p>
Dataset for "Generalizing property prediction of ionic liquids from limited labeled data: a one-stop framework empowered by transfer learning"
<p>Dataset for "Generalizing property prediction of ionic liquids from limited labeled data: a one-stop framework empowered by transfer learning", codes can be found <a href="https://github.com/GuzhongChen/ILTransR">here</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.