Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,805

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,805 results for “Data model”

Learn how ShareScore rates datasets ↗
zenodo36/100

Gridded forecast data over the Mediterranean and the North Sea from ECMWF models

<p>Gridded forecast data over the Mediterranean and the North Sea from ECMWF used in Cavaleri et al JGR 2023</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Modeled waver data for ML

<p>The uploaded data are numerical model generated&nbsp;daily wave height, period,&nbsp;and wind frocings in&nbsp;the Chesapeake Bay.&nbsp;Data are used for the publication &quot;Machine Learning-based Wave Model with High Spatial Resolution in Chesapeake Bay&quot; submitted to the Journal for review.</p>

opencc-byJul 2023View details →
zenodo36/100

ThoughtSource: A central hub for large language model reasoning data (dataset snapshot)

<p><strong>ThoughtSource is a meta-dataset and software library for chain-of-thought reasoning in large language models (LLMs). </strong></p> <p><strong>This repository contains a snapshot of the openly available ThoughtSource datasets.</strong></p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Data on the Model of Loneliness and Smartphone Use Intensity as a Mediator of Self-Control, Emotion Regulation, and Spiritual Meaningfulness in Nomophobia

<p><em>The data collection was conducted in June-July 2022, with 355 participants completing the scales provided through offline paper-pencil surveys in the classroom. The research data was collected from three cities, namely Palembang, Jambi, and Yogyakarta, and consisted of junior high school and high school students, as well as university students. The university student participants were recruited from Ahmad Dahlan University, Jambi University, and Charitas Musi Catholic University.&nbsp;&nbsp;The participants completed the Nomophobia NMP-Q scale (Yildirim &amp; Correia, 2015), the R-UCLA Loneliness Scale (Russell et al., 1980), self-control (Tangney, Baumeister &amp; Boone, 2004), The Difficulties in Emotion Regulation Scale (DERS; Gratz &amp; Roemer, 2004), and Spiritual Meaningfulness which was developed based on the theory of Pargament (2007). All the questionnaires for data collection were in the Indonesian version.</em></p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Data from: Investigating the human and non-obese diabetic mouse MHC class II immunopeptidome using protein language modelling.

<p><strong>Background</strong>: Identifying peptides associated with the major histocompability complex class II (MHCII) is a central task in the evaluation of the immunoregulatory function of therapeutics and drug prototypes. MHCII-peptide presentation prediction has multiple biopharmaceutical applications, including the safety assessment of biologics and engineered derivatives&nbsp;in silico, or the fast progression of antigen-specific immunomodulatory drug discovery programs in immune disease and cancer. This has resulted in the collection of large&ndash;scale data sets on adaptive immune receptor antigenic responses and MHC-associated peptide proteomics. In parallel, recent deep learning algorithmic advances in natural language processing (NLP) and protein language modelling (PLM) have shown potential in leveraging large collections of sequence data and improve MHC presentation prediction. <strong>Methodology</strong>: We trained a compact transformer model (AEGIS) on human and mouse MHCII immunopeptidome data, including a preclinical murine model, and evaluated its performance on the peptide presentation prediction task. <strong>Data</strong>:&nbsp;The data and models used in&nbsp;AEGIS are contained in the uploaded tar files. <strong>Results</strong>:&nbsp;The transformer performs on par with existing deep learning algorithms and that combining datasets from multiple organisms increases model performance (see preprint). We trained variants of the model with and without MHCII information. In both alternatives, the inclusion of peptides presented by the I-Ag7&nbsp;MHC class II molecule expressed by the non-obese diabetic (NOD) mice enabled the&nbsp;in silico&nbsp;prediction of presented peptides in a preclinical type 1 diabetes model organism, which has promising therapeutic applications.</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Data for: Leveraging spatio-temporal genomic breeding value estimates of dry matter yield and herbage quality in ryegrass via random regression models

<p>Joint modeling of correlated multi-environment and multi-harvest data of perennial crop species may offer advantages in prediction schemes and a better understanding of the underlying dynamics in space and time. The goal of the present study was to investigate the relevance of incorporating the longitudinal dimension of within-season multiple measurements of forage perennial ryegrass traits in a reaction norm model setup that additionally accounts for genotype-environment interactions (G×E). Genetic parameters and accuracy of genomic breeding value (gEBV) predictions were investigated by fitting three random regression models (gRRM) using Legendre polynomial functions to the data. Genomic DNA sequencing of family pools of diploid perennial ryegrass was performed using DNA nanoball-based technology and yielded 56,645 single nucleotide polymorphisms which were used to calculate the allele frequency-based genomic relationship matrix. Biomass yield's estimated additive genetic variance and heritability values were higher in later harvests. The additive genetic correlations were moderate to low in early measurements and peaked at intermediates, with fairly stable values across the environmental gradient, except for the initial harvest data collection. This led to the conclusion that complex (G×E) arises from spatial and temporal dimensions in the early season, with lower re-ranking trends thereafter. In general, modeling the temporal dimension with a second-order orthogonal polynomial improved the accuracy of gEBV prediction for nutritive quality traits, but no gain in prediction accuracy was detected for dry matter yield. This study leverages the flexibility and usefulness of gRRM models for perennial ryegrass breeding and can be readily extended to other multi-harvest crops.</p>

opencc-zeroAug 2023View details →
zenodo36/100

Raw data for model implementation examples in paper "Incorporating Detected/Undetected Cooperative and Uncooperative Individuals, and Dynamic Transmission Probabilities in Epidemiological Models"

<p>This dataset contains the historical daily new case numbers referenced in the paper titled &quot;Incorporating Detected/Undetected Cooperative and Uncooperative Individuals, and Dynamic Transmission Probabilities in Epidemiological Models.&quot; These numbers were utilized to estimate the basic reproduction number <em>R<sub>0</sub></em> and the aggregated epidemic control measure <em>K<sub>r</sub>(t)</em> parameters within the model, as shown in the example implementation section of the paper.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Mapping data files to semantic data models using the CaosDB crawler

<p>Data from data acquisition can lead to a high variety of data files on file systems. The figure illustrates that these files can be mapped to semantic data models in the research data management system CaosDB using a customizable crawler.</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Data used in the physical-biogeochemical model of Danjiangkou Reservoir

<p>The&nbsp;date set includes&nbsp;meteorological, hydrological, water quality, and organic carbon loading data obtained in 2009&nbsp;for the Danjiangkou Reservoir in China. The data were used to set the&nbsp;boundary conditions of the&nbsp;physical-biogeochemical model, which was adopted to simulate reservoir methane dynamics in the&nbsp;Danjiangkou Reservoir.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

The second data release from the European Pulsar Timing Array II. Customised pulsar noise models for spatially correlated gravitational waves

<p>Aims: The nanohertz gravitational wave background (GWB) is expected to be an aggregate signal of an ensemble of gravitational waves emitted predominantly by a large population of coalescing supermassive black hole binaries in the centres of merging galaxies. Pulsar tiNanohertz&nbsp;ming arrays (PTAs), which are ensembles of extremely stable pulsars at approximately kiloparsec distances precisely monitored for decades, are the most precise experiments capable of detecting this background. However, the subtle imprints that the GWB induces on pulsar timing data are obscured by many sources of noise that occur on various timescales. These must be carefully modelled and mitigated to increase the sensitivity to the background signal. Methods: In this paper, we present a novel technique to estimate the optimal number of frequency coefficients for modelling achromatic and chromatic noise, while selecting the preferred set of noise models to use for each pulsar. We also incorporated a new model to fit for scattering variations in the Bayesian pulsar timing package temponest. These customised noise models enable a more robust characterisation of single-pulsar noise. We developed a software package based on tempo2 to create realistic simulations of European Pulsar Timing Array (EPTA) datasets that allowed us to test the efficacy of our noise modelling algorithms. Results: Using these techniques, we present an in-depth analysis of the noise properties of 25 millisecond pulsars (MSPs) that form the second data release (DR2) of the EPTA and investigate the effect of incorporating low-frequency data from the Indian Pulsar Timing Array collaboration for a common sample of ten MSPs. We used two packages, enterprise and temponest, to estimate our noise models and compare them with those reported using EPTA DR1. We find that, while in some pulsars we can successfully disentangle chromatic from achromatic noise owing to the wider frequency coverage in DR2, in others the noise models evolve in a much more complicated way. We also find evidence of long-term scattering variations in PSR J1600-3053. Through our simulations, we identify intrinsic biases in our current noise analysis techniques and discuss their effect on GWB searches. The analysis and results discussed in this article directly help to improve the sensitivity to the GWB signal and they are already being used as part of global PTA efforts.</p>

opencc-zeroJun 2023View details →
zenodo36/100

CSPG GridPath Model Input and Output Data

<p>This data repository holds input and output data for the paper Jin XY., Chowdhury, A.K., Cheng CT., and Galelli, S. &ldquo;The unintended consequences of decarbonizing the China Southern Power Grid&rdquo;.&nbsp;See Readme for more details.&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Data set used in article: On the Potential of Reduced Order Models for Wind Farm Control: A Koopman Dynamic Mode Decomposition Approach

<p>Step-wise pitch simulation of two wind turbines interacting using SOWFA. More information in the paper.</p>

opencc-by-4.0Oct 2020View details →
dryad36/100

Data from: Emergent spatial patterns can indicate upcoming regime shifts in a realistic model of coral community

<p class="western"><span>Increased stress on coastal ecosystems, such as coral reefs, seagrasses, kelp forests and other habitats can make them shift towards degraded, often algae-dominated or barren communities. This has already occurred in many places around the world, calling for new approaches to identify where such regime shifts may be triggered. Theoretical work predicts that the spatial structure of habitat-forming species should exhibit changes prior to regime shifts,</span><span><em> </em></span><span>such as an increase in spatial autocorrelation. However, extending this theory to marine systems requires theoretical models connecting field-supported ecological mechanisms to data and spatial patterns at relevant scales. To do so, we built a spatially-explicit model of sub-tropical coral communities based on experiments and long-term datasets from Rapa Nui (Easter Island, Chile), to test whether spatial indicators could signal upcoming regime shifts in coral communities. Spatial indicators anticipated degradation of coral communities following increases in frequency of bleaching events or coral mortality. However, they were generally unable to signal shifts that followed herbivore loss, a widespread and well-researched source of degradation, likely because herbivory, despite being critical for the maintenance of corals, had comparatively little effect on their self-organization. Informative trends were found both under equilibrium and non-equilibrium conditions, but were determined by the type of direct neighbor interactions between corals, which remain relatively poorly documented. These inconsistencies show that while this approach is promising, its application to marine systems will require detailed information about the type of stressor, and filling current gaps in our knowledge of interactions at play in coral communities. </span></p>

opencc-zeroAug 2023View details →
dryad36/100

Data from: 3D shear-wave velocity model of central Makran using ambient-noise adjoint tomography

<p>The Makran subduction zone is unique in its wide onshore thick accretionary prism, and a volcanic arc not parallel to the E-W trend of the Makran accretionary prism. To investigate the internal structure of the accretionary prism, the crustal nature of Jaz Murian Depression, and the trend of the buried trench we have calculated a 3D shear-wave velocity model for a region around the border between eastern and western Makran using ambient-noise adjoint tomography and data from IASBS/CAM Makran temporary seismic network. In close agreement with previous works, our velocity model shows that the onshore accretionary prism consists of a low-velocity zone in the south and a high-velocity zone in the north with an average thickness of accreted sediments of 22 and 30 km, respectively. The young age of the surface rocks of the high-velocity part of the prism suggests the presence of a significant volume of igneous rocks scraped from the subducting oceanic slab. The velocity model indicates a continental crust of ~40 km with a thick sedimentary cover of ~20 km for the eastern part of Jaz Murian Depression. The presence of a NE-SW trending low-velocity region at a depth interval of 40-60 km subparallel with the trend of the volcanic arc, intermediate-depth earthquakes, and geometry of the overriding plate might be related to the trend of the buried trench. This implies that the observed NE-SW trending volcanic arc might be related to the geometry of the buried trench and not the eastward reduction of the subduction angle.</p>

opencc-zeroAug 2023View details →
zenodo36/100

Data for analyses by graphical loglinear Rasch models of the PSFP and the PLCFP subscales of the PSSFP

<p>Data for&nbsp;analyses by graphical loglinear Rasch models of the PSFP and the PLCFP subscales of&nbsp;the PSSFP. Data are from Danish student teachers. Contains the following variables:</p> <p>i1 through i10 are the items from the PSSFP. &nbsp;Response scale is 0 = Never, 1 = Almost Never, 2 = Sometimes, 3 = Fairly Often, 4 = Very Often. &nbsp;Items 4, 5, 7 and 8 are reversed.</p> <p>P_level (level of latest field practice placement): 1 = level I, 2 = level II, 3 = level III</p> <p>T_Progr (teacher education program): 1 = regular, 2 = other</p> <p>Campus: 1 = campus A, 2 = campus B</p> <p>Gender: 1 = female, 2 = male</p> <p>Age: 1 = 25 years and younger, 2 = 26 years and older</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

ChromBPNet models and data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency

<p>This record contains ChromBPNet models and data used to train the models for the paper &quot;Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency&quot; by Nair, Ameen <em>et al</em>.</p> <p>`data` contains bigwigs and regions (peaks + non-peaks) used for training each of the models. See `data/README.txt` for more details.</p> <p><strong>Models:</strong></p> <p><em>Loading the&nbsp;model:</em></p> <p>The models were trained using tf1.14. The models are provided in h5 format for tf1.14 (py3.7) and SavedModel format for tf2.X. tf2.X tested only for py3.8-11, tf2.8-13.</p> <p>To load the models in tf1.14:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model.h5")</code></pre> <p>In tf2:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model_dir")</code></pre> <p>If all fails, you can load the architecture as provided in `model_arch.py` with default parameters (`bpnet_seq` for bias model and `chrombpnet` for chrombpnet model), and then load the weights using `model.load_weights` from the weights provided in the `weights` directory.</p> <p>&nbsp;</p> <p><em>Usage:</em></p> <p>The bias models take as input one-hot sequence of length 2000. It has 2 outputs, a vector of logits of length 2000, and 1 logcounts scalar:</p> <pre><code class="language-python"># seq_one_hot of length B x 2000 x 4 out_bias_logits, out_bias_logcounts = bias_model.predict(seq_one_hot) # out_bias_logits: B x 2000 # out_bias_logcounts: B x 1</code></pre> <p>The ChromBPNet model takes as input a one-hot sequence of length 2000, bias logits of length 2000 and bias log-counts scalar. It has the same output types as the bias model. To run the chrombpnet model to obtain predictions:</p> <pre><code class="language-python">pred_profile, pred_logcounts = chrombpnet_model.predict([seq_one_hot, out_bias_logits, out_bias_logcounts]) # pred_profile: B x 2000 # pred_logcounts: B x 1 </code></pre> <p>If you wish to obtain the &quot;de-biased&quot; predictions (see Methods), simply pass in zeros instead of the bias model predictions as:</p> <pre><code class="language-python">pred_profile_debiased, pred_logcounts_debiased = chrombpnet_model.predict([seq_one_hot, np.zeros((seq_one_hot.shape[0], 2000)), np.zeros((seq_one_hot.shape[0], 1))])</code></pre> <p>To obtain predicted per-base predicted counts (with or without bias):</p> <pre><code class="language-python">pred_per_base_counts = scipy.special.softmax(pred_profile, axis=-1) * (np.exp(pred_logcounts)-1) # pred_per_base_counts: B x 2000 </code></pre> <p>Note that in general predicted counts can&#39;t be compared across models as they are not corrected for sequencing depth.</p> <p>&nbsp;</p> <p><em>Note:</em></p> <p>All bias models used across folds are identical, except for the final intercept term in the counts output (see Methods), that is specific to each cell state, fold combination.</p> <p>&nbsp;</p> <p><em>Folds:</em></p> <p>The splits used for training the different folds are as below:</p> Fold Test Chromosomes Validation Chromosomes 0 chr1 chr8, chr10 1 chr2, chr19 chr1 2 chr3, chr20 chr2, chr19 3 chr6, chr13, chr22 chr3, chr20 4 chr5, chr16, chrY chr6, chr13, chr22 5 chr4, chr15, chr21 chr5, chr16, chrY 6 chr7, chr18, chr14 chr4, chr15, chr21 7 chr11, chr17, chrX chr7, chr18, chr14 8 chr9, chr12 chr11, chr17, chrX 9 chr8, chr10 chr9, chr12 <p>Remaining chromosomes were used as the training chromosome for each fold.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Data for A Physical Model for the Observed Inverse Energy Cascade in Typhoon Boundary Layers

<p>This repository contains dataset for the paper entitled &quot;A Physical Model for the Observed Inverse Energy Cascade in Typhoon Boundary Layers&quot;. The magnitude of inverse energy cascade flux is revised in version 2.0 according to&nbsp;Xia et al. (2009) (https://doi.org/10.1063/1.3275861).&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Data used for modeling in Energy-water-land-CCUS nexus model: carbon dioxide opportunities based on optimized regional development

<p>In this dataset, the data used for modeling technologies in an energy-water-land-CCUS nexus model in Khark Island in Iran, and the main sources for gathering them are presented.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Data accompanying the publication: Measurements and modelling of pore-pressure gradients in the swash zone under large-scale laboratory bichromatic waves

<p>Data accompanying the publication: Measurements and modelling of pore-pressure gradients in the swash zone under large-scale laboratory bichromatic waves. See the README.txt file for more detail.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Data for "Reducing Southern Ocean biases in the FOCI climate model"

<p>Jupyter notebooks and time-averaged data needed to reproduce all plots in &quot;Reducing Southern Ocean biases in the FOCI climate model&quot; submitted to JAMES.</p> <p>Source code modifications needed to compile and run the model is also included.</p> <p>See attached README for more information.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record