Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

36

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

36 results for “generalizability”

Learn how ShareScore rates datasets ↗
zenodo48/100

Toward a Generalizable Machine-Learned Potential for Metal-Organic Frameworks

<ul> <li>This repository contains the dataset used in the publication<br>&nbsp; `Toward Generalizable Machine Learned Potential for Metal-Organic Frameworks` Yue Yifei, Saad Aldin Mohammed, Loh Duane*, Jiang Jianwen*<br>&nbsp;&nbsp;<br>&nbsp; Please each the README.md within each subfolder. For brevity, the data is organized into three sections<br>&nbsp;&nbsp;<br>&nbsp; 1. The dataset in DATASET<br>&nbsp; &nbsp; &nbsp; - The training and testing dataset, including structures of MOFs in extxyz format<br>&nbsp; &nbsp; &nbsp;&nbsp;<br>&nbsp; 2. The training output files and logs in NEQUIP-TRAIN<br>&nbsp; &nbsp; &nbsp; - The conda environment details, training scripts and logs<br>&nbsp; &nbsp; &nbsp; - Also Training and testing metrics in csv files<br>&nbsp;&nbsp;&nbsp;&nbsp; - This is split into two zip files NEQUIP-TRAIN1 and NEQUIP-TRAIN2 due to their size<br>&nbsp; &nbsp; &nbsp;<br>&nbsp; 3. Examples of using the developed models in MD simulations<br>&nbsp; &nbsp; &nbsp; - Including LAMMPS scripts, data file and environment details used in our scalability tests<br>&nbsp; &nbsp; &nbsp; - The complied Nequip-patched LAMMPS version is also provided<br>&nbsp; &nbsp; &nbsp; - Details on how to use our models - we used a default model that is slower but more accurate in our study but faster models are also developed.</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Dataset - DeepWealth: A Generalizable Open-Source Deep Learning Framework using Satellite Images for Well-Being Estimation

<p>This dataset encapsulates the Checkpoints obtained during the training process of the Deep Learning model, which can be used for new estimations.</p> <p>The aim of the DeepWealth package is to provide a generalizable Deep Learning framework for the use of remote sensing in poverty estimation. The combination of Deep Learning and Earth Observation data is increasingly being used to estimate socioeconomic conditions at regional and global scales. The proposed framework aligns with the Sustainable Development Goal SDG1 of ending poverty. The framework provides open-source data, code, and training models (checkpoints) for reproducibility and replicability.</p> <ul> <li>The source code can be found in&nbsp;<a href="https://github.com/PARSECworld/DeepWealth" target="_blank" rel="noopener">https://github.com/PARSECworld/DeepWealth</a></li> <li>The metadata from source code can be found in&nbsp;<a href="https://github.com/PARSECworld/DeepWealth/blob/main/metadata.pdf" target="_blank" rel="noopener">https://github.com/PARSECworld/DeepWealth/blob/main/metadata.pdf</a></li> <li>The paper describing the development of this framework can be found at: Ben Abbes, A., Machicao, J., Corr&ecirc;a, P. L. P., Specht, A., Devillers, R., Ometto, J. P., Kondo, Y., &amp; Mouillot, D. (2024). DeepWealth: A generalizable open-source deep learning framework using satellite images for well-being estimation.&nbsp;<em>SoftwareX</em>, 27, 101785.&nbsp; <a href="https://doi.org/10.1016/j.softx.2024.101785">https://doi.org/10.1016/j.softx.2024.101785</a>&nbsp;</li> </ul>

openmit-licenseJan 2024View details →
zenodo40/100

What is learned in approach-avoidance tasks? On the scope and generalizability of approach-avoidance effects

<p>Previous research has shown that approaching a stimulus makes it more positive, while avoiding a stimulus makes it more negative. The present research demonstrates that approach-avoidance behaviors have the potential to charge stimulus attributes such as color with evaluative meaning. This evaluation carries over to other stimuli with that feature. We address the latter point by assessing the influence of colors that were approached or avoided on the perceived attractiveness of persons wearing those colors. We show that wearing a certain color makes people appear more attractive when this color is associated with approach rather than avoidance. In line with a self-perception account of these effects, we obtained approach-avoidance effects on stimulus attributes only when participants carried out approach-avoidance behaviors towards these colors or imagined doing so. This set of experiments adds to the evaluative learning literature by demonstrating approach-avoidance effects on stimulus attributes and that these effects carry over to new classes of stimuli and new tasks. Moreover, we systematically investigated boundary conditions for these effects. Finally, with this research we introduce an ontogenetic perspective to research into colors and their influence on psychological functioning.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Data for "Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms"

<p>This repository contains the data&nbsp;and external data used by teams in the Kaggle competition &quot;HuBMAP+HPA - Hacking the Human Body&quot; and is part of the paper &quot;Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms&quot;.</p> <p>The directories contain:</p> <p><strong>data.zip:</strong> The training and test data, including metadata, used in the Kaggle competition &quot;HuBMAP + HPA - Hacking the Human Body&quot;.</p> <p><strong>Team_1.zip: </strong>External data used by the first place winning solution.</p> <p><strong>Team_2.zip: </strong>External data used by the second place winning solution.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Trained Models for "Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms"

<p>This repository contains the trained model weights&nbsp;for the baseline model and the winning solutions in the Kaggle competition &quot;HuBMAP+HPA - Hacking the Human Body&quot;, and is part of the paper &quot;Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms&quot;.</p> <p>The directory&nbsp;contains:</p> <p><strong>trained_model_1_weights.zip: </strong>Trained model weights for first place solution (Team 1).</p> <p><strong>trained_model_2_weights.zip:</strong>&nbsp;Trained model weights for second place solution (Team 2).</p> <p><strong>trained_model_3_weights.zip:&nbsp;</strong>Trained model weights for third place solution (Team 3).</p> <p><strong>trained_model_weights_baseline.zip:</strong>&nbsp;Trained model weights for the baseline model.</p>

opencc-by-4.0Jan 2023View details →
dryad40/100

On the cross-population generalizability of gene expression prediction models

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad36/100

Clinical trial generalizability assessment in the big data era: a review

<p><span><span><span>Clinical studies, especially randomized controlled trials, are essential for generating evidence for clinical practice.  However, generalizability is a long-standing concern when applying trial results to real-world patients.  Generalizability assessment is thus important, nevertheless, not consistently practiced.  We performed a systematic scoping review to understand the practice of generalizability assessment.  We identified 187 relevant papers and systematically organized these studies in a taxonomy with three dimensions: (1) data availability (i.e., before or after trial [<i>a priori</i> vs <i>a posteriori</i> generalizability]), (2) result outputs (i.e., score vs non-score), and (3) populations of interest.  We further reported disease areas, underrepresented subgroups, and types of data used to profile target populations.  We observed an increasing trend of generalizability assessments, but less than 30% of studies reported positive generalizability results.  As <i>a priori</i> generalizability can be assessed using only study design information (primarily eligibility criteria), it gives investigators a golden opportunity to adjust the study design before the trial starts.  Nevertheless, less than 40% of the studies in our review assessed <i>a priori</i> generalizability.  With the wide adoption of electronic health records systems, rich real-world patient databases are increasingly available for generalizability assessment; however, informatics tools are lacking to support the adoption of generalizability assessment practice.</span></span></span></p>

opencc-zeroApr 2020View details →
zenodo36/100

The Sampling Threat when Mining Generalizable Inter-Library Usage Patterns

<div> <div> <div> <p>Tool support in software engineering often relies on relationships, regularities, patterns, or rules mined from other users&rsquo; code. Examples include approaches to bug prediction, code recommendation, and code autocompletion. Mining is typically performed on samples of code rather than the entirety of available software projects. While sampling is crucial for scaling data analysis, it might influence the generalization of the mined patterns. This paper focuses on sampling software projects filtered for specific libraries and frameworks, and on mining patterns that connect different libraries. We call these inter-library patterns.</p> <p>We observe that limiting the sample to a specific library may hinder the generalization of inter-library patterns, posing a threat to their use or interpretation. Using a simulation and a real case study, we demonstrate this threat for different sampling methods. Our simulation shows that only when sampling for the disjunction of both libraries involved in the implication of a pattern, the implication generalizes well. Additionally, we demonstrate that real empirical data sampled using the GitHub search API does not behave as expected from our simulation. This identifies a potential threat relevant for many studies that use the GitHub search API for studying inter-library patterns.</p> </div> </div> </div>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Geometric deep learning improves generalizability of MHC-bound peptide predictions

<p>Full dataset and trained models from the manuscript "<strong>Geometric deep learning improves generalizability of MHC-bound peptide predictions</strong>".</p> <p>"outputs_and-BA_data.zip" contains the networks' outputs for each cross-validation experiment and a "full_dataset.csv" containing the initial BA data.<br>Note: this file has been updated (2024/11/26) due to errors in generating some of the previous csvs. In the earlier version, both MLP and CNN outputs reported were wrong. The correct values are now reported in the updated csvs.</p> <p>"trained_models.zip" contains all the trained models parameters</p> <p>"propedia_ssl.zip" contains all the 3D models from propedia used to train the 3D-SSL</p> <p>"pdb.zip" contains 3D models generated in PANDORA and used to train CNN, GNN and EGNN. It amounts to 145665 .pdb files, one for each human binding affinity entry from the initial dataset from O'Donnell et al. The list of entries used to actually train networks after filtering can be found in outputs_and-BA_data.zip", in the "full_dataset.csv" file.&nbsp;</p> <p>&nbsp;</p> <p>CHANGELOG v4:</p> <p>- In outputs_and-BA-data.zip, updated CNN_AlleleClustered_test_crossval.csv and CNN_shuffled_test_crossval.csv. These file had the wrong IDs paired with the network outputs.The IDs and labels are now consistent with the outputs.</p> <p>- Updated reference from the preprint to the published article.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Radarize: Enhancing Radar SLAM with Generalizable Doppler-Based Odometry

<p>This is the dataset release for the <strong><a href="https://www.sigmobile.org/mobisys/2024/" target="_blank" rel="noopener">ACM MobiSys 2024</a> </strong>paper "<strong>Radarize: Enhancing Radar SLAM with Generalizable Doppler-Based Odometry</strong>".</p> <ul> <li><strong>Project Website:</strong>&nbsp;<a href="http://radarize.github.io" target="_blank" rel="noopener">https://radarize.github.io</a></li> <li><strong>Project Code:</strong>&nbsp;<a href="http://github.com/ConnectedSystemsLab/radarize_ae" target="_blank" rel="noopener">github.com/ConnectedSystemsLab/radarize_ae</a></li> </ul> <p>If you found this useful, please cite&nbsp;</p> <pre><code>@inproceedings{sie2024radarize, author = {Sie, Emerson and Wu, Xinyu and Guo, Heyu and Vasisht, Deepak}, title = {Radarize: Enhancing Radar SLAM with Generalizable Doppler-Based Odometry}, booktitle = {The 22nd ACM International Conference on Mobile Systems, Applications, and Services (ACM MobiSys '24)}, year = {2024}, doi = {https://doi.org/10.1145/3643832.3661871}, }</code></pre>

opengpl-3.0-or-laterApr 2024View details →
zenodo36/100

Targeting protein-ligand neosurfaces with a generalizable deep learning tool

<p>Molecular recognition events between proteins drive biological processes in living systems. However, higher levels of mechanistic regulation have emerged, where protein-protein interactions are conditioned to small molecules. Despite recent advances, computational tools for the design of novel chemically-induced protein interactions have remained a challenging task for the field. Here, we present a computational strategy for the design of proteins that target neosurfaces, i.e. surfaces arising from protein-ligand complexes. To do so, we leveraged a geometric deep learning approach based on learned molecular surface representations and experimentally validated binders against three drug-bound protein complexes: Bcl2:Venetoclax, DB3:Progesterone and PDF1:Actinonin. All binders demonstrated high affinities and accurate specificities assessed by mutational and structural characterization. Remarkably, surface fingerprints previously trained only on proteins can be applied to neosurfaces emerging from small molecules, serving as a powerful demonstration of generalizability that is uncommon in other deep learning approaches. We anticipate that the designed chemically-induced protein interactions hold the potential to expand the sensing repertoire and the assembly of new synthetic pathways in engineered cells for innovative drug-controlled cell-based therapies</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Generalizable cone beam CT esophagus segmentation using physics-based data augmentation

<p>This upload contains open source AAPM thoracic auto-segmentation data (http://aapmchallenges.cloudapp.net/competitions/3) augmented with physics-based data augmentation technique introduced in this paper:https://iopscience.iop.org/article/10.1088/1361-6560/abe2eb .</p> <p>The original data contained thoracic planning CT along with organs-at-risk segmentation masks for Esophagus, heart, lungs and spinal cord. The physics-based augmentation pipeline was used to convert planning CT images to pseudoCBCT (psCBCT) images which are routinely used in weekly radiotherapy treatment sessions for cancer patients. Further geometric data augmentations are also applied to convert one planning CT/OAR dataset into 23 perfectly paired CT/psCBCT/OAR pairs.</p> <p>This dataset has been used to train deep learning models for organs-at-risk segmentation from psCBCT images and multitask simultaneous CBCT to CT translation and segmentation tasks.</p>

opencc-by-4.0Jun 2021View details →
dryad36/100

Thermally controlled intein splicing of engineered DNA polymerases provides a robust and generalizable solution for accurate and sensitive molecular diagnostics

<p>DNA polymerases are essential for nucleic acid synthesis, cloning, sequencing and molecular diagnostics technologies. Conditional intein splicing is a powerful tool for controlling enzyme reactions. We have engineered a thermal switch into thermostable DNA polymerases from two structurally distinct polymerase families by inserting a thermally activated intein domain into a surface loop that is integral to the polymerase active site, thereby blocking DNA or RNA template access. The fusion proteins are inactive but retain their structures such that the intein excises during a heat pulse delivered at 70–80°C to generate spliced, active polymerases. This straightforward thermal activation step provides a highly effective, one-component 'hot-start' control of PCR reactions that enables accurate target amplification by minimizing unwanted by-products generated by off-target reactions. In one engineered enzyme, derived from <em>Thermus aquaticus</em> DNA polymerase, both DNA polymerase and reverse transcriptase activities are controlled by the intein, enabling single-reagent amplification of DNA and RNA under hot-start conditions. This engineered polymerase provides high-sensitivity detection for molecular diagnostics applications, amplifying 5–6 copies of the tested DNA and RNA targets with &gt;95% certainty. The design principles used to engineer the inteins can be readily applied to construct other conditionally activated nucleic acid processing enzymes.</p>

opencc-zeroJun 2023View details →
ClinicalTrials.gov36/100

Surgical Critical Care Initiative (SC2i) Tissue and Data Acquisition Protocol (TDAP) in Burn Patients Improving the Robustness and Generalizability of Post-burn Sepsis Prediction With the Post-Burn Se

ClinicalTrials.gov study NCT07249762. IPD Sharing: YES. Countries: 1. Publications: 11.

controlledIPD-YESFeb 2026View details →
dryad36/100

Thermally controlled intein splicing of engineered DNA polymerases provides a robust and generalizable solution for accurate and sensitive molecular diagnostics

Open the record for dataset details and reuse information.

publicJun 2023View details →
dryad36/100

Clinical trial generalizability assessment in the big data era: a review

Open the record for dataset details and reuse information.

publicApr 2020View details →
dryad36/100

Data from: Generalizable physical descriptors of pool boiling heat transfer from unsupervised learning of images

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad32/100

Data from: A generalizable energetics-based model of avian migration to facilitate continental-scale waterbird conservation

Conserving migratory birds is made especially difficult because of movement among spatially disparate locations across the annual cycle. In light of challenges presented by the scale and ecology of migratory birds, successful conservation requires integrating objectives, management, and monitoring across scales, from local management units to ecoregional and flyway administrative boundaries. We present an integrated approach using a spatially explicit energetic-based mechanistic bird migration model useful to conservation decision-making across disparate scales and locations. This model moves a Mallard-like bird (Anas platyrhynchos), through spring and fall migration as a function of caloric gains and losses across a continental-scale energy landscape. We predicted with this model that fall migration, where birds moved from breeding to wintering habitat, took a mean of 27.5 d of flight with a mean seasonal survivorship of 90.5% (95% CI = 89.2%, 91.9%), whereas spring migration took a mean of 23.5 d of flight with mean seasonal survivorship of 93.6% (95% CI = 92.5%, 94.7%). Sensitivity analyses suggested that survival during migration was sensitive to flight speed, flight cost, the amount of energy the animal could carry, and the spatial pattern of energy availability, but generally insensitive to total energy availability per se. Nevertheless, continental patterns in the bird-use days occurred principally in relation to wetland cover and agricultural habitat in the fall. Bird-use days were highest in both spring and fall in the Mississippi Alluvial Valley and along the coast and near-shore environments of South Carolina. Spatial sensitivity analyses suggested that locations nearer to migratory endpoints were less important to survivorship; for instance, removing energy from a 1036 km2 stopover site at a time from the Atlantic Flyway suggested coastal areas between New Jersey and North Carolina, including the Chesapeake Bay and the North Carolina piedmont, are essential locations for efficient migration and increasing survivorship during spring migration but not locations in Ontario and Massachusetts. This sort of spatially explicit information may allow decision-makers to prioritize their conservation actions toward locations most influential to migratory success. Thus, this mechanistic model of avian migration provides a decision-analytic medium integrating the potential consequences of local actions to flyway-scale phenomena.

opencc-zeroDec 2015View details →
zenodo32/100

Targeting protein-ligand neosurfaces using a generalizable deep learning approach [benchmark dataset]

<p>PDB files and processed surface meshes for the binder recovery benchmark.&nbsp;For the larger PDBbind decoy set only the PDB files are provided.</p>

opencc-by-4.0Jun 2024View details →
dryad32/100

Generalizable EHR-R-REDCap pipeline for a national multi-institutional rare tumor patient registry

<p><strong>Objective</strong>: To develop a clinical informatics pipeline designed to capture large-scale structured EHR data for a national patient registry.</p> <p><strong>Materials and Methods</strong>: The EHR-R-REDCap pipeline is implemented using R-statistical software to remap and import structured EHR data into the REDCap-based multi-institutional Merkel Cell Carcinoma (MCC) Patient Registry using an adaptable data dictionary.</p> <p><strong>Results</strong>: Clinical laboratory data were extracted from EPIC Clarity across several participating institutions. Labs were transformed, remapped and imported into the MCC registry using the EHR labs abstraction (eLAB) pipeline. Forty-nine clinical tests encompassing 482,450 results were imported into the registry for 1,109 enrolled MCC patients. Data-quality assessment revealed highly accurate, valid labs. Univariate modeling was performed for labs at baseline on overall survival (N=176) using this clinical informatics pipeline.</p> <p><strong>Conclusion</strong>: We demonstrate feasibility of the facile eLAB workflow. EHR data is successfully transformed, and bulk-loaded/imported into a REDCap-based national registry to execute real-world data analysis and interoperability.</p>

opencc-zeroJan 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record