Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Pruned DNN model data for pruning example code
<p>Pruned DNN model datasets for example codes of neural network pruning.</p> <p>Example pruning codes are published in "https://github.com/FujitsuLaboratories/CAC/tree/main/cac/pruning".</p>
Data from: A national VS30 model for South Korea to combine nationwide dense borehole measurements with ambient seismic noise analysis
<p>The average shear-wave velocity within the top 30 m from the surface, V<sub>S30</sub>, represents site characteristics including the soil classification and site amplification that are essential information for building codes and seismic design. A novel method to determine a V<sub>S30</sub> model based on a composite analysis of borehole standard penetration test numbers (SPT N) and horizontal-to-vertical (H/V) spectral ambient noise ratios is introduced. A national V<sub>S30</sub> model for South Korea is determined using the method. The shear-wave velocity structures beneath 20 nationwide broadband seismic stations are determined using the H/V analysis. The SPT N data are collected from 175,619 nationwide densely-distributed boreholes. The shear-wave velocity models from SPT N values are calibrated for the local reference velocity models from H/V analysis. A representative relationship between the SPT N values and shear-wave velocities is introduced. A national V<sub>S30</sub> model for South Korea is determined using the calibrated SPT N models at the nationwide boreholes. The V<sub>S30</sub> model is verified by comparisons with local field measurements. The proposed model is consistent with the USGS model based on a surface slope analysis. The V<sub>S30</sub> structure presents high correlation with geological and topographic features. The V<sub>S30</sub> values are low in coastal (low topographic) areas, and high in mountain (high topographic) areas. Apparent linear relationship is observed between V<sub>S30</sub> and topography. The western and southeastern coastal regions may be vulnerable to strong seismic shaking.</p>
Infectious disease modelling and the dynamics of the active cases - Data
<p>We developed a model that tries to describe the dynamics of the spread of a disease among a population, in particular the progress of infected active cases. The model is then applied to describe Italy CoViD-19 outbreak and subsequently, we tried to predict possible scenarios.</p>
Data for Model test on the passive failure of slurry shield tunneling in circular-gravel stratum
<p>Data for Model test on the passive failure of slurry shield tunneling in circular-gravel stratum.</p>
Audio samples from generative models trained on the TIMIT speech data.
<p>The snippets include samples and reconstructions. All samples are completely unconditional and utilise only the prior internal representations learned by the models. Reconstructions are computed from a given test audio snippet by first encoding it to a learned representation and then decoding that to a reconstruction of the audio.</p> <p>All models are trained on the TIMIT speech dataset (<a href="https://catalog.ldc.upenn.edu/LDC93s1">https://catalog.ldc.upenn.edu/LDC93s1</a>). Some snippets are from models trained at different temporal resolutions denoted by `s1` and `s64`. We refer to the paper for details.</p> <p>The files include:</p> <ul> <li>`clockwork-vae-s64-reconstruction-*` <ul> <li>Four reconstructions using a two-layered Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`clockwork-vae-s64-sample-*` <ul> <li>Four samples from the prior of a Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`original-*` <ul> <li>Four original samples from TIMIT corresponding in pairs to the reconstructions.</li> </ul> </li> <li>`vrnn-s64-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`vrnn-s1-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`srnn-s64-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`srnn-s1-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s64-sample-*` <ul> <li>Four samples from a WaveNet trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s1-sample-*` <ul> <li>Two samples from a WaveNet trained with temporal resolution s=64.</li> </ul> </li> </ul>
Predicting atrial fibrillation recurrence by combining population data and virtual cohorts of patient-specific left atrial models
<p><strong>Abstract</strong></p> <p><strong>Background: </strong>Current ablation therapy for atrial fibrillation is sub-optimal and long-term response is challenging to predict. Clinical trials identify bedside properties that provide only modest prediction of long-term response in populations, while patient-specific models in small cohorts primarily explain acute response to ablation. We aimed to predict long-term atrial fibrillation recurrence after ablation in large cohorts, by using machine learning to complement biophysical simulations by encoding more inter-individual variability.</p> <p><strong>Methods: </strong>Patient-specific models were constructed for 100 atrial fibrillation patients (43 paroxysmal, 41 persistent, 16 long-standing persistent), undergoing first ablation. Patients were followed for 1-year using ambulatory ECG monitoring. Each patient-specific biophysical model combined differing fibrosis patterns, fibre orientation maps, electrical properties and ablation patterns to capture uncertainty in atrial properties and to test the ability of the tissue to sustain fibrillation. These simulation stress tests of different model variants were post-processed to calculate atrial fibrillation simulation metrics. Machine learning classifiers were trained to predict atrial fibrillation recurrence using features from the patient history, imaging and atrial fibrillation simulation metrics.</p> <p><strong>Results: </strong>We performed 1100 atrial fibrillation ablation simulations across 100 patient-specific models. Models based on simulation stress tests alone showed a maximum accuracy of 0.63 for predicting long-term fibrillation recurrence. Classifiers trained to history, imaging and simulation stress tests (average ten-fold cross-validation area under the curve 0.85 ± 0.09, recall 0.80 ± 0.13, precision 0.74 ± 0.13) outperformed those trained to history and imaging (area under the curve 0.66 ± 0.17), or history alone (area under the curve 0.61 ± 0.14). </p> <p><strong>Conclusion: </strong>A novel computational pipeline accurately predicted long-term atrial fibrillation recurrence in individual patients by combining outcome data with patient-specific acute simulation response. This technique could help to personalise selection for atrial fibrillation ablation.</p> <p><strong>Dataset Description: </strong>We include surface meshes in vtk format, consisting of the nodes, triangular elements, the atrial coordinate fields defined on the nodes, and the endocardial and epicardial fibre fields defined on the elements. </p> <p>We also include universal atrial coordinate fields alpha and beta, which are a lateral-septal coordinate and posterior-anterior coordinate for the LA. More details on the coordinate construction are given in our manuscript and <a href="https://www.ncbi.nlm.nih.gov/pubmed/31026761">https://www.ncbi.nlm.nih.gov/pubmed/31026761</a>. These coordinates can be used for registering datasets. </p> <p><strong>Publication</strong>: https://pubmed.ncbi.nlm.nih.gov/35089057/</p>
Data from: Physiological and transcriptional immune responses of a non-model arthropod to infection with different entomopathogenic groups
<p>Insect immune responses to multiple pathogen groups including viruses, bacteria, fungi, and entomopathogenic nematodes have traditionally been documented in model insects such as <em>Drosophila melanogaster</em>, or medically important insects such as <em>Aedes aegypti</em>. Despite their potential importance in understanding the efficacy of pathogens as biological control agents, these responses are infrequently studied in agriculturally important pests. Additionally, studies often neglect to investigate responses against different pathogen groups, and typically focus on only a single time point during infection. As such, a robust understanding of immune system responses over the time of infection is often lacking. This study was conducted to understand how 3<sup>rd</sup> instar larvae of the major insect pest <em>Helicoverpa zea</em> responded over time to infection by four different pathogenic groups: viruses, bacteria, fungi, and entomopathogenic nematodes. Physiological immune responses were assessed at 4-, 24-, and 48-hours post-infection by measuring hemolymph phenoloxidase concentrations, hemolymph prophenoloxidase concentrations, hemocyte counts, and encapsulation ability. Transcriptional immune responses were measured at 24-, 48-, and 72-hours post-infection by quantifying the expression of <em>PPO2</em>, <em>Argonaute-2</em>, <em>JNK</em>, <em>Dorsal</em>, and <em>Relish</em>. This gene set covers the major known immune pathways: phenoloxidase cascade, siRNA, JNK pathway, Toll pathway, and IMD pathway. Our results indicate <em>H. zea</em> has an extreme immune response to <em>Bacillus thuringiensis</em> bacteria, a mild response to <em>Helicoverpa armigera</em> nucleopolyhedrovirus, and no detectable response to either the fungus <em>Beauveria bassiana</em> or <em>Steinernema carpocapsae </em>nematodes.</p>
Source code and data from informed dispersal model on range expansion
<p>In those files, you can find the source code to run the model and the R codes to process the data and create the figures</p>
Data for Model solution for canopy flows
<p>Data for Model solution for canopy flows</p>
Data from: The relationship between native species richness and exotic species richness or occurrence will always be negative when the total number of species is accounted for in statistical models: A response to Beaury et al.
Beaury et al. (2020) attempt to address the scale dependence of evidence for biotic resistance by including environmental covariates that can account for total species richness. However, this approach will incorrectly estimate relationships, driven by the accuracy of the covariates rather than the true relationship between native and non-native species.
Supplementary material 1 from: Hatami R, Inglis G, Lane SE, Growcott A, Kluza D, Lubarsky C, Jones-Todd C, Seaward K, Robinson AP (2022) Modelling the likelihood of entry of marine non-indigenous species from internationally arriving vessels to maritime ports: a case study using New Zealand data. NeoBiota 72: 183-203. https://doi.org/10.3897/neobiota.72.77266
Supplementary materials
Occurence data for species distribution modelling of wild Coffea canephora
<p><span>The assessment of population vulnerability under climate change is crucial for planning conservation as well as for ensuring food security. <em>Coffea canephora</em> is, in its native habitat, an understory tree that is mainly distributed in the lowland rainforests of tropical Africa. Also known as Robusta, its commercial value constitutes a significant revenue for many human populations in tropical countries. Comparing ecological and genomic vulnerabilities within the species' native range can provide valuable insights about habitat loss and the species' adaptive potential, allowing to identify genotypes that may be act as a resource for varietal improvement. By applying species distribution models, we assessed ecological vulnerability as the decrease in climatic suitability under future climatic conditions from 492 occurrences. We then quantified genomic vulnerability (or risk of maladaptation) as the allelic composition change required to keep pace with predicted climate change. Genomic vulnerability was estimated from genomic environmental correlations throughout the native range. Suitable habitat was predicted to diminish to half its size by 2050, with populations near coastlines and around the Congo River being the most vulnerable. Whole-genome sequencing revealed 165 candidate SNPs associated to climatic adaptation in <em>C. canephora</em>, which were located in genes involved in plant response to biotic and abiotic stressors. Genomic vulnerability was higher for populations in West Africa and in the region at the border between DRC and Uganda. Despite an overall low correlation between genomic and ecological vulnerability at broad scale, these two components of vulnerability overlap spatially in ways that may become damaging. Genomic vulnerability was estimated to be 23% higher in populations where habitat will be lost in 2050 compared to regions where habitat will remain suitable. These results highlight how ecological and genomic vulnerabilities are relevant when planning on how to cope with climate change regarding an economically important species.</span></p>
Supplementary Material: Predictive model using Cross Industry Standard Process for Data Mining
<p>The Supplementary Material of the paper "Supplementary Material: Predictive model using Cross Industry Standard Process for Data Mining" includes: <br> 1) APPENDIX 1: SQL Statements for data extraction. Appendix 2: Interview for operating Staff.<br> 2) The DataSet of the normalized data to define the predictive model.</p>
NZESM & UKESM data for JAMES study on climatic changes associated with a nested ocean model in the region around New Zealand.
<p>NZESM & UKESM data for JAMES study on climatic changes associated with a nested ocean model in the region around New Zealand.</p>
Haplotype-phasing of long-read HiFi data to enhance structural variant detection through a Skip-Gram model
<p>Example dataset for DipPAV</p>
Data for "Impact of warmer sea surface temperature on the global pattern of intense convection: insights from a global storm resolving model"
<p>Data relevant to a manuscript on X-SHiELD, for submission to GRL.</p>
3D-MSNet: A point cloud based deep learning model for untargeted feature detection and quantification in profile LC-HRMS data
<p>Supplementary data of 3D-MSNet</p>
Data from: Empirical estimation of skin resistance to water loss in amphibians: agar evaluation as a non-resistance model to evaporation
<p>Total resistance (R<sub>T</sub>) to evaporative water loss (EWL) in amphibians is given by the sum of the boundary layer (<em>r</em><sub>b</sub>) and the skin resistance (<em>r</em><sub>s</sub>). Thus, <em>r</em><sub>s</sub> can be determined if the <em>r</em><sub>b</sub> component is defined (<em>r</em><sub>s</sub> = R<sub>T</sub> - <em>r</em><sub>b</sub>). The use of agar models has become the standard technique to estimate <em>r</em><sub>b</sub> under the assumption that agar surface imposes no barrier to evaporation (<em>r<sub>s</sub></em> = 0). We evaluated this assumption by determining EWL rates and <em>r</em><sub>b</sub> values from exposed surfaces of free water, a physiological solution mimicking the osmotic properties of a generalized amphibian, and agar gels prepared at various concentrations either using water or physiological solution as diluent. Water evaporation was affected by both, the presence of solutes and agar concentration. Models prepared with agar at 5% concentration in water provided the most practical and appropriate proxy for the estimation of <em>r</em><sub>b</sub>.</p> <p> </p> <p> </p>
Original data for "Multiple co-existing structures of an RNA four-way junction resolved by FRET, SAXS, and integrative modeling"
<p>Experimental single-molecule FRET data (Intensity ratio histograms) and starting structures used for rigid body docking for an RNA four-way junction related to the hairpin ribozyme.</p> <p> </p>
Data from: Incorporating single-step strategy into random regression model to enhance genomic prediction of longitudinal trait
In prediction of genomic values, single-step method has been demonstrated to outperform multi-step methods. In statistical analyses of longitudinal traits, random regression test-day model (RR-TDM) has clear advantages over other models. Our goal in this study was to evaluate the performance of the model integrating both single-step and RR-TDM prediction methods, called single-step random regression test-day model (SS RR-TDM), in comparison with the pedigree-based RR-TDM and genomic best linear unbiased prediction (GBLUP) model. We performed extensive simulations to exploit potential advantages of SS RR-TDM over the other two models under various scenarios with different level of heritability, the number of QTL as well as the selection scheme. SS RR-TDM was found to achieve the highest accuracy and unbiasedness under all scenarios, exhibiting robust prediction ability in longitudinal trait analyses. Moreover, SS RR-TDM showed better persistency of accuracy over generations than GBLUP model. In addition, we also found that the SS RR-TDM had advantages over RR-TDM and GBLUP in terms of a real dataset of human contributed by the GAW18 workshop. The findings in our study firstly proved the feasibility and advantages of the SS RR-TDM, and further enhanced strategies for the genomic prediction of longitudinal traits in the future.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.