Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,481 results for “data processing”

Learn how ShareScore rates datasets ↗
zenodo36/100

Processed data supporting the manuscript "Cutting the sap: first molecular phylogeny of twig-girdler longhorn beetles (Coleoptera: Cerambycidae: Lamiinae: Onciderini) suggests shifts in host plant attack behaviors contributed to morphological evolution"

<div><strong>Processed data supporting the manuscript:&nbsp;</strong>Cutting the sap: first molecular phylogeny of twig-girdler longhorn beetles (Coleoptera: Cerambycidae: Lamiinae: Onciderini) suggests shifts in host plant attack behaviors contributed to morphological evolution</div> <div>&nbsp;</div> <div><strong>By:</strong> Diego de S. Souza 1, 2, Rowan L. K. French 3, Jos&eacute; O. Silva J&uacute;nior 4, Eugenio H. Nearns 5, Luciane Marinoni 4, Ian P. Swift 6, Kelly B. Miller 7, Felix A. H. Sperling 2 &amp; Marcela L. Monn&eacute; 1</div> <div>&nbsp;</div> <div>1 Department of Entomology, National Museum, Federal University of Rio de Janeiro, Rio de Janeiro, Rio de Janeiro, Brazil.</div> <div>2 Department of Biological Sciences, University of Alberta, Edmonton, Alberta, Canada.</div> <div>3 Department of Ecology and Evolutionary Biology, University of Toronto, Toronto, Ontario, Canada.</div> <div>4 Department of Zoology, Federal University of Paran&aacute;, Curitiba, Paran&aacute;, Brazil.</div> <div>5 National Museum of Natural History, Smithsonian Institution, Washington, DC, USA.</div> <div>6 California State Collection of Arthropods, Sacramento, California, USA.</div> <div>7 Department of Biology and Museum of Southwestern Biology, University of New Mexico, Albuquerque, New Mexico, USA.</div> <div>&nbsp;</div> <div>Corresponding author: Diego de S. Souza, dsouza@fieldmuseum.org. Current affiliation: Field Museum of Natural History, Chicago, Illinois, USA.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div><strong>List of Contents:&nbsp;</strong></div> <div>&nbsp;</div> <div><strong>Onciderini_concat_matrix.phy</strong></div> <div>Concatenated matrix (cox1, Wg and CPS) used for the phylogenetic analyses of Onciderini (Coleoptera: Cerambycidae: Lamiinae: Onciderini).&nbsp;</div> <div>&nbsp;</div> <div><strong>PartitionFinder_AICc_best_scheme.txt</strong></div> <div>Results from PartitionFinder v2.1.1, containing the best partitioning scheme for the concatenated matrix of Onciderini, identified using the corrected Akaike Information Criterion (AICc), with model definitions for use in the phylogenetic analyses.</div> <div>&nbsp;</div> <div><strong>RAxML_Onciderini_concat_matrix (zip file)</strong></div> <div>- Onciderini_concat_matrix.phy: concatenated matrix (cox1, Wg and CPS) used in the RAxML phylogenetic analyses of Onciderini (Coleoptera: Cerambycidae: Lamiinae: Onciderini).</div> <div>- Partitions_AICc_RAxML.txt: partitioning scheme used in the RAxML analysis as predefined by PartitionFinder v2.1.1 using the corrected Akaike Information Criterion (AICc).</div> <div>- RAxML_bestTree.Onciderini_concat_matrix_ML: best-scoring maximum likelihood tree inferred by RAxML for the concatenated matrix of Onciderini.</div> <div>- RAxML_bipartitions.Onciderini_concat_matrix_final: bipartitions (clades) of the maximum likelihood tree inferred by RAxML with support values estimated from 1,000 pseudoreplicates.</div> <div>- RAxML_bipartitionsBranchLabels.Onciderini_concat_matrix_final: final maximum likelihood tree inferred by RAxML for the concatenated matrix of Onciderini, with labeled branches showing bootstrap support values.</div> <div>- RAxML_bootstrap.Onciderini_concat_matrix_bootstrap: bootstrap trees generated from a non-parametric bootstrap analysis in RAxML based on 1,000 pseudoreplicates.</div> <div>- RAxML_info.Onciderini_concat_matrix_bootstrap: log file containing details of the bootstrap analysis, including the settings and parameters used in the non-parametric bootstrap runs in RAxML.</div> <div>- RAxML_info.Onciderini_concat_matrix_final: log file summarizing the RAxML analysis, including settings and convergence statistics for the final maximum likelihood tree.</div> <div>- RAxML_info.Onciderini_concat_matrix_ML: log file containing details of the maximum likelihood tree search, including the parameters and models applied during the maximum likelihood analysis conducted by RAxML.</div> <div>- RAxML_log.Onciderini_concat_matrix_ML: log file of the maximum likelihood tree search for the concatenated matrix of Onciderini.</div> <div>- RAxML_parsimonyTree.Onciderini_concat_matrix_ML: parsimony starting tree used by RAxML during the maximum likelihood analysis for the concatenated matrix of Onciderini.</div> <div>- RAxML_result.Onciderini_concat_matrix_ML: maximum likelihood tree inferred by RAxML from the concatenated matrix of Onciderini, summarizing the tree topology and likelihood score for the best tree obtained.</div> <div>&nbsp;</div> <div><strong>BI_AICc_Onciderini_concat_matrix (zip file)</strong></div> <div>- BI_AICc_Onciderini_concat_matrix.nex: nexus file containing the concatenated matrix of Onciderini used for Bayesian Inference (BI), including the best-fit model scheme identified by PartitionFinder and MCMC parameters for running the analysis in MrBayes.</div> <div>- BI_AICc_Onciderini_concat_matrix.nex_r1_r2_combined_consensus.tree: consensus tree from two combined independent Bayesian Inference (BI) runs based on the concatenated matrix of Onciderini, after discarding the first 25% of initial generations as burn-in.</div> <div>- BI_AICc_Onciderini_concat_matrix.nex.run1.p: log file containing parameter values and likelihood scores from the first run of the Bayesian Inference (BI) based on the concatenated matrix of Onciderini.</div> <div>- BI_AICc_Onciderini_concat_matrix.nex.run2.p: log file containing parameter values and likelihood scores from the second run of the Bayesian Inference (BI) based on the concatenated matrix of Onciderini.</div> <div>&nbsp;</div> <div><strong>BEAST2_Onciderini_BD_lognormal (zip file)</strong></div> <div>- BEAUTi_Onciderini_BD_lognormal.xml: XML file generated by BEAUTi for running BEAST2, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a lognormal distribution.</div> <div>- BEAST2_Onciderini_BD_lognormal_run[1-8].log: log files from eight independent runs of BEAST2, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a lognormal distribution.</div> <div>- TreeAnnotator_Onciderini_BD_lognormal_run1-run8_consensus.out: TreeAnnotator output file combining the results of eight BEAST2 runs based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a lognormal distribution.</div> <div>- TreeAnnotator_Onciderini_BD_lognormal_run1-run8_consensus.tre: consensus tree from eight combined BEAST2 runs, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a lognormal distribution, after discarding the first 10% of initial generations as burn-in.</div> <div>&nbsp;</div> <div><strong>BEAST2_Onciderini_BD_exponential (zip file)</strong></div> <div>- BEAUTi_Onciderini_BD_exponential.xml: XML file generated by BEAUTi for running BEAST2, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and an exponential distribution.</div> <div>- BEAST2_Onciderini_BD_exponential_[1-8].log: log files from eight independent runs of BEAST2, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and an exponential distribution.</div> <div>- TreeAnnotator_Onciderini_BD_exponential_run1-run8_consensus.out: TreeAnnotator output file combining the results of eight BEAST2 runs based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and an exponential distribution.</div> <div>- TreeAnnotator_Onciderini_BD_exponential_run1-run8_consensus.tre: consensus tree from eight combined BEAST2 runs, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and an exponential distribution, after discarding the first 10% of initial generations as burn-in.</div> <div>&nbsp;</div> <div><strong>BEAST2_Onciderini_BD_uniform (zip file)</strong></div> <div>- BEAUTi_Onciderini_BD_uniform.xml: XML file generated by BEAUTi for running BEAST2, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a uniform distribution.</div> <div>- BEAST2_Onciderini_BD_uniform_[1-8].log: log files from eight independent runs of BEAST2, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a uniform distribution.</div> <div>- TreeAnnotator_Onciderini_BD_uniform_run1-run8_consensus.out: TreeAnnotator output file combining the results of eight BEAST2 runs based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a uniform distribution.</div> <div>- TreeAnnotator_Onciderini_BD_uniform_run1-run8_consensus.tre: consensus tree from eight combined BEAST2 runs, based on the concatenated matrix of Onciderini, using a birth-death (BD) process model and a uniform distribution, after discarding the first 10% of initial generations as burn-in.</div> <div>&nbsp;</div> <div><strong>Comparative_analyses (zip file)</strong></div> <div><strong>RawData (folder):</strong> raw morphometric and girdling data, plus tree that was later pruned for downstream comparative analyses; these data were used as input for the OncidHeadDimorphism-DatasetPREP-FINAL.R data cleaning script.&nbsp;</div> <div>- Onciderini_BD_lognormal_run1-run8_consensus.nwk: newick version of TreeAnnotator_Onciderini_BD_lognormal_run1-run8_consensus.tre (outputted as a newick file by importing the .tre file into FigTree and exporting in newick format).</div> <div>- Measurements_Onciderini_Raw_Final.csv: individual-level raw morphometric data for Onciderini.</div> <div>- Behav_Matrix_2_states_trimmed_2022-12-15.csv: species-level data on girdling status for Onciderini species, with all Lochmaeocles species classified as girdlers (2 behavioral states across Onciderini species).&nbsp;</div> <div>- Matrix_3_states_trimmed_final.csv: species-level data on girdling status for Onciderini species, with all Lochmaeocles species classified as facultative girdlers (3 behavioral states across Onciderini species).&nbsp;</div> <div>- Matrix_2_states_trimmed_1LochGirdler_Final.csv: species-level data on girdling status for Onciderini species, with only one Lochmaeocles species (L. tessellatus) classified as a girdler (2 behavioral states across Onciderini species).&nbsp;</div> <div>&nbsp;</div> <div><strong>ProcessedData (folder):&nbsp;</strong>filtered data and pruned trees outputted by the OncidHeadDimorphism-DatasetPREP-FINAL.R script</div> <div>- oncid_f_36spp_clean.csv: dataset of species means and log-ratios for morphometric traits in females, plus girdling data; only includes species that have girdling data and are in the phylogenetic tree</div> <div>- oncid_m_42spp_clean.csv: dataset of species means and log-ratios for morphometric traits in males, plus girdling data; only includes species that have girdling data and are in the phylogenetic tree</div> <div>- oncid_sd_35spp_clean.csv: dataset of species means for sexual dimorphism in morphometric traits, plus girdling data; only includes species that have girdling data and are in the phylogenetic tree</div> <div>- oncid_girdlingbehav_allingroupspp_clean.csv: full dataset of girdling behavior for 56 Onciderini species that are in the phylogenetic tree; includes separate columns for the three alternative girdling classification schemes</div> <div>- oncid_tree_behavfull_56spp.nwk: pruned phylogenetic tree for the full girdling dataset (56 species)</div> <div>- oncid_tree_f_36spp.nwk: pruned phylogenetic tree for the female morphometric dataset (36 species)</div> <div>- oncid_tree_m_42spp.nwk: pruned phylogenetic tree for the male morphometric dataset (42 species)</div> <div>- oncid_tree_mf_43spp.nwk: pruned phylogenetic tree for all species with morphometric data for males or females; used for the stochastic character map next to the heatmap plot (Fig 4)</div> <div>- oncid_tree_sd_35spp.nwk: pruned phylogenetic tree for the sexual dimorphism dataset (35 species)</div> <div>&nbsp;</div> <div><strong>FittedModels (folder): </strong>fitted models (mvgls, model comparison analyses, OUM models, simmaps) outputted by the OncidHeadDimorphism-Analysis-FINAL.R script&nbsp;</div> <div>- MacroModelFits_logRtraits_f_36spp-2024-10-12.Rdata: summary of model comparison results for Brownian Motion (BM), single-peak Ornstein-Uhlenbeck (OU), multipeak OU (OUM), and multi-rate Brownian motion (BMM) models (univariate and multivariate) fitted to female morphometric data across 100 stochastic character maps of girdling behavior</div> <div>- MacroModelFits_logRtraits_m_42spp-2024-10-12.Rdata: summary of model comparison results for Brownian Motion (BM), single-peak Ornstein-Uhlenbeck (OU), multipeak OU (OUM), and multi-rate Brownian motion (BMM) models (univariate and multivariate) fitted to male morphometric data across 100 stochastic character maps of girdling behavior</div> <div>- MacroModelFits-SDDI-35spp_2024-10-11.Rdata: summary of model comparison results for Brownian Motion (BM), single-peak Ornstein-Uhlenbeck (OU), multipeak OU (OUM), and multi-rate Brownian motion (BMM) models (univariate and multivariate) fitted to sexual dimorphism data across 100 stochastic character maps of girdling behavior</div> <div>- mvgls-results-headsize-mf-2024-10-12.Rdata: fitted mvgls regression models for male and female traits (analyzed separately)</div> <div>- mvgls-results-sddi-2024-10-12.Rdata: fitted mvgls regression models for sexual dimorphism</div> <div>- OUM_headtraits_f_36spp-2024-10-12.Rdata: fitted OUM models and summary statistics for female head traits</div> <div>- OUM_headtraits_m_42spp-2024-10-12.Rdata: fitted OUM models and summary statistics for male head traits</div> <div>- OUM-SDDI-35spp-2024-10-12.Rdata: fitted OUM models and summary statistics for sexual dimorphism in head traits</div> <div>- simmaps_ard_full_2state.RDS: stochastic character maps of girdling behaviour for all 56 species with girdling data, with two behavioral states (girdling or non-girdling) - all Lochmaeocles species are classified as girdlers</div> <div>- simmaps_ard_full_3state.RDS: stochastic character maps of girdling behaviour for all 56 species with girdling data, with three behavioral states (girdling, non-girdling, or facultative girdling)</div> <div>- simmaps_ard_full_1Loch.RDS: stochastic character maps of girdling behaviour for all 56 species with girdling data, with two behavioral states (girdling or non-girdling) - only one Lochmaeocles species (L. tessellatus) is classified as a girdler</div> <div>- simmaps_ard_m.RDS: stochastic character map for the 42 species used in the analyses of male morphometric traits</div> <div>- simmaps_ard_f.RDS: stochastic character map for the 36 species used in the analyses of female morphometric traits</div> <div>- simmaps_ard.RDS: stochastic character map for the 35 species used in the sexual dimorphism analyses; 2 behavioral states.</div> <div>&nbsp;</div> <div><strong>Rscripts (folder):&nbsp;</strong>R scripts used to process data and run phylogenetic comparative analyses of head size and girdling behavior.</div> <div>- OncidHeadDimorphism-DatasetPREP-FINAL.R: R script used to filter data and prune trees from the RawData folder for downstream phylogenetic comparative analyses; outputs of this script are in the ProcessedData folder.</div> <div>- OncidHeadDimorphism-Analysis-FINAL.R: R script used to analyze data in the ProcessedData folder to answer questions about the origin and evolution of girdling behavior and the relationship between girdling and head size or head size sexual dimorphism; fitted models outputted by this script are in the FittedModels folder</div> <div>- OncidHeadDimorphism-Plots-FINAL.R: R script used to generate plots for the manuscript<br><br></div>

opencc-by-4.0May 2025View details →
zenodo36/100

Geophysical Signals from Magma Propagation: Experimental and Processed Data

<p>This repository contains the experimental data analyzed and interpreted in the manuscript titled "<em>Geophysical signals induced by magma propagation: Insights from analog experiments</em>" by S. Furst, J. Vandemeulebrouck, and V. Pinel. It includes video recordings, timelapse photos, accelerometer data, and deformation data. Additionally, there are three MATLAB scripts for post-processing the timelapse photos following the approach described in the manuscript. The results of the MFP analysis on the accelerometer data, as well as the outcomes from the COMSOL Multiphysics simulations, are also included in the repository.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Processed species-level data for "Global Biodiversity Loss from Outsourced Deforestation"

<p>This repository holds processed species-level data for the following paper:</p> <blockquote> <p>Wiebe, R.A.* and Wilcove, D.S.&nbsp;<em>2025</em>. Global Biodiversity Loss from Outsourced Deforestation.</p> </blockquote> <p>*<a href="mailto:rwiebe@princeton.edu">rwiebe@princeton.edu</a>, Guyot Hall, Princeton University, Princeton, NJ</p> <p>In this analysis, we quantified the range loss to forest-dwelling vertebrates that is attributable to demand for agricultural and forestry products by developed countries in 2001-2015.</p> <p>We used several datasets in this analysis. Data on forest cover and forest loss were sourced from&nbsp;<em>Hansen et al.</em>&nbsp;and publicly available to download online:&nbsp;<a href="https://earthenginepartners.appspot.com/science-2013-global-forest/download_v1.7.html" rel="nofollow">https://earthenginepartners.appspot.com/science-2013-global-forest/download_v1.7.html</a>. Data on range maps for species were sourced from the IUCN and downloaded from the IUCN&rsquo;s API (<a href="https://www.iucnredlist.org/resources/spatial-data-download" rel="nofollow">https://www.iucnredlist.org/resources/spatial-data-download</a>) and its partner BirdLife International&rsquo;s API (<a href="http://datazone.birdlife.org/species/requestdis" rel="nofollow">http://datazone.birdlife.org/species/requestdis</a>). These datasets are made available for academic research on request, and an API key is necessary for their downloads. We sourced data on land use attribution from Nguyen Tien Hoang and Keiichiro Kanemoto, from their 2021 publication &ldquo;Mapping the deforestation footprint of nations reveals growing threat to tropical forests.&rdquo; These data are not publicly available, but the authors kindly provided the data on request.</p> <p>Code necessary to replicate the analysis can be found in a public repository at: https://github.com/AlexWiebe/Outsourced-Biodiversity-Loss.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Data for the publication "Evaluation of climatic impacts of ice nucleating particles through precipitation process with a GCM"

<p>These data are a set of annual-mean values for 5yr simulations using the MIROC6-SPRINTARS global aerosol-climate model with different treatments of precipitation (i.e., diagnostic and prognostic). The outputs include diagnostics from the satellite simulator COSP2.<br>The data are used in the manuscript entitled "Evaluation of climatic impacts of ice nucleating particles through precipitation process with a GCM". All data used in this study are available from the corresponding author upon request.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Processed data for Evidence of horizontal gene transfer and environmental selection impacting antibiotic resistance evolution in soil-dwelling Listeria

<p>Processed/source data for the manuscript Evidence of horizontal gene transfer and environmental selection impacting antibiotic resistance evolution in soil-dwelling <em>Listeria</em>.</p>

opencc-by-4.0Nov 2024View details →
dryad36/100

Data from: Developmental lead exposure has mixed effects on butterfly cognitive processes

While the effects of lead pollution have been well studied in vertebrates, it is unclear to what extent lead may negatively affect insect cognition. Lead pollution in soils can elevate lead in plant tissues, suggesting it could negatively affect neural development of insect herbivores. We used the cabbage white butterfly (Pieris rapae) as a model system to study the effect of lead pollution on insect cognitive processes, which play an important role in how insects locate and handle resources. Cabbage white butterfly larvae were reared on a 4-ppm lead diet, a concentration representative of vegetation in polluted sites; we measured eye size and performance on a foraging assay in adults. Relative to controls, lead-reared butterflies did not differ in time or ability to search for a food reward associated with a less preferred color. Indeed, lead-treated butterflies were more likely to participate in the behavioral assay itself. Lead exposure did not negatively affect survival or body size, and it actually sped up development time. The effects of lead on relative eye size varied with sex: lead tended to reduce eye size in males, but increase eye size in females. These results suggest that low levels of lead pollution may have mixed effects on butterfly vision, but only minimal impacts on performance in foraging tasks, although follow-up work is needed to test whether this result is specific to cabbage whites, which are often associated with disturbed areas.

opencc-zeroDec 2015View details →
dryad36/100

Complex ecological phenotypes on phylogenetic trees: a Markov process model for comparative analysis of multivariate count data

The evolutionary dynamics of complex ecological traits – including multistate representations of diet, habitat, and behavior – remain poorly understood. Reconstructing the tempo, mode, and historical sequence of transitions involving such traits poses many challenges for comparative biologists, owing to their multidimensional nature. Continuous-time Markov chains (CTMC) are commonly used to model ecological niche evolution on phylogenetic trees but are limited by the assumption that taxa are monomorphic and that states are univariate categorical variables. A necessary first step in the analysis of many complex traits is therefore to categorize species into a pre-determined number of univariate ecological states, but this procedure can lead to distortion and loss of information. This approach also confounds interpretation of state assignments with effects of sampling variation because it does not directly incorporate empirical observations for individual species into the statistical inference model. In this study, we develop a Dirichlet-multinomial framework to model resource use evolution on phylogenetic trees. Our approach is expressly designed to model ecological traits that are multidimensional and to account for uncertainty in state assignments of terminal taxa arising from effects of sampling variation. The method uses multivariate count data for individual species to simultaneously infer the number of ecological states, the proportional utilization of different resources by different states, and the phylogenetic distribution of ecological states among living species and their ancestors. The method is general and may be applied to any data expressible as a set of observational counts from different categories.

opencc-zeroApr 2020View details →
dryad36/100

Data from: The roles of non-production vegetation in agroecosystems: a research framework for filling process knowledge gaps in a social-ecological context

<p>1. An ever-expanding human population, climatic changes, and the spread of intensive farming practices is putting increasing pressure on agroecosystems and their inherent biodiversity. Non-production vegetation elements, such as woody patches, riparian margins, and restoration plantings, are vital for conserving agroecosystem biodiversity. Further, such elements are key building blocks that are manipulated via land management, thereby influencing the biotic and abiotic processes that underpin functioning agroecosystems.</p> <p>2. Despite this critical role, there has been a lack of synthesis on which types of vegetation elements drive and/or support ecological processes, and the mechanisms by which this occurs. Using a systematic, quantitative literature review of 342 articles, we asked: what are the effects of non-production vegetation on agroecosystem processes and how are these processes measured within global agroecosystems?</p> <p>3. Woody patches, hedgerows and borders, riparian margins, and shelterbelts were the most studied types of non-production vegetation. The majority (61%) of studies showed positive effects of non-production vegetation on ecological processes, where the presence, level or rate of the studied process was increased or enhanced.</p> <p>4. However, four key research gaps were revealed: (1) most studies (83%) used proxies for, instead of direct measurements of, ecosystem processes related to non-production vegetation; (2) study designs used to investigate non-production vegetation effects on ecosystem processes directly were largely limited to observational comparisons of non-production vegetation types, farm-scale vegetation configurations, and different proximities to vegetation in terms of the effect on ecological processes; relatively few studies used manipulative experiments (3) the relatively few studies directly measuring ecosystem processes were dominated by four process categories: invertebrate biocontrol, predator and natural enemy spillover, animal movement, and ecosystem cycling, and (4) the methods used to directly measure non-production vegetation effects comprised a surprisingly limited set of approaches.</p> <p>5. To fill key research gaps that will inform the use of non-production vegetation to enhance agroecosystem processes, we present a framework for future research that emphasises the need to combine an understanding of human decision making with carefully-designed and targeted investigations into the roles of taxa, ecosystem processes, and landscape heterogeneity related to non-production vegetation, at multiple spatial scales within agroecosystems.</p>

opencc-zeroMay 2020View details →
zenodo36/100

Sample Raw (PPG, Accel, and Gyro) and Processed (steps, calories, sleep, HR, HRV, SPO2, Respiratory Rate, R-R) data over 24 hours

<p>Over the course of 24 hours, we collected raw (Photoplethysmography (PPG), Acceleration, and Gyro) and processed (steps, calories, sleep, HR, HRV, SPO2, Respiratory Rate, R-R) data samples. Biostrap approaches health insights from a data-driven perspective. Our clinical-grade hardware enables users to accurately track SpO2, HRV, RHR, and a variety of other biometrics with confidence.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Dataset: Process Mining for Reliability Modeling of Smart Manufacturing Systems with Reduced Data Requirements

<p>Operational state logs from the Industry 4.0 Lab, University of Southern Denmark.</p> <p>&quot;I4.0Lab_state_log.csv&quot; -&gt; without failures</p> <p>I4.0Lab_state_log_failures.csv -&gt; with failures</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Code and data for publication "pyGRETA, pyCLARA, pyPRIMA: A pre-processing suite to generate flexible model regions for energy system models"

<p>This dataset contains the code of the&nbsp;three pre-processing tools&nbsp;<a href="https://github.com/tum-ens/pyGRETA">pyGRETA</a>,&nbsp; <a href="https://github.com/tum-ens/pyPRIMA">pyPRIMA</a>&nbsp;and&nbsp;<a href="https://github.com/tum-ens/pyCLARA">pyCLARA</a> and an examplary database for the scope of Austria.</p> <p>To run the code with full functionality additional data is needed. Check the documentation of the tools for further information.</p> <p>&nbsp;</p> <p>Sources for data can be found here:&nbsp;</p> <p>pyGRETA: https://pygreta.readthedocs.io/en/stable/user_manual.html#recommended-input-sources</p> <p>pyPRIMA: https://pyprima.readthedocs.io/en/stable/user_manual.html#recommended-input-sources</p> <p>pyCLARA: https://pyclara.readthedocs.io/en/stable/user_manual.html#recommended-input-sources</p>

opencc-by-4.0Sep 2020View details →
dryad36/100

Data from: Processing citizen science- and machine-annotated time-lapse imagery for biologically meaningful metrics

Time-lapse cameras facilitate remote and high-resolution monitoring of wild animal and plant communities, but the image data produced require further processing to be useful. Here we publish pipelines to process raw time-lapse imagery, resulting in count data (number of penguins per image) and 'nearest neighbour distance' measurements. The latter provide useful summaries of colony spatial structure (which can indicate phenological stage) and can be used to detect movement – metrics which could be valuable for a number of different monitoring scenarios, including image capture during aerial surveys. We present two alternative pathways for producing counts: 1) via the Zooniverse citizen science project Penguin Watch and 2) via a computer vision algorithm (Pengbot), and share a comparison of citizen science-, machine learning-, and expert- derived counts. We provide example files for 14 Penguin Watch cameras, generated from 63,070 raw images annotated by 50,445 volunteers. We encourage the use of this large open-source dataset, and the associated processing methodologies, for both ecological studies and continued machine learning and computer vision development.

opencc-zeroMar 2020View details →
zenodo36/100

Pre-Processed Cancer Multi-Omic Data from TCGA and Synthetic Data

<p><strong>ABSTRACT&nbsp;</strong></p> <p>It contains the data of four omic profiles (CNV, mRNA, miRNA, and protein) obtained for BRCA, LGG, and LUAD obtained from the TCGA project.&nbsp;</p> <p>In addition, we provide synthetic data for a mixture of isotropic distributions.</p> <p><strong>Instructions:&nbsp;</strong></p> <p>Cancer data are identified by cancer type (LGG: low-grade glioma, BRCA: breast cancer, and LUAD: lung cancer). The data are scaled by using the minima and maxima of each column so that the values are between 0 and 1. In these files, the columns are the features and the rows correspond to the patients.</p> <p>The summary data contains only the numerical values. The columns are the features and the rows are the observations.</p> <p><strong>Inspiration:</strong></p> <p>This dataset uploaded to U-BRITE for &quot;AI against CANCER DATA SCIENCE HACKATHON&quot;</p> <p>https://cancer.ubrite.org/hackathon-2021/</p> <p><strong>Acknowledgements</strong></p> <p>Diego Salazar, June 20, 2021, &quot;Pre-processed Cancer multi-omic data from TCGA and synthetic data&quot;, IEEE Dataport, doi: https://dx.doi.org/10.21227/pjb8-d090.</p> <p>https://ieee-dataport.org/documents/pre-processed-cancer-multi-omic-data-tcga-and-synthetic-data</p> <p><strong>U-BRITE last update date:</strong>&nbsp;07/21/2021</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Pre-processed IgH repertoire sequencing data from BioProject PRJNA748239

<p><strong>Data Processing</strong></p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p>software_versions&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p>quality_thresholds&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p>paired_reads_assembly&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p>primer_match_cutoffs&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p>consensus_building&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p>collapsing_method&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;IMGT</p> <p>&nbsp;</p> <p>Format</p> <p>&nbsp;</p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>C_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Isotype subclass</p> <p><strong>SEQUENCE_ID</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Sequence identifier</p> <p><strong>V_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V segment gene and allele</p> <p><strong>D_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;D segment gene and allele</p> <p><strong>J_CALL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;J segment gene and allele</p> <p><strong>JUNCTION_LENGTH</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Junction length</p> <p><strong>CONSCOUNT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;UMI count for the given unique sequence</p> <p><strong>ISOTYPE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Constant region primer (isotype)</p> <p><strong>MU_COUNT_CDR_R</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of replacement mutations in CDR region</p> <p><strong>MU_COUNT_CDR_S</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of silent mutations in CDR region</p> <p><strong>MU_COUNT_FWR_R</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of replacement mutations in FWR region</p> <p><strong>MU_COUNT_FWR_S</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Number of silent mutations in FWR region</p> <p><strong>MUT_TOTAL</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Total number of mutations in V gene&nbsp;</p> <p><strong>NP_LENGTH</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Total number of N and P additions<strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong></p> <p><strong>SEQUENCE_INPUT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Full length sequence</p> <p><strong>SEQUENCE_IMGT</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Gapped IMGT sequence</p> <p><strong>V_GERM_START_VDJ</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;position of the first nucleotide in ungapped V germline sequence alignment</p> <p><strong>JUNCTION</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Junction nucleotide sequence</p> <p><strong>GERMLINE_IMGT_D_MASK</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;IMGT-gapped germline nucleotide sequence with ns masking the NP1-D-NP2 regions</p> <p><strong>CDR3_AA_GRAVY</strong>&nbsp; &nbsp; &nbsp; &nbsp;CDR3 hydrophobicity</p> <p><strong>CDR3_AA_BULK</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; CDR3 bulkiness</p> <p><strong>CDR3_AA_ALIPHATIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Normalized aliphatic index</p> <p><strong>CDR3_AA_POLARITY</strong>&nbsp; &nbsp; &nbsp; &nbsp; CDR3 polarity</p> <p><strong>CDR3_AA_CHARGE</strong>&nbsp; &nbsp; &nbsp; &nbsp; normalised net&nbsp;charge</p> <p><strong>CDR3_AA_BASIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Basic side chain residue content</p> <p><strong>CDR3_AA_ACIDIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Acidic side chain residue content</p> <p><strong>CDR3_AA_AROMATIC</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;aromatic side chain conten</p> <p><strong>Subset</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Defined B cell subset&nbsp;</p> <p><strong>Repertoire</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;R/S ratio in CDR region</p> <p><strong>R_SFWR</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;R/S ratio in FWR region</p> <p><strong>V_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V segment gene</p> <p><strong>D_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;D segment gene</p> <p><strong>J_GENE</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;J segment gene</p> <p><strong>V_FAM</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;V family gene</p> <p><strong>Run</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;ID of sequencing run</p> <p><strong>Sex</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Sex of the Subject</p> <p><strong>Age</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Age of the subject</p> <p><strong>UNIQUE_ID</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;Subject identifier&nbsp;</p> <p><strong>SAMPLE</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Sample identifier, linking back to raw data</p> <p><strong>Bcellno</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Number of input B cells</p> <p><strong>Cells</strong>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Cell type&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Toy datasets for GTN materials about Metabolomics data processing

<p>These are the datasets needed for the GTN tutorial about Data processing of Metabolomics MS-based data.</p> <p>The dataMatrix and corresponding sampleMetadata files are simulated data, specially created from scratch for the GTN tutorial.</p> <p>The variableMetadata file&#39;s columns of m/z and rt characteristics have been filled with values picked from a table obtained by an extraction process of some untargetted MS raw files. MS raw files used are from https://zenodo.org/record/3757956; extraction was performed using XCMS with parameters that might not be optimal.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Processed data for the "Property-Based Testing of Web APIs" paper

<p>Processed data for the &quot;Property-Based Testing of Web APIs&quot; paper.&nbsp; Each directory in the archive consists of:</p> <p>- metadata.json. Metadata about a test run - tested fuzzer name, run duration, etc</p> <p>- fuzzer.json&nbsp;- Structured fuzzer output</p> <p>-&nbsp;deduplicated_cases.json - Deduplicated reported failures, when fuzzers provide it</p> <p>- sentry.json&nbsp;- Cleaned Sentry events for this run</p> <p>- target.json&nbsp;- Parsed stdout for Gitlab &amp; Disease.sh targets that were tested without Sentry integration</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

16S processed data and shell processing scripts for "Epithelial-myeloid exchange of MHCII constrains immunity and microbiota composition"

<p>Main repo for MHCII on IECs 16S processing code.</p>

openmit-licenseSep 2021View details →
zenodo36/100

Pre-processed IgH receptor repertoire data from MS patients after aHSCT from BioProject PRJNA763367

<p><strong>Data Processing</strong></p> <p>Samples were demultiplexed via their Illumina indices, and processed using the Immcantation toolkit(1,2).&nbsp;Raw fastq files were filtered based on a quality score threshold of 20. Paired reads were joined if they had a minimum length of 10 nt, maximum error rate of 0.3 and a significance threshold of 0.0001. Reads with identical UMI were collapsed to a consensus sequence. Reads with identical full-length sequence and identical constant primer but differing UMI were further collapsed. Sequences were then submitted to IgBlast (3) for VDJ assignment and sequence annotation. Constant region sequences were mapped to germline using Stampy(4). The number and type of V gene mutations was calculated using the shazam R package.(2)</p> <p>&nbsp;</p> <p><strong>software_versions</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;pRESTO:0.5.3,Change-O:0.3.4,IgBlast 1.6.1, stampy1.0.21. shazam0.1.8</p> <p><strong>quality_thresholds</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;FilterSeq.py pRESTO Q&gt;20</p> <p><strong>paired_reads_assembly</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;AssemblePairs.py pRESTO minlen 10 maxerror 0.3 alpha 0.0001</p> <p><strong>primer_match_cutoffs</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;MaskPrimers.py pRESTO C primer &amp; V primer maxerror 0.2</p> <p><strong>consensus_building</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;BuildConsensus.py pRESTO maxerror 0.1 maxgap 0.5</p> <p><strong>collapsing_method</strong>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;CollapseSeq.py pRESTO</p> <p><strong>germline_database&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>IMGT</p> <p>&nbsp;</p> <p><strong>Format</strong></p> <p>Processed sequences are provided in a tab delimited file format, including the following annotations:</p> <p>&nbsp;</p> <p><strong>ISOTYPE_SUBCLASS &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Isotype subclass</p> <p><strong>SEQUENCE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sequence identifier</p> <p><strong>JUNCTION_LENGTH&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction length</p> <p><strong>CONSCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Raw read count from which UMI consensus sequences were generated, summed over all UMIs for the given unique sequence.</p> <p><strong>DUPCOUNT&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>UMI count for the given unique sequence</p> <p><strong>ISOTYPE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Constant region primer (isotype)</p> <p><strong>MUT_TOTAL&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Total number of mutations in V gene&nbsp;</p> <p><strong>SAMPLE&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</strong>Sample identifier, linking back to raw data</p> <p><strong>JUNCTION&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Junction nucleotide sequence</p> <p><strong>Protein_seq &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Amino acid sequence</p> <p><strong>CDR3_AA_GRAVY&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>CDR3 hydrophobicity index</p> <p><strong>CDR3_AA_BULK &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 bulkiness</p> <p><strong>CDR3_AA_ALIPHATIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 aliphatic index</p> <p><strong>CDR3_AA_POLARITY &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 polarity</p> <p><strong>CDR3_AA_CHARGE &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 normalized net charge</p> <p><strong>CDR3_AA_BASIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 basic side chain residue content</p> <p><strong>CDR3_AA_ACIDIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 acidic side chain residue content</p> <p><strong>CDR3_AA_AROMATIC &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>CDR3 aromatic side chain content</p> <p><strong>Subset&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell subset&nbsp;</p> <p><strong>Repertoire&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Defined B cell repertoire (Naive, Memory IgM/IgD, IgA, IgG)</p> <p><strong>R_SCDR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in CDR region</p> <p><strong>R_SFWR&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>R/S ratio in FWR region</p> <p><strong>V_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V segment gene</p> <p><strong>D_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>D segment gene</p> <p><strong>J_GENE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>J segment gene</p> <p><strong>V_FAM&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>V family gene</p> <p><strong>Clust_REPRES&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster representative</p> <p><strong>Clust_SIZE&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Cluster size</p> <p><strong>Sex&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sex of the Subject</p> <p><strong>UNIQUE_ID&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;</strong>Sample identifier&nbsp;</p> <p><strong>Bcellno &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Input B cell number</p> <p><strong>Days_posttx &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Sampling time point relative to transplantation</p> <p><strong>Age_at_tx &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Age of the subject (at aHSCT)</p> <p><strong>Disease &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>MS subtype</p> <p><strong>Last_therapy &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Last therapy prior to aHSCT</p> <p><strong>Disease_duration &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Disease duration</p> <p><strong>CMV_reactivation &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Cytomegalovirus reactivation</p> <p><strong>Month_label &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Month post-aHSCT inverval bin</p> <p><strong>Patient_label &nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp; </strong>Subject identifier</p> <pre> &nbsp;</pre> <p><strong>References</strong></p> <p>1.&nbsp;Vander Heiden, J. A., G. Yaari, M. Uduman, J. N. H. Stern, K. C. O&rsquo;Connor, D. A. Hafler, F. Vigneault, and S. H. Kleinstein.&nbsp;2014. PRESTO: A toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires.&nbsp;<em>Bioinformatics</em>30: 1930&ndash;1932.</p> <p>2. Gupta, N. T., J. A. Vander Heiden, M. Uduman, D. Gadala-Maria, G. Yaari, and S. H. Kleinstein.&nbsp;2015. Change-O: A toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data.&nbsp;<em>Bioinformatics</em>31: 3356&ndash;3358.</p> <p>3. Ye, J., N. Ma, T. L. Madden, and J. M. Ostell. 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool.&nbsp;<em>Nucleic Acids Res.</em>41.</p> <p>4. Lunter, G., and M. Goodson. 2011. Stampy: A statistical algorithm for sensitive and fast mapping of Illumina sequence reads.&nbsp;<em>Genome Res.</em>21: 936&ndash;939.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Processed data from Repair-seq screens of double-strand breaks

<p>Processed data from Repair-seq screens of double-strand breaks induced&nbsp;by SpCas9 and AsCas12a in the presence or absence of oligonucleotide homology donors.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Scripts for post-processing Delft3d output data and figures for manuscript 'Longitudinal scour-bar pattern in estuaries'

<p>The 7z file contains two folders, one named &#39;mat&#39; contains the matlab scripts for post-processing Delft3D output data and plotting, the other named &#39;Figures&#39; contains main outputs for the manuscript &#39;Longitudinal scour-bar pattern in estuaries&#39;.&nbsp;</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record