Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
357
datasets available to search
ShareScore release 0.9.0
Dataset results
357 results for “supplementary information”
Supplementary Information for the manuscript: "Gene expression evolution is predicted by stronger selection at more pleiotropic genes"
<p>This repository contains the supplementary information for the manuscript "<em>Gene expression evolution is predicted by stronger selection at more pleiotropic genes</em>" (<a href="https://doi.org/10.1101/2024.07.22.604294">https://doi.org/10.1101/2024.07.22.604294</a>).</p> <p>The supplementary data "data.tar.gz" is related to the Github repository <a href="https://github.com/charlesrocabert/Koch-et-al-Gene-expression-evolution-is-predictable-and-driven-by-indirect-selection-pressures">https://github.com/charlesrocabert/Koch-et-al-Predictability-of-Gene-Expression</a>.</p> <h2>Content of the repository</h2> <ul> <li><strong>Script S1. </strong>BSFG estimates (50.9 MB).</li> <li><strong>Script S2.</strong> WGCNA analysis and results G1 (99.6 MB).</li> <li><strong>Data S1.</strong> Datasets resulting from the global genomics analysis of the output of the transcriptomics pipeline (5.9 MB).</li> <li><strong>Data S2.</strong> VCF file containing the 566,296 quality-checked SNPs (1.1 GB).</li> <li><strong>Data S3.</strong> VCF file containing the 358,142 SNPs with a call rate higher or equal to 50% (964.3 MB).</li> <li><strong>Data S4.</strong> VCF file containing the imputed genotypes (964.3 MB).</li> <li><strong>Data S5.</strong> Results of the imputation tests (479.5 kB).</li> <li><strong>Data S6.</strong> VCF file containing imputed genotypes where SNPs with a minor allele frequency lower than 0.05 have been filtered out (MAF $\leq$ 0.05) (176.0 MB).</li> <li><strong>Data S7.</strong> List of the 1,273 significant eQTL associations (for gene expression levels or relative fitness as phenotypes) (61.2 kB).</li> <li><strong>Data S8.</strong> List of allele frequency changes (AFCs) for every markers in HD environment, for lines L1, L2, L3, L5, L6, Mx1 and Mx2 (32.4 MB).</li> <li><strong>Data S9.</strong> List of all SNPs indicating if their AFC is significantly higher in each line and their degree of parallelism (2.4 MB).</li> <li><strong>Data S10.</strong> Excel file containing the results of the gene functional enrichment analysis of the hub and eQTL carrier genes (33.7 KB).</li> <li><strong>data.tar.gz.</strong> Dataset mandatory to re-run the genomics analysis (see <a href="https://github.com/charlesrocabert/Koch-et-al-Gene-expression-evolution-is-predictable-and-driven-by-indirect-selection-pressures">https://github.com/charlesrocabert/Koch-et-al-Predictability-of-Gene-Expression</a>) (7.8 GB).</li> </ul> <p> </p>
Supplementary Information: Video data files: Gray et al. Caught on camera: Ocelot, Leopardus pardalis (Mammalia: Felidae) predation on foam nests of Savage's thin-toed frog, Leptodactylus savagei (Amphibia: Leptodactylidae)
<p>Supplementary Information: Original camera trap footage for Gray, Ibáñez, Barrios & Potvin: Caught on camera: Ocelot, <i>Leopardus pardalis </i>(Mammalia: Felidae) predation on foam nests of Savage's thin-toed frog, <i>Leptodactylus savagei</i> (Amphibia: Leptodactylidae). </p><p> </p>
Supplementary Information BeXY
<p>These are all the input files that were used in the publication.</p><p>S1 Table: contains a list of all samples, their respective publication and ENA project number analyzed in this paper (there are 4 sheets for whole-genome sequencing of ancient humans (WGS), high-depth subset of ancient WGS, 1240k target enrichment capture samples of ancient humans, and modern human WGS from the Simons Genome Diversity Project).</p><p>Capture1240k: contains the counts and the number of targets per scaffold for all ancient humans sequenced with 1240k target enrichment capture sequencing (see S1 table)</p><p>downsampling_aneuploids: contains all downsampled counts for simulated aneuploid invididuals as well as scripts to generate those. Also contains our implementation of the method seGMM.</p><p>downsampling_trisomy: contains all downsampled counts for simulated invididuals with trisomy 21 (factor21_1.5) and without trisomy 21 (factor_21_1) as well as scripts to generate those.</p><p>downsampling_euploids: contains all downsampled counts for euploid individuals as well as scripts to generate those.</p><p>lowQualityReference: contains all downsampled counts for the simulated low-quality reference genome assembly, as well as the scripts that were used to generate such an assembly based on the human reference genome.</p><p>WGS_ancient: contains the counts and the chromosome lengths for all ancient humans sequenced with whole-genome shotgun sequencing (see S1 table)</p><p>WGS_ancient_samples_lt_2x: contains the counts and the chromosome lengths for the 116 high-depth WGS samples (see S1 table)</p><p>WGS_modern_SGDP: contains the counts and the chromosome lengths for the 276 modern human samples downloaded from the Simons Genome Diversity Project (see S1 table). wgs_modern_SGDP_original_counts_withoutSeqTypes.txt corresponds to the counts obtained from the downloaded CRAM files, without considering differences in sequencing types. wgs_modern_SGDP_original_counts_withSeqTypes.txt corresponds to the same counts, but the sequencing type per sample is specified in the second column. wgs_modern_SGDP_filteredMQ30_counts.txt corresponds to the counts obtained by filtering on a mapping quality of 30.</p><p>non_model_organisms_posterior_probabilities_t: contains the posterior probabilities for each scaffold to be autosomal, Y-linked, X-linked or different as inferred by BeXY, for each of the six non-model organism species published in Nursyifa et al. 2021 ( <a href="https://doi.org/10.1111/1755-0998.13491">https://doi.org/10.1111/1755-0998.13491</a>).</p><p> </p>
The datasets used in the Supplementary Information
<p>The datasets used in the Supplementary Information. For additional details, please refer to the ReadMe file within each subfolder.</p>
Supplementary material 1 from: Sosa-López JR, Díaz Bernal NN, Padilla E, Briones-Salas M (2023) Analysis of the effects of habitat characteristics, human disturbance and prey on felids presence using long-term community monitoring information. Nature Conservation 53: 279-295. https://doi.org/10.3897/natureconservation.53.104135
Sampling sites and dates on which the camera-traps were installed and Generalized linear mixed models (GLMM) for Puma, Bobcat and Margay
Supplementary Information: RAPP-containing arrest peptides induce translational stalling by short circuiting the ribosomal peptidyltransferase activity
<p><span><span>The Simulation_files folder contains initial coordinates, input files and output coordinates of the Molecular Dynamics (MD) simulations of wild type and mutants of the Apdp peptide in the ribosome ( region of 35 </span></span>Å around the peptide).</p> <p><span><span>The Uncharged_N-terminal_residues folder includes residue topologies (.rtp file) and hydrogen data bank (.hdb) of the uncharged terminal Alanine and Proline (used for modelling Ala134 and Pro134).</span></span></p> <p><span><span>The Figures folder contains rmsd, distances, rmsf values, and projections on the most dominant conformational modes sampled by the Ala 132 backbone atoms obtained from the MD trajectories.</span></span><span><br></span></p>
Supplementary Information: CHAPTER 3 - Classification of genomic features of plant-associated bacteria using machine learning
<p>Appendix A- List of all bacterial genomes used in orthologous genes clustering in the feature extraction step and in the further steps to build and test classifiers’ models. The list includes the isolation source information and the related category for the genome classification and features selection purposes.</p> <p>Appendix B - Distribution of genomes by phylum, family, and genus among the categories defined according to bacteria lifestyle association.</p> <p>Appendix C - Enriched orthogroups by genus according to each enrichment test (Material and Methods). Values for each test are "Y" (enriched), "N" (not enriched), or "Untested" (clusters were untested when there was insufficient phylogenetic signal, they were too small or were found in all genomes).</p> <p>Appendix D - Classification performance of random forest and logistic regression techniques applied to genus-specific datasets of genomic features (orthogroups) using both matrices from gene count number and presence/absence values. Sensitivity is a measure of how well a test identifies true positives; Specificity: is a measure how well a test or model avoids false positives; Positive Predictive Value (Pos. Pred. Value): The probability that a positive prediction is correct; Negative Predictive Value (Neg. Pred. Value): The probability that a negative prediction is correct; Precision: The accuracy of positive predictions; Recall (Sensitivity): The ability to find all relevant cases; F1 Score: A combined measure of precision and recall; Prevalence: The proportion of positive cases in the total; Detection Rate: The proportion of true positive cases identified; Detection Prevalence: The proportion of positive predictions; Balanced Accuracy: An average of sensitivity and specificity; Area Under the Curve (AUC): The overall performance of the model in distinguishing between positive and negative cases.</p> <p>Appendix E - Orthogroups assigned with predicted COGs as an important feature for classifying plant-associated genomes. COG categories: A - RNA processing and modification; B - Chromatin structure and dynamics; C - Energy production and conversion; D - Cell cycle control, cell division, chromosome partitioning; E - Amino acid transport and metabolism; F - Nucleotide transport and metabolism; G - Carbohydrate transport and metabolism; H - Coenzyme transport and metabolism; I - Lipid transport and metabolism; J - Translation, ribosomal structure and biogenesis; K - Transcription; L - Replication, recombination and repair; M - Cell wall/membrane/envelope biogenesis; N - Cell motility; O - Posttranslational modification, protein turnover, chaperones; P - Inorganic ion transport and metabolism; Q - Secondary metabolites biosynthesis, transport and catabolism; R - General function prediction only; S - Function unknown; T - Signal transduction mechanisms; U - Intracellular trafficking, secretion, and vesicular transport; V - Defense mechanisms; W - Extracellular structures; X - Mobilome: prophages, transposons; Y - Nuclear structure; Z - Cytoskeleton.</p>
Supplementary information
<p><strong>Supplementary information</strong></p> <p> </p> <p>The changes between version 1 to version 2 are described in README file.</p>
Supplementary Information "Key drivers and pressures of global water scarcity hotspots"
<p>Supplementary information of "Key drivers and pressures of global water scarcity hotspots".</p> <ul> <li>Supplementary table A - SCOPUS search strings</li> <li>Supplementary table B - DPSIR definitions</li> <li>Supplementary table C - Data resources</li> <li>Supplementary figures D - Lineplots of historical data analysis</li> <li>Supplementary table E - DPSIR case study results</li> <li>Supplementary figures F - Circular barplots DPSIR analysis per hotspot</li> </ul>
Supplementary Information of Multistage Provincial Power Expansion Planning and Market Design toward Energy Transition: A Case Study of Xinjiang in China
<p>Supplementary Information of Multistage Provincial Power Expansion Planning and Market Design toward Energy Transition: A Case Study of Xinjiang in China</p>
Supplementary Material for the Paper: "Mining Architectural Information: A Systematic Mapping Study"
<p>The supplementary material is composed of three files:</p> <p><strong>1. SelectedStudies.pdf</strong><br>contains the 104 studies selected in this systematic mapping study.</p> <p><strong>2. PublicationVenues.pdf</strong></p> <p>contains the major publication venues of the selected studies.</p> <p><strong>3. Dataset.xlsx</strong></p> <p>contains the following Excel sheets:</p> <p><em><strong>(1) Architecting Activities</strong></em></p> <p>contains the 11 architecting activities, mined architectural information, datasets and sources used in the selected studies, and relevant studies.</p> <p><em><strong>(2) Automatic Approaches</strong></em></p> <p>contains tasks, automatic approaches, used sources, mined architectural information, and relevant studies.</p> <p><em><strong>(3) Semi-automatic Approaches</strong></em></p> <p>contains tasks, semi-automatic approaches, used sources, mined architectural information, and relevant studies.</p> <p><em><strong>(4) Manual Approaches</strong></em></p> <p>contains tasks, manual approaches, used sources, mined architectural information, and relevant studies.</p> <p><em><strong>(5) General Tools </strong></em></p> <p>contains general tool names, description of each tool, URL of each tool (if provided), and relevant studies.</p> <p><em><strong>(6) Automatic Tools</strong></em></p> <p>contains tasks, automatic tool names, used sources, mined architectural information, URL of each tool (if provided), and relevant studies.</p> <p><em><strong>(7) Semi-automatic Tools</strong></em></p> <p>contains tasks, semi-automatic tool, used sources, mined architectural information, URL of each tool (if provided), and relevant studies.</p>
Supplementary Information: Multimodal binding and inhibition of bacterial ribosomes by the 2 antimicrobial peptides Api137 and Api88
<div> <div>This dataset contains important data files for the MD simulation that are part of this publication.</div> <div> </div> <div>The "simulations" directory contains Gromacs parameter files (.mdp) and the run input files (.tpr) as well as the final coordinate files (.gro) of each individual production simulation.</div> <div> </div> <div>The directory "figure3" contains the raw data used to create Figure 3 in the manuscript.</div> <div> </div> <div>The subdirectory "a" contains the data for the PCA projection plot in subfigure 3a. It includes projections of the simulation ensembles of Api88 conformation I-III on to the two dominant conformational modes (.xvg) and the respective extreme conformations (.pdb). The projections of the three initial models and the optimized structure set are also included.</div> <div> </div> <div>Subdirectory "b" contains a numpy array (.npy) with the data for the correlation heatmap in subfigure 3b.</div> <div> </div> <div>Subdirectory "c" contains the results of several correlation-optimization searches. Each directory "N#_maps", where # is to be replaced by the number of structures in the set, contains the search results for N correlation-optimized structures in the Api88 trajectories in the form of a pickled python dictionary (state.pkl). The dictionary has the following keys:</div> <ul> <li>used: Already used sets of MD structures (frozenset)</li> <li>selection: Structure set selected in the last iteration (set)</li> <li>weights: weights of each structure in the selected structure set (numpy array)</li> <li>iteration: Counter of the last iteration (int)</li> </ul> <div> </div> <div>The directory "supplentary_figure_correlation_time" contains the data for a plot of the optimized correlation coefficient as a function of simulation time. The results of the optimization algorithms (as pickled python objects) are included in the subdirectories with the associated trajectory length as a name.</div> </div> <p> </p>
Supplementary Information for 'Smoke in your eyes…'
<p>The file contains pdf version of Numbers spreadsheets used for data analysis in the paper 'Smoke in your eyes: An investigation of the effects of wind power on weather trends and climate using time series analysis'.</p> <p>The supplementary data have been updated to include recalculation of the CCFs using a Python script incorporating standard function np.correlate. Errors in sheets 3-5-2 and 3-5-3 have been corrected, leading to correction of fig. 16 in the main text.</p>
Supplementary information provided with Murray et al.: Discovery of an Antarctic ascidian-associated uncultivated Verrucomicrobia with antimelanoma palmerolide biosynthetic potential
<p><span>The Antarctic marine ecosystem harbors a wealth of biological and chemical innovation that has risen in concert over millennia since the isolation of the continent and formation of the Antarctic circumpolar current. Scientific inquiry into the novelty of marine natural products produced by Antarctic benthic invertebrates led to the discovery of a bioactive macrolide, palmerolide A, that has specific activity against melanoma and holds considerable promise as an anticancer therapeutic. While this compound was isolated from the Antarctic ascidian <i>Synoicum adareanum</i>, its biosynthesis has since been hypothesized to be microbially mediated, given structural similarities to microbially-produced hybrid non-ribosomal peptide-polyketide macrolides. Here, we describe a metagenome-enabled investigation aimed at identifying the biosynthetic gene cluster (BGC) and palmerolide A-producing organism. A 74 Kbp candidate BGC encoding the multi-modular enzymatic machinery (hybrid Type I-<i>trans</i>-AT polyketide synthase-non-ribosomal peptide synthetase and tailoring functional domains) was identified and found to harbor key features predicted as necessary for palmerolide A biosynthesis. Surveys of ascidian microbiome samples targeting the candidate BGC revealed a high correlation between palmerolide-gene targets and a single 16S rRNA gene variant (R=0.83 – 0.99). Through repeated rounds of metagenome sequencing followed by binning contigs into metagenome-assembled genomes, we were able to retrieve a near-complete genome (10 contigs) of the BGC-producing organism, a novel verrucomicrobium within the <i>Opitutaceae</i> family that we propose here as <i>Candidatus</i> Synoicihabitans palmerolidicus. The refined genome assembly harbors five highly similar BGC copies, along with structural and functional features that shed light on the host-associated nature of this unique bacterium.</span></p>
Supplementary information for Paleoceanographic changes in the late Pliocene promoted rapid diversification in pelagic seabirds
<p><b>Aim: </b>Paleoceanographic changes can act as drivers of diversification and speciation, even in highly mobile marine organisms. Shearwaters are a group of globally distributed and highly mobile pelagic seabirds. Despite a recent well resolved phylogeny, shearwaters have controversial species limits, and show periods of both slow and rapid diversification. Here, we explore the role of paleoceanographic changes on the diversification and speciation in these highly mobile pelagic seabirds. We investigate shearwater biogeography and the evolution of a key phenotypic trait, body size, and we assess the validity of the current taxonomy of the group.</p> <p><b>Location:</b> Worldwide.</p> <p><b>Taxa:</b> Shearwaters (Order Procellariiformes, Family Procellariidae, Genera <i>Ardenna</i>, <i>Calonectris</i> and <i>Puffinus</i>).</p> <p><b>Methods: </b>We generated genomic data (double-digest restriction site-associated DNA) for almost all extant shearwater species to infer a time-calibrated species tree. We estimated ancestral ranges and evaluated the roles of founder events, vicariance and surface ocean currents in driving shearwater diversification. We performed phylogenetic generalized least squares to identify potential predictors of variability in body size along the phylogeny. To assess the validity of the current taxonomy of the group, we analysed genomic patterns of recent shared ancestry and differentiation among shearwater taxa.</p> <p><b>Results:</b> We identified a period of high dispersal and rapid speciation during the Late Pliocene - early Pleistocene. Species dispersal appears to be favoured by surface ocean currents, and biogeographic reconstructions support founder events as the main mode of speciation in these highly mobile pelagic seabirds. Body mass shows significant associations with life strategies and local conditions. The current taxonomy shows some incongruences with the patterns of genomic divergence.</p> <p><b>Main Conclusions: </b>A reduction of neritic areas during the Pliocene seems to have driven global extinctions of shearwater species, followed by a subsequent burst of speciation and dispersal probably promoted by Plio-Pleistocene climatic shifts. Our findings extend our understanding on the drivers of speciation and dispersal of highly mobile pelagic seabirds and shed new light on the important role of paleoceanographic events.</p>
Hölzchen et al .(2022) Supplement A. Supplementary information
<p>This dataset contains Supplement A (supplementary information) for Hölzchen et al. (2022) "Estimating crossing success of human agents across sea straits out of Africa in the Late Pleistocene".</p>
Opportunities for seasonal forecasting to support water management outside the tropics: Supplementary information
<p>Supporting information for Jackson-Blake et al. (2022), 'Opportunities for seasonal forecasting to support water management outside the tropics', HESS</p> <p>Contents (section numbers refer to sections in the HESS paper):</p> <ul> <li>SI1_CatchmentMaps.pdf: Catchment maps for each of the five case study sites.</li> <li>SI2_FirstSurvey_HistoricEvent.pdf: Questions asked as part of the first assessment exercise described in Section 3. In this exercise, stakeholders were asked to pick a historic season when there were problems in their catchment/lake, and then seasonal forecasts were produced for that season. Stakeholders were then asked a set of questions designed to assess: (1) their interpretation of the forecasts, and (2) to what extent they thought forecasts would have been useful for them. Full questions and answers are provided in this pdf.</li> <li>SI3_WindowsOfOpportunity_Responses.xlsx: Full responses to the questionnaire described in Sections 2.2 and 3.2.</li> </ul>
Supplementary Information of Amplicon-based nanopore sequencing of patients with COVID-19 omicron (B.1.1.529) variant from India
<p><strong>We report sequencing of omicron variants from SARS-CoV-2 in 75 patients, using Nanopore long-read sequencing chemistry. We highlight the nature of mutations in spike glycoprotein that are unique and common to other populations.</strong></p> <p> </p>
Supplementary information for the paper: "Barriers and drivers to implement innovative business models towards sustainable urban mobility"
<p>The dataset presents excerpts from the interviews used as main source of empirical evidence in the coding process for the identification of barriers and solutions adopted by the companies in the paper: “Barriers and drivers to implement innovative business models towards sustainable urban mobility”.</p>
Supplementary Information and Data for The Effect of Compositional Heterogeneity on the Martensite Start Temperature of a High Strength Steel During Rapid Austenitisation and Cooling
<p>The supplementary information and supporting data for "The Effect of Compositional Heterogeneity on the Martensite Start Temperature of a High Strength Steel During Rapid Austenitisation and Cooling" in the proceedings of the 42nd Risø International Symposium on Materials Science. The images used in the paper are also included. The files for the montaged EBSD datasets are not included but can be made available upon request. Contact the corresponding author: mark.taylor-5@manchester.ac.uk</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.