Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Genetic data and underlying taxa and GenBank sources of diatoms used in phylogenetic analysis for the diatom genus Nupela
<p>Supplementary material for the manuscript: Kulikovskiy M., Maltsev Y., Glushchenko A., Gusev E., Kapustin D., Kuznetsova I., Kociolek J.P. Preliminary molecular phylogeny of the diatom genus <em>Nupela</em> with the description of a new species and consideration of the interrelationships of taxa in the suborder Neidiineae D.G. Mann sensu E.J. Cox. Fottea</p> <p>Molecular investigation of diatom genera <em>Nupela</em> and <em>Brachysira</em> is conducted using strains from Indonesia and Vietnam. New species from the genus <em>Nupela indonesica</em> sp. nov. is described using combined approach. <em>Nupela lesothensis</em> (Schoeman) Lange-Bertalot is investigated using molecular data too. Phylogenetic analysis shows that <em>Nupela</em> and <em>Brachysira</em> are not closest genera. Morphology of <em>Nupela</em> and it differences from <em>Brachysira</em> is discussed. The genus <em>Nupela</em> is differs from all other diatom taxa by having coalescent hymenes ouside of areolae but not inside. Facultative development of raphe between different <em>Nupela</em> species is discussed.<br> SUPPLEMENT S1. Taxa and DNA sequence data used in phylogenetic analysis.<br> SUPPLEMENT S2. Final alignment of 2-gene DNA sequence data used for phylogenetic analysis in FASTA format.<br> SUPPLEMENT S3. Maximum Likelihood tree of <em>Nupela</em> species (indicated in bold) constructed from a concatenated alignment of 163 partial rbcL and partial 18S rDNA sequences of 1806 characters. Values near the horizontal lines (slash) are bootstrap support from RAxML analyses (<50 are not shown). Species from the centric diatoms were used as an outgroup. Families indicated according COX (2015).</p>
COVID 19 SARS COV2 targets and small molecule data including insilico analysis
<p>Welcome to the repository for the COVID-19 research data.</p> <p>Corresponding Author: Girinath G. Pillai and few experts</p> <p>Co-authors: Team of experts, scholars and students</p> <p>To join dedicated Slack Discussion : <a href="https://join.slack.com/t/nyroindia/shared_invite/zt-ejes216c-QZzEK_G5tNKIjewbVj2IPA">https://join.slack.com/t/nyroindia/shared_invite/zt-ejes216c-QZzEK_G5tNKIjewbVj2IPA</a></p> <p>We commit to conduct research analysis and all the findings and data will be open and anyone can use or help us improve the data.</p> <p>The parameters for checkpoints are:</p> <p>A) Pharmacophore Modelling - i) generate pharmacophore reference maps from XRay crystal geometry, ii) Generate all possible conformers of the dataset molecules for screening.</p> <p>B) Virtual Screening - i) highest docking score within the dataset, ii) lowest clashes (interligand or intraligand), iii) interactions with key amino acid residues based on literature reports, PROSITE server and pocket finding algorithm like DoGSite or CASTp iv) optimal LE values and v) satisfactory interactions between small molecules and amino acids.</p> <p>C) Selection - i) binding affinity range predictions, lowest among the dataset, ii) free binding energy calculation considering desolvation terms, lowest among the dataset and iii) torsion analysis - coverage of bonds in CSD database.</p> <p>D) Optimization - i) pharmacokinetic properties to be generated from selected hits and an optimal balance of properties to be considered for candidate selection criteria.</p> <p>E) For novel lead molecules - i) chemical space exploration on building blocks could be carried out, ii) on-demand synthesis and procurement.</p> <p>If you prefer you could always cite <a href="https://github.com/giribio/COVID19">https://github.com/giribio/COVID19</a></p> <p>Feel free to create any issues in Github or feel free to contact me via Slack for any queries.</p> <p>Thanks and let us fight against COVID-19 in all possible ways.</p>
R code and data to reproduce figures from the "Multivariate autoregressive modelling and conditional simulation for temporal uncertainty analysis of an urban water system in Luxembourg" paper
<p>This repository contains the R code and data to reproduce figures from the "Multivariate autoregressive modelling and conditional simulation for temporal uncertainty analysis of an urban water system in Luxembourg" paper.</p>
Archäologische Chronologie und historische Interpretation: Die Merowingerzeit in Süddeutschland (Correspondence Analysis Data Set)
<p>This data set is a supplement to the book "Archäologische Chronologie und historische Interpretation: Die Merowingerzeit in Süddeutschland" (De Gruyter, 2016) and comprises the archaeological data and the results of the correspondence analysis of Merovingian-period graves from southern Germany and their chronological classification. The data sets for female and male burials can be downloaded as PDF, EXCEL and CSV files.</p>
Bibliographic data and analysis of COVID-19 research outputs from Imperial College London 16.01.2020-02.04.2020
<p>Bibliographic data and analysis of 41 research outputs, including reports/preprints/published articles/code, identified as having Imperial authorship and being relevant to COVID-19, published between 16.01.2020 - 02.04.2020. </p> <p>Related report can be found at: Price RC and Ozkan YA. 13 weeks in a pandemic: a descriptive study of Imperial College London’s COVID-19 publications. Imperial College London (April 2020), https://doi.org/10.25561/77970</p>
Data from: Arm waving in stylophoran echinoderms: three-dimensional mobility analysis illuminates cornute locomotion
<p>The locomotion strategies of fossil invertebrates are typically interpreted on the basis of morphological descriptions. However, it has been shown that homologous structures with disparate morphologies in extant invertebrates do not necessarily correlate with differences in their locomotory capability. Here, we present a new methodology for analysing locomotion in fossil invertebrates with a rigid skeleton through an investigation of a cornute stylophoran, an extinct fossil echinoderm with enigmatic morphology that has made its mode of locomotion difficult to reconstruct. We determined the range of motion of a stylophoran arm based on digitized three-dimensional morphology of an early Ordovician form, <i>Phyllocystis crassimarginata</i>. Our analysis showed that efficient arm-forward epifaunal locomotion based on dorsoventral movements, as previously hypothesized for cornute stylophorans, was not possible for this taxon; locomotion driven primarily by lateral movement of the proximal aulacophore was more likely. 3D digital modelling provides an objective and rigorous methodology for illuminating the movement capabilities and locomotion strategies of fossil invertebrates.</p>
Data for "Tuning parameters of dimensionality reduction methods for single-cell RNA-seq analysis"
<p>The files named <code>df_scran.csv</code>, <code>df_seurat.csv</code>, <code>df_zinbwave.csv</code>, <code>df_dca.csv</code>, and <code>df_scvi.csv</code> contain one row per configuration that we ran successfully.</p> <p>The files named <code>DATASET.METHOD.h5ad</code> are encoded with anndata <code>v0.7.0</code> (be careful as they are not readable with previous versions) and contain 100 embeddings each. The embeddings are in the <code>obsm</code> attribute of the object. All the embeddings can be listed with the <code>obsm_keys()</code> method. The name of the embedding contains the parameters used to generate that embedding and are written like that <code>method=zinbwave.dims=10.epsilon=1000.features=300.gene_covariate=0</code>.</p> <p> </p> <p>For questions on this dataset please contact fraimundo@google.com</p>
SSH data set used for Rossby Wave Analysis, extraction from ORCA12.L46-MJM189 DRAKKAR simulation
<p>This data set corresponds to the Sea Surface Heigh (SSH) silmulated by the NEMO ocean circulation model, under the ORCA12.L46-MJM189 configuration, developped in the frame of the DRAKKAR project. This particular data set is an extraction from the native numerical grid, covering the area between 38N and 40N in the North Altantic ocean, for the period 1970 to 2015. The data are concatenated in a single file with 5-days average of SSH. The corresponding metrics for this sub domain are also present in this netcdf file. This subset was used in Watelet et al. (2020) submitted paper, dealing with Rossby waves analysis.</p>
Data and statistical analysis for: Bacterial nanotubes are a manifestation of cell death
<p>Contains all data and code to reproduce the statistical analysis in the Supplementary file 2 for the paper "Bacterial nanotubes are a manifestation of cell death" to be published in Nature Communications.</p> <p><strong>Contents:</strong></p> <ul> <li>statistical_analysis.Rmd is the main document written in R Markdown</li> <li>statistical_analysis.html is a compiled version of statistical_analysis.Rmd showing all the computed results.</li> <li>The Source data.xls file contains raw data used for the analysis - the sheet names indiciate the figure they refer to. See the main file for code that can read the data.</li> </ul> <p>The code can also be accessed at <a href="https://github.com/cas-bioinf/nanotubes-death">https://github.com/cas-bioinf/nanotubes-death</a></p>
2019 H1B Petition Data Analysis
<p>In the United States, H1B is one of the Non-immigrant visas. It allows foreign workers to work in the States temporarily in specialty occupations such as IT, medical or business industries. It also requires workers to have a bachelor's or higher degree. Until 2020, more than 580K people are working under H1B in the USA. Upon the current immigration laws, the United States Citizenship and Immigration Services issues 65K H1B visas per year. </p> <p>This project aims to help the users understand how H1B visa affects the IT industry in the United States. An overview of H1B visa by positions, salary, top 10 sponsors, and worksite states. Last, this project dives into an in-depth analysis of the denied H1B cases by job titles and salaries. </p> <p>Contents:</p> <p>- Overview of Top 30 Job Titles With Most Filed H1B Cases<br> - Number of IT Job H1B Cases by States<br> - Top 10 Sponsors With Most Filed H1B Cases (IT Job ONLY)<br> - Number of Top 12 IT Job Employment Comparison : H1B vs National <br> - Avg Salary Comparison: H1B vs National <br> - H1B Case Status Percentages <br> - Deneid H1B Cases<br> - Salary vs Job titles <br> - 5 Numbers Summary: Certified vs Denied <br> - Certified Rate vs Avg Salary : Top 10 Sponsors<br> - Relationship between Salary and Certified Rate</p>
Tara Pacific 18S-based coral host genetic analysis data release version 1
<p>This dataset contains 4 tables and 3 sets of figures related to the primary analysis of the 18S metabarcoding sequencing output. This dataset is only concerned with the identity of the coral host (i.e. not additional protist diversity). The samples included in this dataset have a 'sample-material_label' value of 'CORAL' and 'sampling-protocol_label' value of 'SEQ-CS4L'. They represent the coral samples collected at all 32 of the islands visited in the Tara Pacific expedition.</p>
Appendix: Data Analysis and Machine Learning Experiments
<p>The plots and statistics generated for the data analysis are given in this data set.<br> </p> <p>Furthermore, this data set contains the models, feature sets, scaler, prediction results and visualizations for the machine learning experiments conducted.</p> <ol> <li>Reproduction Experiment</li> <li>Multiple Commit Thresholds Experiment</li> <li>Imbalanced Training Experiment</li> </ol>
Data and analysis scripts associated with the paper 'Long-term experimental evolution of HIV-1 reveals effects of environment and mutational history''
<p><em>Eva Bons, Christine Leemann, Karin J. Metzner, Roland R. Regoes</em></p> <p>This repository contains all the data and analysis scripts associated with the paper 'Long-term experimental evolution of HIV-1 reveals effects of environment and mutational history'</p> <p>See the readme after unpacking the .zip for a description of the files</p>
Training dataset: DIA data analysis of a HEK/Ecoli Spike-in dataset using OpenSwathWorkflow
<p>The eight raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 °C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in two different ratios and four replicates of each Spike/in ratio were measured:</p> <p>Sample HEK E.coli MS method<br> Sample1 2.5 0.15 DIA<br> Sample2 2.5 0.15 DIA<br> Sample3 2.5 0.15 DIA<br> Sample4 2.5 0.15 DIA<br> Sample5 2.5 0.80 DIA<br> Sample6 2.5 0.80 DIA<br> Sample7 2.5 0.80 DIA<br> Sample8 2.5 0.80 DIA</p> <p>Additionally, iRT peptides were added and 1µg of each samples was measured using data independent acquisition with a Q-Exactive Plus mass spectrometer. Briefly, a scan range from 400-1000 m/Z was first covered by an MS1 scan followed by 25 consecutive MS2 scans (each 24 m/z broad). In the next cycle another MS1 scan was acquired followd by 26 MS2 scans (also 24m/z broad) in which the window centers were shifted by 50% compared to the previous cycle of MS2 scans. The resulting raw files contain overlapping MS2 scans.</p> <p>Besides the eight raw files, we uploaded a spectral library, a transition list for the iRT peptides as well as an sample annotation file.<br> Additionally, we uploaded the Galaxy PyProphet score training result files: PyProphet score report and PyProphet score.</p>
Data and scripts for the analysis of fruit flies' species distributions in Reunion island
<p>Data and scripts supporting the analyses of the joint species distributions of eight Tephritids species in La Réunion island.</p>
Supplementary data: What millimeter-wavelength radar reflectivity reveals about snowfall: An information-centric analysis
<p>This dataset includes supplementary data used in the analyses described in Wood, N. B., and T. S. L'Ecuyer, 2020: What millimeter-wavelength radar reflectivity reveals about snowfall: An information-centric analysis. Atmospheric Measurement Techniques, doi:10.5194/amt-2020-216.</p>
Data set 1 discourse analysis BRAD research project
<p>Discourse analysis data set with excerpts of press articles generated in the coding (coded with keywords ‘Brexit’ and ‘deportations’). This data set connects to the WP3 of the BRAD research project.</p>
Experiment data in support of "Segmentation analysis and the recovery of queuing parameters via the Wasserstein distance: a study of administrative data for patients with chronic obstructive pulmonary disease"
<p>This archive contains a ZIP archive, `data.zip`, that itself contains the data used in the final sections of the paper. The remainder of the paper's supporting files are available at <a href="https://github.com/daffidwilde/copd-paper/">github.com/daffidwilde/copd-paper/</a></p> <p>The ZIP archive is structured as follows:</p> <ul> <li>There is a directory, `wasserstein`, for the parameter sweep described in the model construction section of the paper. Its contents are: (i) a file, `main.csv`, describing each parameter and their maximal Wasserstein distance to the observed data, and (ii) three directories, `best`, `median` and `worst`, each containing the simulated queuing results (in `main.csv`) from that sweep with the best, median and worst found parameter sets, respectively (in `params.txt`).</li> <li>The remaining three directories correspond to the experiments conducted in the final section of the paper. Each directory contains two files: (i) `system_times.csv` which holds trial parameters and system time records for every patient to pass through the model in that experiment, and (ii) `utilisations.csv` which holds trial parameters and utilisations for each server in the model for that experiment.</li> </ul>
Data from: An analysis of mating biases in trees
<p><span><span><span><span><span><span><span><span><span><span><span>Assortative mating is a deviation from random mating based on phenotypic similarity. As it is much better studied in animals than in plants, we investigate for trees whether kinship of realized mating pairs deviates from what is expected from the set of potential mates and use this information to infer mating biases that may result from kin recognition and/or assortative mating. Our analysis covers twenty species of trees for which microsatellite data is available for adult populations (potential mates) as well as seed arrays. We test whether mean relatedness of observed mating pairs deviates from null expectations that only take pollen dispersal distances into account (estimated from the same dataset). This allows to identify elevated as well as reduced kinship among realized mating pairs, indicative of positive and negative assortative mating, respectively. The test is also able to distinguish elevated biparental inbreeding that occurs solely as a result of related pairs growing closer to each other from further assortativeness.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>Assortative mating in trees appears potentially common but not ubiquitous: nine data sets show mating bias with elevated inbreeding, nine do not deviate significantly from the null expectation, and two show mating bias with reduced inbreeding.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>While our datasets lack direct information on phenology, our investigation of the phenological literature for each species identifies flowering phenology as a potential driver of positive assortative mating (leading to elevated inbreeding) in trees. Since active kin recognition provides an alternative hypothesis for these patterns, we encourage further investigations on the processes and traits that influence mating patterns in trees.</span></span></span></span></span></span></span></span></span></span></span></p>
Data files for INS data analysis course at github.com/pace-neutrons/edatc
<p>Large data files for use in the inelastic neutron training course hosted at https://github.com/pace-neutrons/edatc</p> <p>Included are:</p> <ul> <li>UPd3 measured on MERLIN (M. D. Le et al., 2008)</li> <li>bcc-Iron measured on MAPS (T. G. Perring et al., 2010)</li> <li>CuGeO3 measured on MERLIN (H. C. Walker et al., 2014)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.